forum

1495 threads · Page 27 of 30

How should I evaluate a fine-tuned model to see if it does better compared to the base model ? 7 messages
The training loss doesnt change much 79 messages
How to structure the data before training ? 71 messages
did unsloth support bnb_config? 4 messages
Base model gets changed when merging 9 messages
How to choose base model when merging? 6 messages
Model taking too long to get a response for a prompt 21 messages
Data Formatting 24 messages
Solved
Your opinion about parameters to get more concise answers 10 messages
Gemma 27B DPO Impossible due to defaulting to `/models/llama.py` for PeftModelForCausalLM_fast_forwa 2 messages
Gemma 2: Can't merge 27B with added tokens 12 messages
Open
How do I fine tune llama 3 as a complete noob? 3 messages
Open
nvidia-smi returned non-zero exit status 18 19 messages
Why does Full-Fine Tuning work with Unsloth models? 13 messages
Sawtooth loss curve when using `group_by_length`? 31 messages
Open
Fine tuning phi-3-mini on Cognitive Computations - dolphin datasets. 17 messages
Gemma 2: torch._dynamo.exc.BackendCompilerFailed 2 messages
Solved
How to create FastLanguageModel from already existing adapters 2 messages
EOS Token error 4 messages
Open
kaggle error 3 messages
Open
Recommendations On Document Chunking 11 messages
Jupyter notebook transformers training outputs file is empty 38 messages
Using other `transformers` models alongside an Unsloth model 4 messages
Solved
Merge of model, efficiency, language and function calling 3 messages
Open
phi-3-small finetuning support 2 messages
Open
continued pretraining notebook EoS token missing? 3 messages
Open
fine tuned llama3 7b instruct, when importing into ollama it is all a complete hallucination 2 messages
Continued pre-training on GCP 2 messages
How do I format a dataset? 2 messages
Continued Pretraining Question 32 messages
how to perform DPO on gemma/llama3 5 messages
Deciding how to format dataset 7 messages
Open
Can unsloth work with DITTO ? 3 messages
Flash Atten Loading Forever 28 messages
Open
Can i retrain my model? 7 messages
Open
Testing model performance before training 4 messages
Mixtral-8x7b not NotImplementedError! 8 messages
Qwen2: Some weights of were not initialized from the model checkpoint 12 messages
Model not working on replicate 4 messages
saving checkpoints with config files 12 messages
qwen2-1.5b crashing when starting training on kaggle 12 messages
continued pretraining dataset preperation 19 messages
from unsloth import FastLanguageModel error 5 messages
Open
max_seq_length = 2048, can this value be set to bigger number for model unsloth/llama-3-8b-bnb-4bit 7 messages
When to use rsLoRA/use_rslora? Relation to alpha? 12 messages
Open
ShareGPT is "from":"human" and "from":"gpt" not "from":"user" and "from":"assistant" 3 messages
EOS token never being predicted 10 messages
Open
Resuming CPT training results in high loss 51 messages
finetune model repeate my question over and over 7 messages
Open
Language translation 5 messages