forum
How should I evaluate a fine-tuned model to see if it does better compared to the base model ?
The training loss doesnt change much
How to structure the data before training ?
did unsloth support bnb_config?
Base model gets changed when merging
How to choose base model when merging?
Model taking too long to get a response for a prompt
Data Formatting
Solved
Your opinion about parameters to get more concise answers
Gemma 27B DPO Impossible due to defaulting to `/models/llama.py` for PeftModelForCausalLM_fast_forwa
Gemma 2: Can't merge 27B with added tokens
Open
How do I fine tune llama 3 as a complete noob?
Open
nvidia-smi returned non-zero exit status 18
Why does Full-Fine Tuning work with Unsloth models?
Sawtooth loss curve when using `group_by_length`?
Open
Fine tuning phi-3-mini on Cognitive Computations - dolphin datasets.
Gemma 2: torch._dynamo.exc.BackendCompilerFailed
Solved
How to create FastLanguageModel from already existing adapters
EOS Token error
Open
kaggle error
Open
Recommendations On Document Chunking
Jupyter notebook transformers training outputs file is empty
Using other `transformers` models alongside an Unsloth model
Solved
Merge of model, efficiency, language and function calling
Open
phi-3-small finetuning support
Open
continued pretraining notebook EoS token missing?
Open
fine tuned llama3 7b instruct, when importing into ollama it is all a complete hallucination
Continued pre-training on GCP
How do I format a dataset?
Continued Pretraining Question
how to perform DPO on gemma/llama3
Deciding how to format dataset
Open
Can unsloth work with DITTO ?
Flash Atten Loading Forever
Open
Can i retrain my model?
Open
Testing model performance before training
Mixtral-8x7b not NotImplementedError!
Qwen2: Some weights of were not initialized from the model checkpoint
Model not working on replicate
saving checkpoints with config files
qwen2-1.5b crashing when starting training on kaggle
continued pretraining dataset preperation
from unsloth import FastLanguageModel error
Open
max_seq_length = 2048, can this value be set to bigger number for model unsloth/llama-3-8b-bnb-4bit
When to use rsLoRA/use_rslora? Relation to alpha?
Open
ShareGPT is "from":"human" and "from":"gpt" not "from":"user" and "from":"assistant"
EOS token never being predicted
Open
Resuming CPT training results in high loss
finetune model repeate my question over and over
Open
Language translation