forum
ValueError: bf16 mixed precision requires PyTorch >= 1.10 and a supported device.bf16
windows 11 ,rtx 3090, qwen3 moe
Unsloth's gradient checkpointing crashing during GRPO training on evaluation
Gemma 3 regression?
help
Unable to convert model from HF to GGUF format
name 'psutil' is not defined
Solved
text classification
Convert finetuned ministral-3b to gguf failed
Extremely High VRAM Usage with Ministral 3 3B VLM QLoRA
Open
Model works with lora, but if lora is merged, the model starts making gibberish output
Open
Ministral model can't load - no Config file
Solved
Packing Compatibility with Ministral
Open
Building docker image in DGX Spark
Open
Pytorch + xformers compatibility issues on DGX Spark docker image build
Open
Can QAT aware trained phi-4 model using unsloth and torchao collab packages be converted to gguf?
Open
Confused on model differences
Nemotron-3-Nano-30B-A3B-FP8
train_on_responses_only does not work when using "packing"
Open
Does packing not work with multigpu?
Open
Error when upgrading to the latest unsloth version
Solved
does unsloth support transformers v5?
Qwen3VL (or any VLM) OOM from large dataset
qwen3 vl 235b crash
Fine-tune on multi GPUs
Unable to export Ministral model to GGUF
Open
Documentation Improvement: Common Mistakes
Open
Trouble pushing unsloth/Ministral-3-3B-Instruct-2512 to my HF account
VyvoTTS: ImportError: libnvshmem_host.so.3
Open
Wondering what to consider if training a binary classifier
Can not finetune whisper model with error about AssertionError: expected size 128==128, stride 1500
How to serve unsloth/gpt-oss-120b-unsloth-bnb-4bit on vllm?
Open
Synth Reasoning Gen
Open
Finetuning Ministral-3-3B-Instruct-2512
Open
Loading ft version (Gelato) of Qwen3VL into unsloth
Fast Inference vllm GRPO tuning on Qwen3VL
Open
What is the purchase price of Unsloth Enterprise
Open
Hosting 4bit SFT/RFT using Vllm on DGX Sparks
PaddleOCR-VL sft error: `Unsloth: Failed to make input require gradients!`
Solved
Is this the correct way load a trained model?
Image Size for finetuning qwen3-VL-2B
Open
eval loss and train loss always identical
answer cropped after finetune
Batch size increase causes training time to increase (inverse scaling) on H100
Open
How to use Fsdp via unsloth
Open
Problem with GGUF conversion
flashinfer issues
Solved
Use qwen3vl thinking model to get final output after thinking
Can I fine-tune Gemma-3-12B on a single RTX 3060 12 GB with Unsloth?
Subject: Request for Guidance on Fine-Tuning GPT-OSS-20B Using COT + Structured DNA Reasoning Prompt