forum
Can't fine tune Phi 4 using GRPO
Solved
Qwen3 Coder Next on 2x RTX 3090
GLM-4.7-Flash not starting on vLLM
Qwen3-Coder-Next on 4x5060ti OOM
Open
gpt-oss RL takes 41H in DGX rather than 4H as given in docs
Open
install unsloth on colab
Solved
Unsloth Quant Request: Step 3.5 Flash
How do I finetune QLoRA on direct Windows installation?
Open
How to inference fine-tuned GPT-OSS-20B saved in mxfp4?
Open
How to enable reasoning_content for Kimi-k2.5?
Set non-thinking mode during training Qwen3 GRPO
Fine-tuning takes very long on H200 and 1B params
Solved
Docker image for B200?
Data collection framework best practices?
Suggested dataset format for full-FT
Open
qwen-image-2512 comfy-ui vs stable-diffusion.cpp
Open
Multi turn environments like OpenSpiel using RL
How to train vision model with IterableDataset?
Solved
Ollama glm 4.7 flash reap gguf
[Unsloth Docker] Can't import 'FastSentenceTransformer'
Cannot install unsloth
GLM-4.7 degraded Quality
GLM-4.7-FLASH WEIRDNESS IN OLLAMA
Open
Unsloth GLM-4.7 Flash Finetuning & RL Request
GRPO on A100 80GB GPU resulting in 0.02 it/s.
Issue with Multi-GPU Setup for RL Task Using GRPO
Is it possible to run GLM-4.7-Flash on 8GB VRAM + 32GB RAM locally or on a Colab notebook 16GB VRAM?
Open
Unsloth Quantization Request: GLM-4.7 Flash
Extraction with Fine-tuned LLM
grad_norm: 0 and LR: 0
80gb VRAM usage for Gemma 3 4b 32k Context / Max Sequence
Open
Can't get Unsloth to work
Support for transformers v5
TRL custom rollout func support
Enterprise plan
Endless thinking with unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF
How do I only upload the LORA adapters to huggingface using Unsloth?
Nemotron UD 2 XL and UD 3 XL are the same size?
Can't create conda environment for unsloth because of xformers [SOLVED]
Unsloth Quantization Request: GLM-4.7-REAP (from Cerebras)
nan loss: FT Qwen3 QAT
No config file found - are you sure the `model_name` is correct?
Open
Fine tuning nemotron 3 nano with load_in_4bit doesn't work
CISPO / SAPO Loss Support for GRPO?
How can I cleanly fine tune unsloth/Ministral-3-3B-Instruct-2512 without vision ?
rubbish response from unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF:UD-Q4_K_XL
unable to covert Qwen4b-instruct to GGUF
Open
RLHF need help chosing correct methods
Open
unsloth/Qwen3-VL-30B-A3B-Instruct quantization issue
GRPO gpt-oss-20b 100k context OOM & TRL/vLLM Version Hell