Learning-Supervised-Finetuning-and-Reinforcement-Learning-on-Math-LLMs
Implementing Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) for Qwen3 and DeepSeek-Math models. Includes experimental code, training logs, and insights on improving mathematical reasoning in LLMs.
// repository documentation
Was this content helpful?
(0 ratings)
