Learning-Supervised-Finetuning-and-Reinforcement-Learning-on-Math-LLMs

(★ 10)

Implementing Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) for Qwen3 and DeepSeek-Math models. Includes experimental code, training logs, and insights on improving mathematical reasoning in LLMs.

Learning-Supervised-Finetuning-and-Reinforcement-Learning-on-Math-LLMs Latest Version Download

Download Latest Version (.zip)
// repository documentation