Learning-Supervised-Finetuning-and-Reinforcement-Learning-on-Math-LLMs

Implementing Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) for Qwen3 and DeepSeek-Math models. Includes experimental code, training logs, and insights on improving mathematical reasoning in LLMs.

// repository documentation