RLHF_in_notebooks

★ 252 Open GitHub ↗

RLHF (Supervised fine-tuning, reward model, and PPO) step-by-step in 3 Jupyter notebooks

// repository documentation