RLHF_in_notebooks
RLHF (Supervised fine-tuning, reward model, and PPO) step-by-step in 3 Jupyter notebooks
// repository documentation
Was this content helpful?
(0 ratings)