hh-rlhf

(★ 1,857)

Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"

  • .gitattributes
  • LICENSE
  • README.md
// repository documentation