qwen2.5-1.5b-grpo-starter

(★ 9)

A simple example of using GRPO (Group Relative Policy Optimization) trainer from TRL library to fine-tune Qwen2.5-1.5B model on the TL;DR summarization dataset.

qwen2.5-1.5b-grpo-starter 최신버젼 다운로드

최종 버전 다운로드 (.zip)
// repository documentation