qwen2.5-1.5b-grpo-starter
A simple example of using GRPO (Group Relative Policy Optimization) trainer from TRL library to fine-tune Qwen2.5-1.5B model on the TL;DR summarization dataset.
qwen2.5-1.5b-grpo-starter 최신버젼 다운로드
최종 버전 다운로드 (.zip)// repository documentation
Was this content helpful?
(0 ratings)
