OPD
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
File Explorer
Download Latest Version (.zip)- test.json
- test.parquet
- test.parquet_first_row.json
- test.json
- test.parquet
- test.parquet_first_row.json
- test.json
- test.parquet
- test.parquet_first_row.json
- train.json
- train.parquet
- train.parquet_first_row.json
- train.json
- train.parquet
- train.parquet_first_row.json
- test.json
- test.parquet
- test.parquet_first_row.json
- train.json
- train.parquet
- train.parquet_first_row.json
- test.json
- test.parquet
- test.parquet_first_row.json
- train.json
- train.jsonl
- train.parquet
- train.parquet_first_row.json
- test.json
- test.parquet
- test.parquet_first_row.json
- test.json
- test.parquet
- test.parquet_first_row.json
- test.json
- test.parquet
- test.parquet_first_row.json
- test_pre.json
- test.json
- test.parquet
- test.parquet_first_row.json
- test.json
- test.parquet
- test.parquet_first_row.json
- train.json
- train.parquet
- train.parquet_first_row.json
- dapo-math-17k-processed.parquet
- dapo-math-17k.parquet
- DeepMath_deduped.parquet
- OpenThoughts3_opd.parquet
- opd_teaser.png
- 1-bug-report.yml
- 2-feature-request.yml
- config.yml
- docker.yml
- docs.yml
- label_issue.yml
- publish.yml
- tests.yml
- tests_cuda.yml
- tests_npu.yml
- CODE_OF_CONDUCT.md
- CONTRIBUTING.md
- copilot-instructions.md
- instructions-v0.md
- instructions-v1.md
- PULL_REQUEST_TEMPLATE.md
- SECURITY.md
- serpapi.svg
- warp.jpg
- colab.svg
- discord.svg
- dsw.svg
- lab4ai.svg
- online.svg
- logo.png
- 1.jpg
- 1.mp3
- 1.mp4
- 2.avi
- 2.jpg
- 2.wav
- 3.flac
- 3.jpg
- 3.mp4
- 4.mp3
- 4.mp4
- alpaca_en_demo.json
- alpaca_zh_demo.json
- c4_demo.jsonl
- dataset_info.json
- dpo_en_demo.json
- dpo_zh_demo.json
- glaive_toolcall_en_demo.json
- glaive_toolcall_zh_demo.json
- identity.json
- kto_en_demo.json
- mllm_audio_demo.json
- mllm_demo.json
- mllm_video_audio_demo.json
- mllm_video_demo.json
- README.md
- README_zh.md
- reason_tool_use_demo_50.jsonl
- v1_dpo_demo.jsonl
- v1_dpo_demo.yaml
- v1_sft_demo.jsonl
- v1_sft_demo.yaml
- wiki_demo.txt
- docker-compose.yml
- Dockerfile
- Dockerfile.base
- Dockerfile.megatron
- README.md
- docker-compose.yml
- Dockerfile
- docker-compose.yml
- Dockerfile
- lang-switcher.css
- switcher.js
- custom-kernels.md
- fused-operators.md
- triton.md
- deepspeed.md
- fsdp.md
- parallel-dp-tp-ep-sp-cp.md
- lora.md
- quantization.md
- data-processing.md
- data-engine.md
- model-engine.md
- trainer.md
- initialization.md
- kernels.md
- rendering.md
- data-plugins.md
- data-argument.md
- model-argument.md
- sample-argument.md
- training-argument.md
- deploy.md
- dpo.md
- sft.md
- conf.py
- getting-started.md
- index.rst
- installation.md
- llamaboard-web-ui.md
- custom-kernels.md
- fused-operators.md
- triton.md
- deepspeed.md
- fsdp.md
- parallel-dp-tp-ep-sp-cp.md
- lora.md
- quantization.md
- data-processing.md
- data-engine.md
- model-engine.md
- trainer.md
- initialization.md
- kernels.md
- rendering.md
- data-plugins.md
- data-argument.md
- model-argument.md
- sample-argument.md
- training-argument.md
- deploy.md
- dpo.md
- sft.md
- conf.py
- getting-started.md
- index.rst
- installation.md
- llamaboard-web-ui.md
- conf.py
- make.bat
- Makefile
- requirements.txt
- fsdp2_config.yaml
- fsdp_config.yaml
- fsdp_config_multiple_nodes.yaml
- fsdp_config_offload.yaml
- qwen3_full_sft_fsdp2.yaml
- qwen3moe_full_sft_fsdp.yaml
- qwen3vlmoe_full_sft_fsdp2.yaml
- qwen3vlmoe_lora_sft_fsdp.yaml
- ds_z0_config.json
- ds_z2_autotp_config.json
- ds_z2_config.json
- ds_z2_offload_config.json
- ds_z3_config.json
- ds_z3_fp8_config.json
- ds_z3_offload_config.json
- qwen2_full_sft.yaml
- llama3_full_sft.yaml
- llama2_full_asft.yaml
- qwen2_full_asft.yaml
- llama3_full_sft.yaml
- qwen2_full_sft.yaml
- qwen25_05b_eaft_full.yaml
- llama3_fp8_deepspeed_sft.yaml
- llama3_fp8_fsdp_sft.yaml
- llama3_lora_sft.yaml
- train.sh
- llama3_full_sft.yaml
- expand.sh
- llama3_freeze_sft.yaml
- llama3_lora_sft.yaml
- llama3_full_sft.yaml
- tokens_cfg.yaml
- qwen2_full_sft.yaml
- llama3_lora_predict.yaml
- llama3_oft_sft.yaml
- qwen2_5vl_oft_sft.yaml
- init.sh
- llama3_lora_sft.yaml
- llama3_oft_sft_awq.yaml
- llama3_oft_sft_bnb_npu.yaml
- llama3_oft_sft_gptq.yaml
- qwen3.yaml
- qwen3_full_sft.yaml
- qwen3_lora_sft.yaml
- qwen3vl.yaml
- deepseek2_lora_sft_kt.yaml
- deepseek3_kt.yaml
- deepseek3_lora_sft_kt.yaml
- qwen3moe_lora_sft_kt.yaml
- DeepSeek-V2-Chat-sft-amx.yaml
- DeepSeek-V2-Chat.yaml
- DeepSeek-V2-Lite-Chat-sft-amx-multi-gpu.yaml
- DeepSeek-V2-Lite-Chat-sft-amx.yaml
- DeepSeek-V2-Lite-Chat-sft.yaml
- DeepSeek-V2-Lite-Chat.yaml
- DeepSeek-V3-Chat-amx.yaml
- DeepSeek-V3-Chat-sft-amx-multi-gpu-4.yaml
- DeepSeek-V3-Chat-sft-amx-multi-gpu.yaml
- DeepSeek-V3-Chat-sft-amx.yaml
- Qwen3Moe-sft-amx.yaml
- deepseek2_lora_sft_kt.yaml
- deepseek3_lora_sft_kt.yaml
- qwen3moe_lora_sft_kt.yaml
- qwen2_vl_full.yaml
- qwen3_moe_full.yaml
- qwen3_full_sft.yaml
- qwen3_gptq.yaml
- qwen3_lora_sft.yaml
- qwen3vl_lora_sft.yaml
- qwen3_base_full_sft.yaml
- qwen3_full_sft.yaml
- qwen3vl_full_sft.yaml
- qwen3_lora_dpo.yaml
- qwen3_lora_kto.yaml
- qwen3_lora_pretrain.yaml
- qwen3_lora_reward.yaml
- qwen3_lora_sft.sh
- qwen3_lora_sft.yaml
- qwen3_lora_sft_ds3.yaml
- qwen3_lora_sft_ray.yaml
- qwen3_preprocess.yaml
- qwen3vl_lora_dpo.yaml
- qwen3vl_lora_sft.yaml
- llama3_lora_sft_aqlm.yaml
- llama3_lora_sft_awq.yaml
- llama3_lora_sft_gptq.yaml
- qwen3_lora_sft_bnb_npu.yaml
- qwen3_lora_sft_otfq.yaml
- train_freeze_sft.yaml
- train_full_deepspeed.yaml
- train_full_fsdp2.yaml
- export_lora.yaml
- train_lora_sft.yaml
- quantization.yaml
- README.md
- README_zh.md
- adam-mini.txt
- apollo.txt
- aqlm.txt
- badam.txt
- bitsandbytes.txt
- deepspeed.txt
- dev.txt
- eetq.txt
- fp8-te.txt
- fp8.txt
- galore.txt
- gptq.txt
- hqq.txt
- liger-kernel.txt
- metrics.txt
- minicpm-v.txt
- npu.txt
- openmind.txt
- sglang.txt
- swanlab.txt
- vllm.txt
- test_image.py
- test_toolcall.py
- llamafy_baichuan2.py
- llamafy_qwen.py
- tiny_llama4.py
- tiny_qwen3.py
- cal_flops.py
- cal_lr.py
- cal_mfu.py
- cal_ppl.py
- length_cdf.py
- bench_qwen.py
- eval_bleu_rouge.py
- hf2dcp.py
- llama_pro.py
- loftq_init.py
- megatron_merge.py
- pissa_init.py
- qwen_omni_merge.py
- vllm_infer.py
- __init__.py
- app.py
- chat.py
- common.py
- protocol.py
- __init__.py
- base_engine.py
- chat_model.py
- hf_engine.py
- kt_engine.py
- sglang_engine.py
- vllm_engine.py
- __init__.py
- feedback.py
- pairwise.py
- pretrain.py
- processor_utils.py
- supervised.py
- unsupervised.py
- __init__.py
- collator.py
- converter.py
- data_utils.py
- formatter.py
- loader.py
- mm_plugin.py
- parser.py
- template.py
- tool_utils.py
- __init__.py
- evaluator.py
- template.py
- __init__.py
- constants.py
- env.py
- logging.py
- misc.py
- packages.py
- ploting.py
- __init__.py
- data_args.py
- evaluation_args.py
- finetuning_args.py
- generating_args.py
- model_args.py
- parser.py
- training_args.py
- __init__.py
- attention.py
- checkpointing.py
- embedding.py
- ktransformers.py
- kv_cache.py
- liger_kernel.py
- longlora.py
- misc.py
- mod.py
- moe.py
- packing.py
- quantization.py
- rope.py
- unsloth.py
- valuehead.py
- visual.py
- __init__.py
- adapter.py
- loader.py
- patcher.py
- __init__.py
- muon.py
- __init__.py
- __init__.py
- ktrainer.py
- trainer.py
- workflow.py
- __init__.py
- trainer.py
- workflow.py
- __init__.py
- trainer.py
- workflow.py
- __init__.py
- ppo_utils.py
- trainer.py
- workflow.py
- __init__.py
- trainer.py
- workflow.py
- __init__.py
- metric.py
- trainer.py
- workflow.py
- __init__.py
- metric.py
- trainer.py
- workflow.py
- __init__.py
- callbacks.py
- fp8_utils.py
- test_utils.py
- trainer_utils.py
- tuner.py
- __init__.py
- helper.py
- interface.py
- profiler.py
- __init__.py
- arg_parser.py
- arg_utils.py
- data_args.py
- model_args.py
- sample_args.py
- training_args.py
- __init__.py
- batching.py
- callback.py
- inference_engine.py
- rendering.py
- __init__.py
- base_sampler.py
- base_trainer.py
- data_engine.py
- model_engine.py
- __init__.py
- converter.py
- loader.py
- __init__.py
- npu_fused_moe.py
- npu_swiglu.py
- __init__.py
- npu_rms_norm.py
- __init__.py
- npu_rope.py
- __init__.py
- __init__.py
- base.py
- interface.py
- registry.py
- __init__.py
- qwen3.py
- qwen3_nothink.py
- __init__.py
- add_token.py
- initialization.py
- peft.py
- quantization.py
- rendering.py
- __init__.py
- vllm.py
- __init__.py
- deepspeed.py
- fsdp2.py
- hub.py
- __init__.py
- batching.py
- lr_scheduler.py
- optimizer.py
- __init__.py
- cli_sampler.py
- __init__.py
- dpo_trainer.py
- rm_trainer.py
- sft_trainer.py
- __init__.py
- constants.py
- dtype.py
- env.py
- helper.py
- logging.py
- objects.py
- packages.py
- plugin.py
- pytest.py
- types.py
- __init__.py
- launcher.py
- __init__.py
- chatbot.py
- data.py
- eval.py
- export.py
- footer.py
- infer.py
- top.py
- train.py
- __init__.py
- chatter.py
- common.py
- control.py
- css.py
- engine.py
- interface.py
- locales.py
- manager.py
- runner.py
- __init__.py
- cli.py
- launcher.py
- api.py
- train.py
- webui.py
- test_feedback.py
- test_pairwise.py
- test_processor_utils.py
- test_supervised.py
- test_unsupervised.py
- test_collator.py
- test_converter.py
- test_formatter.py
- test_loader.py
- test_mm_plugin.py
- test_template.py
- test_chat.py
- test_sglang.py
- test_train.py
- test_eval_template.py
- test_add_tokens.py
- test_attention.py
- test_checkpointing.py
- test_misc.py
- test_packing.py
- test_visual.py
- test_base.py
- test_freeze.py
- test_full.py
- test_lora.py
- test_pissa.py
- test_sft_trainer.py
- check_license.py
- conftest.py
- version.txt
- test_interface.py
- test_args_parser.py
- test_batching.py
- test_rendering.py
- test_data_engine.py
- test_model_loader.py
- test_converter.py
- test_init_plugin.py
- test_kernel_plugin.py
- test_peft.py
- test_quantization_plugin.py
- test_fsdp2.py
- test_cli_sampler.py
- test_fsdp2_sft_trainer.py
- conftest.py
- .dockerignore
- .env.local
- .gitattributes
- .gitignore
- .pre-commit-config.yaml
- =15.0.0
- CITATION.cff
- constraints.txt
- LICENSE
- Makefile
- MANIFEST.in
- pyproject.toml
- README.md
- README_zh.md
- dedup_deepmath.py
- vllm_rollout.py
- test.parquet
- test.parquet
- test.parquet
- test.parquet
- test.parquet
- test.parquet
- test.parquet
- test.parquet
- test.parquet
- gen_vllm.py
- grade.py
- utils.py
- config.yaml
- bug-report.yml
- config.yml
- feature-request.yml
- e2e_eval_aime24.yml
- e2e_ppo_trainer.yml
- e2e_ppo_trainer_megatron_sglang.yml
- e2e_prime.yml
- e2e_spin.yml
- e2e_sppo.yml
- check-pr-title.yml
- checkpoint_converter.yml
- cpu_unit_tests.yml
- doc.yml
- docker-build-ascend-a2.yml
- docker-build-ascend-a3.yml
- docker-validate-ascend.yml
- e2e_ascend.yml
- e2e_dapo.yml
- e2e_fully_async_policy.yml
- e2e_genrm_remote.yml
- e2e_one_step_off_policy.yml
- e2e_ppo_trainer.yml
- e2e_ppo_trainer_megatron_sglang.yml
- e2e_ppo_trainer_megatron_sglang_2.yml
- e2e_ppo_trainer_megatron_vllm.yml
- e2e_ppo_trainer_megatron_vllm_2.yml
- e2e_sft.yml
- gpu_unit_tests.yml
- model.yml
- pre-commit.yml
- README.md
- reward_model_sglang.yml
- reward_model_vllm.yml
- sanity.yml
- scorecard.yml
- secrets_scan.yml
- sgl.yml
- type-coverage-check.yml
- vllm.yml
- CODEOWNERS
- dependabot.yml
- PULL_REQUEST_TEMPLATE.md
- Dockerfile.ascend_8.2.rc1_a2
- Dockerfile.ascend_8.2.rc1_a3
- Dockerfile.app.sglang.vllm.mcore0.12
- Dockerfile.app.sglang.vllm.mcore0.12.deepep
- Dockerfile.app.sglang.vllm.mcore0.13.preview
- Dockerfile.app.vllm.mcore0.12
- Dockerfile.app.vllm.mcore0.12.deepep
- Dockerfile.app.vllm.mcore0.13.preview
- Dockerfile.base
- README.md
- Dockerfile.app.sglang0.4.10.post2.mcore0.13
- Dockerfile.app.sglang0.4.9.post6.mcore0.13
- Dockerfile.app.vllm.mcore0.13
- Dockerfile.app.vllm.mcore0.15
- Dockerfile.base.torch2.7.1
- README.md
- Dockerfile.app.sglang.mcore0.12
- Dockerfile.app.sglang.mcore0.13.preview
- Dockerfile.base
- README.md
- Dockerfile.app.sglang.megatron
- Dockerfile.base
- README.md
- Dockerfile.app.sglang
- Dockerfile.base
- Dockerfile.vllm011.mcore_gpt-oss
- Apptainerfile.rocm
- Dockerfile.extention.awsefa
- Dockerfile.ngc.vllm
- Dockerfile.ngc.vllm0.8
- Dockerfile.ngc.vllm0.8.sagemaker
- Dockerfile.rocm
- Dockerfile.rocm7
- Dockerfile.rocm_verl-0.3.0.post1
- Dockerfile.rocm_verl-0.4.1
- Dockerfile.sglang
- Dockerfile.stable.vllm011
- Dockerfile.vemlp.vllm.te
- Dockerfile.vllm.sglang.megatron.deepseek
- README.md
- resizable-sidebar.js
- runllm-widget.js
- custom.css
- logo.png
- agent_loop.rst
- attention_implementation.rst
- checkpoint.rst
- dpo_extension.rst
- fsdp_extension.rst
- fully_async.md
- megatron_extension.rst
- one_step_off.md
- placement.rst
- ppo_lora.rst
- reward_loop.rst
- rollout_corr.md
- rollout_corr_math.md
- rollout_skip.rst
- rollout_trace.rst
- rope.rst
- baseline.md
- collabllm.md
- dapo.md
- entropy.md
- gpg.md
- grpo.md
- opo.md
- ppo.md
- spin.md
- sppo.md
- amd_build_dockerfile_page.rst
- amd_vllm_page.rst
- data.rst
- single_controller.rst
- trainer.rst
- utils.rst
- ascend_profiling_en.rst
- ascend_profiling_zh.rst
- ascend_quick_start.rst
- ascend_sglang_quick_start.rst
- dockerfile_build_guidance.rst
- transfer_queue.md
- config.rst
- gsm8k_example.rst
- multi_modal_example.rst
- ppo_code_architecture.rst
- sandbox_fusion_example.rst
- skypilot_examples.rst
- faq.rst
- best_practices.rst
- device_tuning.rst
- dpsk.md
- nsight_profiling.md
- perf_tuning.rst
- verl_profiler_system.md
- prepare_data.rst
- reward_function.rst
- interaction_system.rst
- multiturn.rst
- sandbox_fusion.rst
- search_tool_example.rst
- agentic_rl.rst
- install.rst
- more_resources.rst
- multinode.rst
- quickstart.rst
- ray_debug_tutorial.rst
- fsdp_workers.rst
- megatron_workers.rst
- model_engine.rst
- ray_trainer.rst
- sglang_worker.rst
- conf.py
- hybrid_flow.rst
- index.rst
- Makefile
- README.md
- README_vllm0.7.md
- README_vllm0.8.md
- requirements-docs.txt
- single_controller.rst
- aime2024_multiturn_w_tool.py
- dapo_multiturn_w_tool.py
- full_hh_rlhf.py
- geo3k.py
- geo3k_multiturn_w_tool.py
- gsm8k.py
- gsm8k_multiturn_sft.py
- gsm8k_multiturn_w_interaction.py
- gsm8k_multiturn_w_tool.py
- gsm8k_tool_agent_loop.py
- hellaswag.py
- math_dataset.py
- multiturn.py
- preprocess_search_r1_dataset.py
- run_deepseek7b_mutli_node.sh
- run_deepseek_v2_lite_math.sh
- README.md
- run_qwen2_5-7b_math.sh
- test_dapo_7b_math.sh
- test_dapo_qwen3_30b_math.sh
- gpg.md
- run_qwen2-7b_math.sh
- run_qwen2-7b_math_megatron.sh
- README.md
- run_deepseek671b_math_megatron_80gb.sh
- run_deepseek671b_math_megatron_96gb.sh
- run_deepseek7b_llm.sh
- run_deepseek7b_llm_math.sh
- run_deepseek7b_llm_math_megatron.sh
- run_deepseek7b_llm_seq_balance.sh
- run_glm41v_9b.sh
- run_gptoss_20b.sh
- run_minicpmo2_6.sh
- run_mistral13b_skyworkrm_hhrlhf.sh
- run_moonlight16b_math_megatron.sh
- run_qwen2-7b.sh
- run_qwen2-7b_math.sh
- run_qwen2-7b_math_megatron.sh
- run_qwen2-7b_seq_balance.sh
- run_qwen2-7b_seq_balance_math_megatron.sh
- run_qwen2-7b_sgl_megatron.sh
- run_qwen2_5-3b_gsm8k_grpo_lora.sh
- run_qwen2_5-3b_gsm8k_grpo_lora_from_adapter.sh
- run_qwen2_5-7b_math_megatron_diff_tp.sh
- run_qwen2_5_32b_grpo_npu.sh
- run_qwen2_5_7b_grpo_discrete_prof_npu.sh
- run_qwen2_5_7b_grpo_e2e_prof_npu.sh
- run_qwen2_5_7b_grpo_npu.sh
- run_qwen2_5_vl-7b-megatron.sh
- run_qwen2_5_vl-7b-sglang.sh
- run_qwen2_5_vl-7b.sh
- run_qwen2_5_vl-7b_freeze_vision.sh
- run_qwen2_5_vl-7b_lora.sh
- run_qwen2_5_vl-7b_seq_balance.sh
- run_qwen2_5_vl_32b_npu.sh
- run_qwen2_5_vl_3b_npu.sh
- run_qwen2_5_vl_7b_npu.sh
- run_qwen3-235b_megatron_96gb.sh
- run_qwen3-32b_npu.sh
- run_qwen3-8b.sh
- run_qwen3-8b_npu.sh
- run_qwen3_8b_grpo_sglang_1k_spmd_npu.sh
- run_qwen3_8b_grpo_sglang_32k_spmd_npu.sh
- run_qwen3_vl-235b-megatron.sh
- run_qwen3_vl-30b-megatron.sh
- run_qwen3_vl-8b-megatron.sh
- run_qwen3moe-30b_megatron_96gb.sh
- run_seed_oss_36b.sh
- on_policy_distillation.sh
- README.md
- run_deepseek7b_llm.sh
- run_deepseek7b_llm_modelscope.sh
- run_deepseek7b_llm_pfppo.sh
- run_deepseek7b_llm_sandbox_fusion.sh
- run_deepseek7b_llm_sp2.sh
- run_deepseek_full_hh_rlhf.sh
- run_deepseek_math_gsm8k_megatron.sh
- run_deepseek_math_gsm8k_megatron_nsys.sh
- run_gemma.sh
- run_moonlight16b_a3b_gsm8k_megatron.sh
- run_qwen1.5_moe_a2.7b-gsm8k_megatron.sh
- run_qwen2-7b_math_gsm8k_megatron.sh
- run_qwen2-7b_rm.sh
- run_qwen2-7b_rm_seq_balance.sh
- run_qwen2-7b_rm_seq_balance_fused_kernels.sh
- run_qwen2-7b_rm_seq_balance_nsys.sh
- run_qwen2-7b_seq_balance.sh
- run_qwen2-7b_sglang_seq_balance.sh
- run_qwen2.5-32b.sh
- run_qwen3-8b_npu.sh
- tutorial.ipynb
- run_qwen2-7b_math_rf.sh
- run_qwen2-7b_math_rf_baseline.sh
- run_qwen2.5-3b_seq_balance.sh
- run_qwen2.5-7b_seq_balance.sh
- run_qwen2-7b.sh
- README.md
- run_with_rollout_corr.sh
- run_deepseek_6b7.sh
- run_gemma_2b.sh
- run_gemma_7b.sh
- run_qwen3_1b7_sft.sh
- run_qwen3_8b_sft_peft_sp2_npu.sh
- run_qwen_05_peft.sh
- run_qwen_05_sp2.sh
- run_qwen_05_sp2_liger.sh
- run_seed_oss_36b_sft.sh
- run_qwen_05_sp2.sh
- gsm8k_interaction_config.yaml
- geo3k_tool_config.yaml
- gsm8k_tool_config.yaml
- mcp_server.json
- mcp_tool_config.yaml
- sandbox_fusion_tool_config.yaml
- search_tool_config.yaml
- geo3k_multiturn_grpo.yaml
- geo3k_multiturn_megatron_grpo.yaml
- gsm8k_multiturn_grpo.yaml
- gsm8k_multiturn_grpo_server.yaml
- gsm8k_multiturn_grpo_w_interaction.yaml
- gsm8k_multiturn_megatron_grpo.yaml
- retool_multiturn_grpo.yaml
- search_multiturn_grpo.yaml
- search_multiturn_grpo_one_step_off.yaml
- run_qwen2.5-3b_geo3k_multiturn.sh
- run_qwen2.5-3b_geo3k_multiturn_4xgpu.sh
- run_qwen2.5-3b_megatron_geo3k_multiturn.sh
- download.py
- retrieval_server.py
- run_qwen2.5-3b_instruct_search_multiturn.sh
- README.md
- run_qwen0.5b_gsm8k_multiturn_curriculum.sh
- run_qwen2.5-0.5b_gsm8k_multiturn_w_interaction.sh
- run_qwen2.5-3b_gsm8k_multiturn.sh
- run_qwen2.5-3b_gsm8k_multiturn_4xgpu.sh
- run_qwen2.5-3b_gsm8k_multiturn_4xgpu_server.sh
- run_qwen2.5-3b_gsm8k_multiturn_server.sh
- run_qwen2.5-3b_gsm8k_multiturn_vllm_fsdp.sh
- run_qwen2.5-3b_gsm8k_tool_agent_mlflow.sh
- run_qwen2.5-3b_megatron_gsm8k_multiturn.sh
- run_qwen3-4b_gsm8k_multiturn.sh
- run_qwen3_4b_dapo_multiturn.sh
- README.md
- verl-grpo.yaml
- verl-multiturn-tools.yaml
- verl-ppo.yaml
- ray_on_slurm.slurm
- ppo_trainer_split.yaml
- main_ppo_split.py
- README.md
- run_deepseek7b_llm.sh
- split_monkey_patch.py
- qwen2-0.5b_grpo-lora_1_h100_fsdp_vllm.sh
- qwen2-1.5b_grpo-lora_1_h100_fsdp_vllm.sh
- qwen2-14b_grpo-lora_2_h100_fsdp_vllm.sh
- qwen2_14b_grpo_4_h800_fsdp_vllm.sh
- qwen2-32b_grpo-lora_4_h100_fsdp_vllm.sh
- qwen2_32B_grpo_8_h20_megatron_vllm.sh
- qwen2-3b_grpo-lora_1_h100_fsdp_vllm.sh
- qwen2-70b_grpo_32_h20_fsdp_vllm.sh
- qwen2-70b_grpo_32_h800_fsdp_vllm.sh
- qwen2-72b_grpo-lora_8_h100_fsdp_vllm.sh
- qwen2-7b_grpo-lora_1_h100_fsdp_vllm.sh
- qwen2-7b_grpo_2_h800_fsdp_vllm.sh
- agent_loop_tutorial.ipynb
- sandbox.py
- create_dataset.py
- README.md
- reward_function.py
- train_grpo.sh
- train_sft.sh
- agent.yaml
- collabllm_interaction_config.yaml
- accuracy.py
- bleu_score.py
- interactivity.py
- pass_rate.py
- token_amount.py
- collabllm_agent_loop.py
- collabllm_interation.py
- process_dataset.py
- README.md
- reward_function.py
- train_rl_collabllm.sh
- train_sft_collabllm.sh
- utils.py
- dapo_megatron_trainer.yaml
- dapo_trainer.yaml
- dapo_ray_trainer.py
- main_dapo.py
- prepare_dapo_data.sh
- README.md
- run_dapo_early_qwen2.5_32b.sh
- run_dapo_qwen2.5_32b.sh
- run_dapo_qwen2.5_32b_npu.sh
- run_dapo_qwen2.5_32b_rollout_corr.sh
- run_dapo_qwen2.5_7b_npu.sh
- run_dapo_qwen3_14b_base_npu.sh
- run_dapo_qwen3_8b_base_npu.sh
- run_dapo_qwen3_moe_30b_base_fsdp_npu.sh
- run_dapo_qwen3_moe_30b_megatron_npu.sh
- run_dapo_wo_ds_qwen2.5_32b.sh
- runtime_env.yaml
- test_dapo_7b.sh
- test_dapo_7b_math.sh
- test_dapo_7b_math_lora.sh
- test_dapo_7b_math_megatron.sh
- test_dapo_dspk_671b_megatron_96gb.sh
- test_dapo_glm_air_megatron.sh
- test_dapo_qwen3_30b_math.sh
- test_dapo_qwen3_30b_math_single_node.sh
- deepeyes_multiturn_grpo.yaml
- image_zoom_in_tool_config.yaml
- deepeyes.py
- README.md
- run_deepeyes_grpo.sh
- entropy_trainer.yaml
- __init__.py
- grader.py
- math_normalize.py
- __init__.py
- 32b_clip_cov.sh
- 32b_kl_cov.sh
- 32b_kl_cov_mininbsz.sh
- 7b_clip_cov.sh
- 7b_kl_cov.sh
- entropy_ray_trainer.py
- main_entropy.py
- README.md
- reward.py
- rm_config.yaml
- prepare_fapo_data.py
- README.md
- reward_fn_genrm.py
- reward_fn_reasoning.py
- reward_fn_reasoning_remote.py
- run_baseline_32b.sh
- run_baseline_7b.sh
- run_fapo_32b.sh
- run_fapo_32b_remote.sh
- run_fapo_7b.sh
- run_fapo_7b_remote.sh
- run_fapo_genrm_train.sh
- runtime_env.yaml
- flowrl_trainer.yaml
- file.svg
- flowrl.pdf
- flowrl.png
- prepare_data.sh
- prepare_model.sh
- __init__.py
- flowrl_actor.py
- flowrl_fsdp_worker.py
- flowrl_ray_trainer.py
- FLOWRL_SIMPLE_GUIDE.md
- main_flowrl.py
- README.md
- run_flowrl_qwen2.5_7b.sh
- __init__.py
- agent_loop.py
- partial_single_turn_agent_loop.py
- fully_async_ppo_megatron_trainer.yaml
- fully_async_ppo_trainer.yaml
- dapo_7b_math_fsdp2_16_16.sh
- dapo_7b_math_fsdp2_32_32.sh
- dapo_7b_math_fsdp2_4_12.sh
- dapo_7b_math_fsdp2_4_4.sh
- dapo_7b_math_fsdp2_64_64.sh
- dapo_7b_math_fsdp2_64_64_mis.sh
- dapo_7b_math_fsdp2_8_8.sh
- geo3k_qwen25vl_7b_megatron_4_4.sh
- runtime_env.yaml
- simple_streaming_demo.py
- __init__.py
- vllm_async_server.py
- detach_utils.py
- fsdp2_utils.py
- fsdp_workers.py
- fully_async_main.py
- fully_async_rollouter.py
- fully_async_trainer.py
- megatron_worker.py
- message_queue.py
- param_sync.py
- ray_trainer.py
- README.md
- README_zh.md
- README.md
- reward_function.py
- run_genrm_remote.sh
- on_policy_distill_trainer.yaml
- runtime_env.yaml
- __init__.py
- client.py
- join_server.sh
- proxy.py
- start_server.sh
- utils.py
- vllm_engine.py
- worker.py
- main_gkd.py
- megatron_kl_loss.py
- megatron_utils.py
- megatron_workers.py
- ray_trainer.py
- README.md
- run_moonlight_dsv3_training.sh
- teacher_utils.py
- test_qwen.sh
- test_qwen_sglang.sh
- test_teacher_server.py
- test_gspo_3b_math.sh
- test_gspo_3b_math_slurm.sh
- test_gspo_qwen30b_a3b_ep.sh
- README.md
- reward_fn.py
- run_3b.sh
- run_7b.sh
- agent.yaml
- create_dataset.py
- math_expression.py
- README.md
- run_gpt_oss_20b_bf16.sh
- run_qwen2.5_3b.sh
- __init__.py
- chat_model.py
- react_agent_loop.py
- test_react_agent_loop.py
- rl_dataset.py
- one_step_off_ppo_megatron_trainer.yaml
- one_step_off_ppo_trainer.yaml
- dapo_7b_math_fsdp2_4_12.sh
- dapo_7b_math_fsdp2_colocate.sh
- dapo_7b_math_fsdp2_sglang_4_12.sh
- dapo_7b_math_fsdp2_sglang_colocate.sh
- dapo_7b_math_megatron_4_12.sh
- dapo_7b_math_megatron_colocate.sh
- distributed_util.py
- fsdp_workers.py
- grpo_0.6b_gsm8k_fsdp2_2_6.sh
- grpo_0.6b_gsm8k_fsdp2_sglang_2_6.sh
- grpo_3b_gsm8k_fsdp2_2_6.sh
- main_ppo.py
- megatron_workers.py
- ray_trainer.py
- README.md
- sglang_sharding_manager.py
- utils.py
- vllm_sharding_manager.py
- compute_score.py
- prepare_eval_dataset.py
- prepare_nvidia-OpenMathReasoning_sft.py
- README.md
- run_eval.sh
- run_generation.sh
- run_sft_qwen3_8b.sh
- prime_trainer.yaml
- __init__.py
- main_prime.py
- prime_core_algos.py
- prime_dp_rm.py
- prime_fsdp_workers.py
- prime_ray_trainer.py
- run_prime_qwen.sh
- run_prime_qwen_code.sh
- evaluation.yaml
- __init__.py
- gpqa.py
- livecodebench.py
- math_reward.py
- __init__.py
- data_process.py
- main_eval.py
- README.md
- reward_score.py
- run_r1_distill_qwen.sh
- README.md
- retool.py
- retool_sft_preprocess.py
- run_gpt_oss_ppo.sh
- run_qwen2-32b_dapo.sh
- run_qwen2-32b_ppo.sh
- run_qwen2-32b_sft.sh
- run_qwen2_7b_dapo.sh
- run_qwen2_7b_sft.sh
- run_qwen2_7b_sft_npu.sh
- sandbox_fusion_tool_config.yaml
- spin_trainer.yaml
- core_algos.py
- dp_actor.py
- fsdp_workers.py
- main_spin.py
- README.md
- run_spin.sh
- spin_trainer.py
- utils.py
- sppo_trainer.yaml
- __init__.py
- config.py
- dp_actor.py
- main_sppo.py
- README.md
- run_qwen2.5-7b_rm.sh
- sppo_ray_trainer.py
- sppo_worker.py
- transfer_queue_ppo_trainer.yaml
- agent_loop.py
- main_ppo.py
- ray_trainer.py
- run_qwen3-8b_transferqueue_npu.sh
- README.md
- __init__.py
- converter_hf_to_mcore.py
- diagnose.py
- generate_trainer_config.sh
- init_random_model.py
- install_vllm_sglang_mcore.sh
- legacy_model_merger.py
- print_cfg.py
- rollout_viewer.py
- agent_utils.py
- qwen_vl_tool_chat_template.jinja2
- test_agent_loop_reward.py
- test_agent_loop_reward_model.py
- test_basic_agent_loop.py
- test_gpt_oss_tool_parser.py
- test_multi_modal.py
- test_standalone_rollout.py
- reward_fn.py
- test_agent_loop_reward_manager.py
- test_reward_model.py
- __init__.py
- test_gsm8k_interaction.py
- test_interaction_registry.py
- test_engine.py
- test_transformer.py
- test_transformers_ulysses.py
- test_decorator.py
- main.py
- client.py
- README.md
- run.sh
- server.py
- __init__.py
- test_auto_padding_on_cpu.py
- test_colocated_workers.py
- test_colocated_workers_fused.py
- test_data_transfer.py
- test_decorator_on_cpu.py
- test_device_mesh_register.py
- test_driverfunc_to_worker.py
- test_fused_workers_on_cpu.py
- test_high_level_scheduling_api.py
- test_nested_worker.py
- test_ray_collectives.py
- test_ray_local_envs_on_cpu.py
- test_ray_utils_on_cpu.py
- test_rvdz.py
- test_worker_group_basics.py
- test_worker_group_torch.py
- README.md
- run_all.sh
- test_fsdp_ckpt.py
- test_mcore_config_converter.py
- test_tensor_dict.py
- __init__.py
- task.py
- tokenizer.py
- __init__.py
- run_gen_qwen05.sh
- run_gen_qwen05_server.sh
- qwen2moe_minimal.json
- run_function_reward.sh
- run_model_reward.sh
- run_single_gpu.sh
- run_single_gpu_with_engine.sh
- compare_sft_engine_results.py
- run_sft.sh
- run_sft_engine_gsm8k.sh
- test_sft_engine_all.sh
- test_sp_loss_match.py
- __init__.py
- check_custom_rwd_fn.py
- check_results.py
- README.md
- run_dapo.sh
- run_fully_async_policy.sh
- run_genrm_remote.sh
- run_geo3k_fsdp_sgl_multiturn_w_tool.sh
- run_grpo_lora_with_merge.sh
- run_gsm8k_fsdp_sgl_multiturn_sf_tool.sh
- run_gsm8k_fsdp_sgl_multiturn_w_tool.sh
- run_one_step_off_policy.sh
- run_ppo_trainer_megatron.sh
- run_prime.sh
- run_r1_distill_qwen_aime24_eval.sh
- run_spin.sh
- run_sppo.sh
- run_test.sh
- run_qwen2_5_05b_dapo.sh
- run_qwen2_5_05b_grpo.sh
- run_qwen2_5_05b_grpo_mindspeed.sh
- run_qwen2_5_05b_sft_peft_sp2.sh
- run_qwen2_5_vl_3b_npu.sh
- run_qwen3_06b_ppo.sh
- check_api_docs.py
- check_dataproto_usage.py
- check_device_api_usage.py
- check_docs_time_info.py
- check_docstrings.py
- check_license.py
- check_pr_description.py
- check_pr_title.py
- test_config_docs.py
- test_import.py
- type_coverage_check.py
- validate_imported_docs.py
- validate_structure.py
- README.md
- test_memory_buffers.py
- __init__.py
- legacy_ppo_megatron_trainer.yaml
- legacy_ppo_trainer.yaml
- test_algo_config_on_cpu.py
- test_legacy_config_on_cpu.py
- __init__.py
- test_core_algos_on_cpu.py
- test_metric_utils_on_cpu.py
- test_rollout_corr.py
- test_rollout_corr_integration.py
- __init__.py
- test_esi_save_ckpt_on_cpu.py
- test_create_rl_sampler_on_cpu.py
- test_multiturn_sft_dataset_on_cpu.py
- test_rl_collate_fn_on_cpu.py
- test_rl_dataset_on_cpu.py
- test_sft_dataset_on_cpu.py
- test_metrics.py
- test_pipeline_parallel.py
- test_sandbox_fusion_on_cpu.py
- test_sandbox_on_cpu.py
- _test_module.py
- test_activation_offload.py
- test_config_on_cpu.py
- test_flops_counter.py
- test_fs_on_cpu.py
- test_groupwise.py
- test_import_utils_on_cpu.py
- test_linear_cross_entropy.py
- test_mlflow_key_sanitization.py
- test_model_on_cpu.py
- test_nvtx_profile.py
- test_rollout_skip_on_cpu.py
- test_rollout_trace_on_cpu.py
- test_seqlen_balancing.py
- test_special_linear_cross_entropy_tp.py
- test_special_mstx_profile.py
- test_temp_env_on_cpu.py
- test_timeout_decorator_cpu.py
- test_torch_functional.py
- test_special_dp_actor.py
- test_actor_config_on_cpu.py
- test_critic_config_on_cpu.py
- test_engine_config_on_cpu.py
- test_optim_config_on_cpu.py
- test_special_dp_critic.py
- test_registry_on_cpu.py
- vllm_async_rollout.py
- mcp_server.json
- mcp_tool_config
- sandbox_fusion_tool_config
- search_tool_config
- test_http_server_engine.py
- run_fsdp_vllm.py
- test_vllm_model_rope_scaling.py
- test_vllm_spmd.py
- test_hf_rollout.py
- test_sglang_async_rollout_mcp_tools.py
- test_sglang_async_rollout_multimodal_delta.py
- test_sglang_async_rollout_search_tools.py
- test_sglang_async_rollout_sf_tools.py
- test_sglang_async_rollout_w_interaction.py
- test_sglang_async_rollout_w_tools.py
- test_sglang_async_rollout_w_tools_token_out.py
- test_sglang_multi_interaction.py
- test_sglang_rollout_sharding_manager.py
- test_sglang_spmd.py
- utils_sglang.py
- test_fsdp_attn_implementation.py
- test_fsdp_workers.py
- __init__.py
- kill_github_tests.sh
- README.md
- test_base_config_on_cpu.py
- test_protocol_on_cpu.py
- test_protocol_v2_on_cpu.py
- __init__.py
- agent_loop.py
- single_turn_agent_loop.py
- tool_agent_loop.py
- tool_parser.py
- utils.py
- __init__.py
- sampler.py
- __init__.py
- dynamicgen_dataset.py
- __init__.py
- base.py
- dapo.py
- naive.py
- registry.py
- naive_router.py
- sglang_router.py
- __init__.py
- reward_manager.py
- reward_model.py
- __init__.py
- __init__.py
- interaction_registry.py
- __init__.py
- base.py
- gsm8k_interaction.py
- weather_interaction.py
- __init__.py
- __main__.py
- base_model_merger.py
- fsdp_model_merger.py
- megatron_model_merger.py
- __init__.py
- llama_loader.py
- llama_loader_depracated.py
- llama_saver.py
- __init__.py
- parallel_attention.py
- parallel_decoder.py
- parallel_linear.py
- parallel_mlp.py
- parallel_rmsnorm.py
- __init__.py
- modeling_llama_megatron.py
- __init__.py
- __init__.py
- attention.py
- model.py
- rope_utils.py
- vision_config.py
- vision_model.py
- vision_transformer_block.py
- __init__.py
- config_converter.py
- loader.py
- mbridge.py
- model_forward.py
- model_forward_1f1b_overlap.py
- model_forward_fused.py
- model_initializer.py
- patch_v012.py
- readme.md
- registry.py
- saver.py
- util.py
- weight_converter.py
- __init__.py
- qwen2_loader.py
- qwen2_loader_depracated.py
- qwen2_saver.py
- __init__.py
- parallel_attention.py
- parallel_decoder.py
- parallel_linear.py
- parallel_mlp.py
- parallel_rmsnorm.py
- __init__.py
- modeling_qwen2_megatron.py
- __init__.py
- __init__.py
- apertus.py
- dense_common.py
- glm4v.py
- kimi_vl.py
- llama.py
- monkey_patch.py
- npu_patch.py
- qwen2.py
- qwen2_vl.py
- qwen3_vl.py
- __init__.py
- README.md
- registry.py
- weight_loader_registry.py
- __init__.py
- decorator.py
- worker.py
- worker_group.py
- __init__.py
- base.py
- __init__.py
- __init__.py
- parallel_state.py
- __init__.py
- state_dict.py
- __init__.py
- _state_dict_utils.py
- __init__.py
- __init__.py
- __init__.py
- McpClientManager.py
- utils.py
- __init__.py
- search_r1_like_utils.py
- tool_registry.py
- __init__.py
- base_tool.py
- geo3k_tool.py
- gsm8k_tool.py
- image_zoom_in_tool.py
- mcp_base_tool.py
- mcp_search_tool.py
- sandbox_fusion_tools.py
- schemas.py
- search_tool.py
- actor.yaml
- dp_actor.yaml
- megatron_actor.yaml
- rollout_correction.yaml
- critic.yaml
- dp_critic.yaml
- megatron_critic.yaml
- legacy_data.yaml
- fsdp.yaml
- megatron.yaml
- hf_model.yaml
- npu_profile.yaml
- fsdp.yaml
- megatron.yaml
- dp_ref.yaml
- megatron_ref.yaml
- ref.yaml
- dp_reward_model.yaml
- megatron_reward_model.yaml
- reward_model.yaml
- rollout.yaml
- __init__.py
- _generated_ppo_megatron_trainer.yaml
- _generated_ppo_trainer.yaml
- algorithm.py
- config.py
- evaluation.yaml
- generation.yaml
- ppo_megatron_trainer.yaml
- ppo_trainer.yaml
- sft_trainer.yaml
- sft_trainer_engine.yaml
- __init__.py
- core_algos.py
- metric_utils.py
- ray_trainer.py
- reward.py
- rollout_corr_helper.py
- utils.py
- __init__.py
- constants_ppo.py
- fsdp_sft_trainer.py
- main_eval.py
- main_generation.py
- main_generation_server.py
- main_ppo.py
- runtime_env.yaml
- sft_trainer.py
- __init__.py
- checkpoint_handler.py
- checkpoint_manager.py
- fsdp_checkpoint_manager.py
- megatron_checkpoint_manager.py
- __init__.py
- dataset_utils.py
- multiturn_sft_dataset.py
- README.md
- rl_dataset.py
- rm_dataset.py
- sft_dataset.py
- vision_utils.py
- __init__.py
- metrics.py
- performance.py
- trajectory_tracker.py
- __init__.py
- torch_functional.py
- __init__.py
- kernels.py
- linear_cross_entropy.py
- __init__.py
- aggregate_logger.py
- __init__.py
- dist_checkpointing.py
- memory.py
- optimizer.py
- pipeline_parallel.py
- sequence_parallel.py
- tensor_parallel.py
- __init__.py
- utils.py
- __init__.py
- config.py
- empty_annotations.py
- mstx_profile.py
- nvtx_profile.py
- performance.py
- profile.py
- __init__.py
- ray_backend.py
- __init__.py
- README.md
- testing_util.py
- utils.py
- __init__.py
- grader.py
- math_normalize.py
- __init__.py
- utils.py
- __init__.py
- grader.py
- math_normalize.py
- math_utils.py
- __init__.py
- geo3k.py
- gsm8k.py
- math_batch.py
- math_dapo.py
- math_reward.py
- math_verify.py
- search_r1_like_qa_em.py
- __init__.py
- patch.py
- utils.py
- __init__.py
- activation_offload.py
- attention_utils.py
- config.py
- device.py
- distributed.py
- flops_counter.py
- fs.py
- fsdp_utils.py
- groupwise.py
- hdfs_io.py
- import_utils.py
- logging_utils.py
- megatron_utils.py
- memory_buffer.py
- memory_utils.py
- model.py
- net_utils.py
- npu_utils.py
- py_functional.py
- ray_utils.py
- rollout_skip.py
- rollout_trace.py
- seqlen_balancing.py
- tensordict_utils.py
- tokenizer.py
- torch_dtypes.py
- torch_functional.py
- tracking.py
- transferqueue_utils.py
- transformers_compat.py
- ulysses.py
- version
- __init__.py
- base.py
- dp_actor.py
- megatron_actor.py
- __init__.py
- actor.py
- critic.py
- engine.py
- model.py
- optimizer.py
- reward_model.py
- rollout.py
- __init__.py
- base.py
- dp_critic.py
- megatron_critic.py
- __init__.py
- transformer_impl.py
- utils.py
- __init__.py
- transformer_impl.py
- utils.py
- __init__.py
- transformer_impl.py
- __init__.py
- base.py
- utils.py
- __init__.py
- abstract.py
- batch.py
- dapo.py
- naive.py
- prime.py
- registry.py
- __init__.py
- reward_model.py
- __init__.py
- base.py
- __init__.py
- losses.py
- padding.py
- __init__.py
- actor.py
- critic.py
- hybrid_engine.py
- __init__.py
- naive_rollout.py
- __init__.py
- async_sglang_server.py
- http_server_engine.py
- sglang_rollout.py
- utils.py
- __init__.py
- utils.py
- vllm_async_server.py
- vllm_rollout_spmd.py
- __init__.py
- base.py
- hf_rollout.py
- replica.py
- schemas.py
- tokenizer.py
- utils.py
- __init__.py
- base.py
- fsdp_sglang.py
- fsdp_ulysses.py
- fsdp_vllm.py
- megatron_sglang.py
- megatron_vllm.py
- __init__.py
- fsdp_workers.py
- megatron_workers.py
- __init__.py
- AGENTS.md
- base_config.py
- protocol.py
- py.typed
- .gitignore
- .pre-commit-config.yaml
- .readthedocs.yaml
- CONTRIBUTING.md
- LICENSE
- Notice.txt
- pyproject.toml
- README.md
- requirements-cuda.txt
- requirements-npu.txt
- requirements.txt
- requirements_sglang.txt
- requirements_transferqueue.txt
- setup.py
- opd.sh
- README.md
- grpo.sh
- on_policy_distillation.sh
- README.md
# Installation Guide
git clone https://github.com/thunlp/OPD
Downloads the entire project code from GitHub to your computer.
cd OPD
Moves into the project folder you just downloaded.
2. Official Install Script
Easy Recommended- Python 3 Python is required to use pip.
pip install math-verify
Installs the package published on PyPI directly β no need to clone the source.
Pulled directly from this repo's README.
3. Docker
Easy- Git Needed to download the project code from GitHub.
- Docker Desktop Needed to build and run containers. Install it and keep it running in the background.
docker compose -f LlamaFactory/docker/docker-cuda/docker-compose.yml up -d --build
Runs the command against the services defined in the compose file.
4. Python
Easypip install math-verify
Installs the package published on PyPI directly β no need to clone the source.
pip install -e .
Installs the Python libraries listed in requirements.txt (or similar).
pip install -r requirements/metrics.txt
Installs the Python libraries listed in requirements.txt (or similar).
Pulled directly from this repo's README.
5. Make
Medium- Git Needed to download the project code from GitHub.
- Make Usually pre-installed on Linux/macOS. On Windows, install separately (e.g. via MSYS2 or WSL).
cd LlamaFactory
This project's files live in a subfolder, so move into it first.
make
Compiles the code based on the generated build configuration to produce an executable.
