Relax
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
파일 탐색기
최종 버전 다운로드 (.zip)- agents
- commands
- skills
- agents
- commands
- skills
- agents
- commands
- skills
- bug_report.md
- config.yml
- feature_request.md
- ci.yml
- deploy-docs.yml
- CODEOWNERS
- PULL_REQUEST_TEMPLATE.md
- SECURITY.md
- CODEOWNERS
- algorithm-expert.md
- fsdp-expert.md
- launcher-expert.md
- megatron-expert.md
- ray-expert.md
- commit.md
- skills
- check_conflict_markers.py
- copyright.py
- format_ascii_boxes.py
- gitleaks_tracked.py
- qwen35-35B-A3B-dapo-precision.png
- qwen35-35B-A3B-sft-precision.png
- qwen35-35B-A3B-vl-mmrl-precision.png
- qwen35-9B-dapo-precision.png
- qwen35-9B-vl-mmrl-precision.png
- arch.png
- Relax.jpg
- env.yaml
- megatron.patch
- megatron-bridge.patch
- megatron.patch
- mindspeed-bridge.patch
- mindspeed.patch
- sglang-npu.patch
- torch-memory-saver.patch
- megatron.patch
- sglang.patch
- sglang_per_pos_topk.patch
- 20251218-3714d81d.patch
- 20260506-85bced0ae.patch
- v0.5.12.post1.patch
- v0.5.15.post1.patch
- v0.5.9.patch
- Dockerfile
- Dockerfile.npu
- npu-training.md
- README.md
- xpu-instructions.md
- AsciiBackground.vue
- AsciiFireworks.vue
- AsciiHero.vue
- CallToAction.vue
- custom.css
- FeatureGrid.vue
- FeatureShowcase.vue
- index.ts
- MarqueeStrip.vue
- SwaggerUI.vue
- config.mts
- tsconfig.json
- .keep
- apprise_notification.md
- autoscaler_k8s_keda.md
- dataset_design.md
- design.md
- distributed_checkpoint_service.md
- dynamic-context-parallel.md
- elastic_rollout.md
- external-model-integration.md
- fully_async.md
- generative_reward_model.md
- health_check_manager.md
- how_to_contribute.md
- metrics_service_usage.md
- on-policy-distillation.md
- restructure-proposal.md
- actor-fwd.md
- actor.md
- genrm.md
- overview.md
- rollout.md
- algorithms.md
- deepeyes.md
- generative-reward-model.md
- low-precision-training.md
- on-policy-distillation.md
- agentic-rollout.md
- architecture.md
- autoscaler-k8s-keda.md
- configuration.md
- customize-training.md
- dataset-design.md
- debugging.md
- distributed-checkpoint.md
- dynamic-context-parallel.md
- elastic-rollout.md
- external-model-integration.md
- fully-async-training.md
- health-check-manager.md
- how-to-contribute.md
- hybrid-training.md
- installation.md
- introduction.md
- low-rank-adaptation-training.md
- metrics-service-detailed.md
- model-conversion.md
- mtp-rl-training.md
- notification-system.md
- oom-troubleshooting.md
- performance-tuning.md
- ppo-training.md
- quick-start.md
- reinforce-plus-plus-training-report.md
- reinforce-plus-plus.md
- rollout-result-viewer.md
- s3-model-loading.md
- sft-training.md
- update-weights-pipeline.md
- index.md
- agent_app.svg
- architecture.svg
- inflight_request.svg
- lifecycle.svg
- partial_rollout.svg
- partial_rollout_sequence.svg
- warmup.svg
- warmup_flow.svg
- actor.json
- actor_fwd.json
- genrm.json
- rollout.json
- evaluation_paired_reward_differences.csv
- evaluation_reward.svg
- evaluation_summary_by_algorithm.csv
- evaluation_summary_by_run.csv
- evidence_index.csv
- training_advantage_std_curve.svg
- training_kl_loss_curve.svg
- training_loss_curve.svg
- training_metrics_long.csv
- training_response_length_curve.svg
- training_reward_curve.svg
- training_summary_by_algorithm.csv
- training_summary_by_run.csv
- training_throughput_curve.svg
- training_truncation_curve.svg
- autoscaler_monitor.png
- deepeyes-h800.png
- favicon-16.png
- favicon-32.png
- favicon.png
- logo.jpg
- rednote-logo.png
- relax-viewer.png
- timeline_demo.png
- actor-fwd.md
- actor.md
- genrm.md
- overview.md
- rollout.md
- algorithms.md
- deepeyes.md
- generative-reward-model.md
- low-precision-training.md
- on-policy-distillation.md
- agentic-rollout.md
- architecture.md
- autoscaler-k8s-keda.md
- configuration.md
- customize-training.md
- dataset-design.md
- debugging.md
- distributed-checkpoint.md
- dynamic-context-parallel.md
- elastic-rollout.md
- external-model-integration.md
- fully-async-training.md
- health-check-manager.md
- how-to-contribute.md
- hybrid-training.md
- installation.md
- introduction.md
- low-rank-adaptation-training.md
- metrics-service-detailed.md
- model-conversion.md
- mtp-rl-training.md
- notification-system.md
- oom-troubleshooting.md
- performance-tuning.md
- ppo-training.md
- quick-start.md
- reinforce-plus-plus-training-report.md
- reinforce-plus-plus.md
- rollout-result-viewer.md
- s3-model-loading.md
- sft-training.md
- update-weights-pipeline.md
- index.md
- .gitignore
- deploy-docs.sh
- DEPLOYMENT_SUMMARY.md
- Dockerfile
- fix-chunk-names.js
- GET_STARTED.md
- index.md
- QUICK_REFERENCE.md
- README.md
- start-docs.sh
- README.md
- run-qwen3-0.6B-1xgpu-reinforce-plus-plus.sh
- run-qwen3-0.6B-1xgpu-rloo.sh
- run-qwen35-9B-8xgpu-openr1mm-cispo-async.sh
- run_relax_deepseek_r1_distill_qwen_7b_sft_8xGPU.sh
- run_relax_qwen3_4b_thinking_sft_8xGPU.sh
- __init__.py
- processor.py
- __init__.py
- base_env.py
- deepeyes_config.yaml
- env_deepeyes.py
- processor_patch_utils.py
- README.md
- reward_deepeyes.py
- reward_deepeyes_genrm.py
- rollout.py
- run_deepeyes.sh
- run_deepeyes_fp16.sh
- run_deepeyes_genrm.sh
- run_deepeyes_pr.sh
- run_deepeyes_qwen35_9B_async.sh
- run_deepeyes_r3.sh
- sglang_judge_service.sh
- __init__.py
- agent.py
- deepeyes_config.yaml
- env_deepeyes.py
- __init__.py
- README.md
- reward_deepeyes.py
- run_agent_app.sh
- run_deepeyes_agentic.sh
- run_deepeyes_agentic_pr.sh
- run_deepeyes_agentic_qwen35_9B_async.sh
- run_deepeyes_agentic_qwen35_9B_async_sticky.sh
- sglang_judge_service.sh
- __init__.py
- _retry.py
- apptainer_jupyter_backend.py
- __init__.py
- base.py
- exceptions.py
- executor.py
- retry.py
- __init__.py
- agent.py
- deepeyes_v2_config.yaml
- env_deepeyes_v2.py
- prompt.py
- search_utils.py
- .gitignore
- apptainer_config.yaml
- deepeyes_v2_kernel.def
- verify_kernel.py
- __init__.py
- cache_convert.py
- rl_data_convert.py
- build_smoke_parquet.py
- prepare.sh
- run_single_session.py
- smoke.sh
- .gitignore
- __init__.py
- env.sh.example
- PITFALLS.md
- README.md
- reward_deepeyes_v2.py
- run_agent_app.sh
- run_deepeyes_v2_agentic.sh
- sglang_judge_service.sh
- __init__.py
- architecture-diagram-prompts.md
- post_process_genrm_swap.py
- README.md
- run-qwen3-4B-8xgpu-async.sh
- run-qwen3-4B-8xgpu-colocated.sh
- run-qwen35-35B-A3B-16xgpu-genrm-397B-defer.sh
- run-qwen35-35B-A3B-16xgpu-genrm-397B-split.sh
- __init__.py
- agent_client.sh
- agent_config.yaml
- agent_server.py
- check_r2e_sifs.py
- README.md
- run_mini_swe_agent.sh
- setup_r2e_data_and_sifs.sh
- __init__.py
- client.py
- protocol.py
- result.py
- calendar-training-entropy-loss.png
- calendar-training-mfu.png
- calendar-training-raw-reward.png
- PITFAIL.md
- prepare_calendar.sh
- README.md
- run-qwen3-4B-8xgpu-nemo-gym-calendar.sh
- start_calendar_gym.sh
- verify_calendar.py
- PITFAIL.md
- prepare_gsm8k.sh
- README.md
- run-qwen3-4B-8xgpu-nemo-gym.sh
- start_gsm8k_gym.sh
- verify_gsm8k.py
- r2e-training-entropy-grad-norm.png
- r2e-training-mfu-rollout-time.png
- r2e-training-raw-reward.png
- r2e-training-rollout-turns.png
- PITFAIL.md
- prepare_r2e_gym.py
- prepare_r2e_gym.sh
- README.md
- run-qwen35-9B-8xgpu-nemo-gym-r2e.sh
- start_r2e_gym_local.sh
- start_r2e_gym_remote.sh
- submit_r2e_gym.sh
- verify_r2e_gym_trial.py
- PITFAIL.md
- prepare_workplace_assistant.sh
- README.md
- run-qwen3-4B-2xgpu-nemo-gym-workplace.sh
- run-qwen3-4B-8xgpu-nemo-gym-workplace.sh
- start_workplace_assistant_gym.sh
- start_workplace_assistant_gym_remote.sh
- verify_workplace_assistant.py
- verify_workplace_assistant_trial.py
- convert_dataset.py
- run-qwen3-4B-8xgpu-nemo-gym.sh
- run_agent_app.sh
- run_gateway.sh
- run_training.sh
- relax_gateway_model.yaml
- app.py
- pyproject.toml
- README.md
- nemo_gym_http_error_traceback.patch
- openhands_platform_release.patch
- openhands_r2e_runtime.patch
- openhands_r2e_runtime_setup.patch
- openhands_rollout_prefix.patch
- openhands_setup_portability.patch
- swe_agents_r2e.patch
- workplace_assistant_cleanup.patch
- sitecustomize.py
- __init__.py
- app.py
- callback_provider.py
- config.py
- Dockerfile
- README.md
- registry.py
- run_adapter.py
- verbose_logging.py
- __init__.py
- conftest.py
- test_client.py
- test_convert_dataset.py
- test_gateway_app.py
- test_gateway_config.py
- test_gateway_service.py
- test_prepare_r2e_gym.py
- test_protocol.py
- test_recipe_paths.py
- test_result.py
- test_run_adapter.py
- test_run_agent_app.py
- test_verbose_logging.py
- __init__.py
- README.md
- RUNBOOK.md
- RUNBOOK_zh.md
- __init__.py
- agent.py
- config_tw.yaml
- env_alfworld.py
- prepare_data.py
- README.md
- reward_alfworld.py
- run-alfworld-grpo-qwen35-35B-A3B-8xgpu.sh
- run-alfworld-opd-qwen35-35B-A3B-8xgpu.sh
- run_agent_app.sh
- __init__.py
- agent.py
- config.yaml
- env_client.py
- prompt.py
- server.py
- prepare_data.py
- README.md
- reward_webshop.py
- run-webshop-grpo-qwen35-35B-A3B-8xgpu.sh
- run-webshop-opd-qwen35-35B-A3B-8xgpu.sh
- run_agent_app.sh
- run_webshop_server.sh
- run-opd-qwen35-35B-A3B-8xgpu-colocate.sh
- prepare_data.py
- README.md
- run-mopd-qwen3-vl-2b-8xgpu-colocate.sh
- run-mopd-qwen35-35ba3b-16xgpu-colocate.sh
- run-mopd-qwen35-9b-8xgpu-colocate.sh
- prepare_data.py
- run-opd-qwen3.5_35ba3b-122ba10b-128xgpu-colocate.sh
- run-vision-opd-qwen3.5_9b-35ba3b-8xgpu-2teacher-colocate.sh
- run-vision-opd-qwen3.5_9b-35ba3b-8xgpu-colocate.sh
- README.md
- __init__.py
- __init__.py
- prepare.py
- reward.py
- runtime.py
- transfer.py
- __init__.py
- ipc.py
- __init__.py
- service.py
- state.py
- __init__.py
- profile.py
- rollout.py
- fake_int4_quant_cuda.cu
- setup.py
- __init__.py
- fp8_kernel.py
- __init__.py
- chunked_grad_coalesce_patch.py
- __init__.py
- padding_remover.py
- quantizer_compressed_tensors.py
- quantizer_fp8.py
- __init__.py
- deepseekv3.py
- glm4.py
- glm4moe.py
- llama.py
- mimo.py
- qwen2.py
- qwen3_5.py
- qwen3_next.py
- qwen3_omni_moe.py
- qwen3_vl.py
- qwen3_vl_moe.py
- qwen3moe.py
- __init__.py
- bridge_converter.py
- common.py
- hf_weight_iterator_base.py
- hf_weight_iterator_bridge.py
- hf_weight_iterator_direct.py
- lora_adapter_sync.py
- train_offload.py
- update_weight_from_distributed.py
- update_weight_from_tensor.py
- __init__.py
- actor.py
- arguments.py
- checkpoint.py
- ci_utils.py
- collective_utils.py
- conditional_branch_sync.py
- cp_utils.py
- data.py
- initialize.py
- loss.py
- misc_utils.py
- model.py
- model_provider.py
- sglang.py
- streaming_schedules.py
- __init__.py
- arguments.py
- routing_replay_patch.py
- sglang_engine.py
- __init__.py
- __init__.py
- actor.py
- actor_fwd.py
- advantages.py
- base.py
- critic.py
- genrm.py
- rollout.py
- sft.py
- __init__.py
- controller.py
- node_group_affinity.py
- optional_roles.py
- registry.py
- service.py
- __init__.py
- base.py
- device_direct.py
- __init__.py
- engine.py
- __init__.py
- service.py
- topology.py
- __init__.py
- config.py
- metrics.py
- utils.py
- __init__.py
- actor_group.py
- genrm.py
- placement_group.py
- ray_actor.py
- rollout.py
- rollout_validation.py
- teacher_manager.py
- train_actor.py
- utils.py
- __init__.py
- coordination.py
- __init__.py
- base_types.py
- dynamic_sampling_filters.py
- __init__.py
- dapo_genrm.py
- deepscaler.py
- f1.py
- geo3k.py
- gpqa.py
- ifbench.py
- math_dapo_utils.py
- math_utils.py
- mopd.py
- multiple_choice.py
- openr1mm.py
- registry.py
- __init__.py
- base_types.py
- data_source.py
- forge_load.py
- on_policy_distillation.py
- request_permit.py
- sglang_rollout.py
- __init__.py
- radix_tree.py
- radix_tree_middleware.py
- __init__.py
- router.py
- __init__.py
- chat_template.py
- chat_template_patch.py
- multimodal.py
- qwen_chat_template_patch.py
- sample.py
- streaming.py
- __init__.py
- ppl.py
- runner.py
- __init__.py
- loop.py
- runner.py
- __init__.py
- bootstrap.py
- debug_print.py
- runtime.py
- __init__.py
- __init__.py
- deploy_metrics_service.py
- train.py
- visualize.py
- __init__.py
- bridge.py
- model.py
- provider.py
- __init__.py
- model.py
- processor.py
- __init__.py
- chat_template.jinja
- configuration.py
- vision.py
- __init__.py
- indexer.py
- sparse_mla.py
- tilelang_indexer_bwd.py
- tilelang_indexer_fwd.py
- tilelang_sparse_mla_bwd.py
- tilelang_sparse_mla_fwd.py
- __init__.py
- dsa_attention.py
- glm5_bridge.py
- glm5_provider.py
- __init__.py
- model.py
- rope.py
- text_model.py
- transformer_block.py
- transformer_config.py
- utils.py
- __init__.py
- qwen3_omni_bridge.py
- qwen3_omni_provider.py
- __init__.py
- __init__.py
- autoscaler.yaml
- autoscaler_service.py
- config.py
- metrics_collector.py
- monitor.py
- scaling_decision.py
- __init__.py
- data.py
- data_utils.py
- mask_utils.py
- processing_utils.py
- processor_pool.py
- seqlen_balancing.py
- stream_dataloader.py
- streaming_dataset.py
- __init__.py
- send_to_sglang.py
- __init__.py
- command_utils.py
- typer_utils.py
- __init__.py
- apprise.py
- clearml.py
- tensorboard.py
- wandb.py
- __init__.py
- client.py
- metric_checker.py
- metric_utils.py
- metrics_service_adapter.py
- service.py
- timeline_trace.py
- __init__.py
- audio_utils.py
- config.py
- image_utils.py
- process.py
- stats.py
- video_utils.py
- __init__.py
- opd_main_worker.py
- opd_opsd_worker.py
- opd_sglang_patch.py
- opd_utils.py
- __init__.py
- convert_moe_int4_to_bf16.py
- fp8.py
- fp8_checkpoint.py
- __init__.py
- data_fields.py
- eval_config.py
- flops_counter.py
- ppo_utils.py
- routing_replay.py
- tensor_backper.py
- train_dump_utils.py
- train_metric_utils.py
- __init__.py
- __main__.py
- server.py
- templates.py
- tui.py
- __init__.py
- arguments.py
- async_utils.py
- checkpoint_write_patch.py
- device.py
- distributed_utils.py
- env.py
- genrm_client.py
- health_monitor.py
- health_system.py
- hf_export.py
- hf_page_cache.py
- http_utils.py
- log_style.py
- logging_utils.py
- megatron_bridge_utils.py
- megatron_peft_utils.py
- memory_utils.py
- misc.py
- model_source.py
- profile_utils.py
- reload_utils.py
- reloadable_process_group.py
- rocm_checkpoint_writer.py
- rotate_ckpt.py
- s3_model_loader.py
- sft_utils.py
- timer.py
- tracking_utils.py
- types.py
- utils.py
- __init__.py
- _version.py
- benchmark_request_permit.py
- benchmark.sh
- compare_sglang_megatron_dotsocr.py
- compare_sglang_megatron_dotsocr_packed.py
- run-compare-dotsocr-packed.sh
- run-compare-dotsocr.sh
- test_nccl_comms.py
- local-klx.sh
- local-npu.sh
- local.sh
- ray-job-npu.sh
- ray-job.sh
- runtime-env-klx.sh
- spmd-multinode.sh
- dotsocr2.sh
- glm4.7-30B-A3B.sh
- glm5-744B-A40B.sh
- kimi-k2.6.sh
- qwen3-0.6B.sh
- qwen3-30B-A3B.sh
- qwen3-32B.sh
- qwen3-4B.sh
- qwen3-8B.sh
- qwen3-omni-30B-A3B.sh
- qwen3-vl-2B.sh
- qwen3-vl-30B-A3B.sh
- qwen3-vl-4B.sh
- qwen3-vl-8B.sh
- qwen35-122B-A10B.sh
- qwen35-27B.sh
- qwen35-35B-A3B.sh
- qwen35-397B-A17B.sh
- qwen35-4B.sh
- qwen35-9B.sh
- qwen36-27B.sh
- qwen36-35B-A3B.sh
- __init__.py
- _pyspy_dump.sh
- convert_fp8_to_bf16.py
- convert_hf_to_fp8.py
- convert_hf_to_int4.py
- convert_moe_int4_to_bf16.py
- convert_torch_dist_to_hf_bridge.py
- generate_openapi.py
- kill_for_ray.sh
- model_upload_hook.py
- process_aime.py
- process_avqa.py
- process_nextqa.py
- process_openr1.py
- process_tool_chat.py
- repro_megatron_bridge_load.py
- repro_qwen35_moe_bridge_load_tp4pp2.sh
- run_on_each_ray_node.py
- run-qwen3-4B-8xgpu-genrm.sh
- run-qwen3-4B-pr-8xgpu.sh
- run-dotsocr2-8xgpu-hybrid.sh
- run-dotsocr2-8xgpu.sh
- run-kimi-k2.6-256xgpu-int4.sh
- run-qwen3-30B-A3B-omni-16xgpu-async.sh
- run-qwen3-30B-A3B-omni-16xgpu-video.sh
- run-qwen3-30B-A3B-omni-16xgpu.sh
- run-qwen3-vl-30B-A3B-16xgpu-async.sh
- run-qwen3-vl-30B-A3B-8xgpu.sh
- run-qwen3-vl-4B-2xgpu.sh
- run-qwen3-vl-4B-8xgpu.sh
- run-qwen3-vl-4B-geo3k-8xgpu.sh
- run-qwen35-27B-8xgpu-openr1mm-colocate.sh
- run-qwen35-27B-8xgpu-openr1mm-hybrid-async.sh
- run-qwen35-35B-A3B-8xgpu-openr1mm-r3.sh
- run-qwen35-35B-A3B-8xgpu-vpp.sh
- run-qwen35-35B-A3B-8xklx.sh
- run-qwen35-397B-A17B-128xgpu.sh
- run-qwen35-9B-8xgpu-openr1mm-async.sh
- run-qwen35-9B-8xgpu-openr1mm-hybrid-async.sh
- run-qwen35-9B-8xgpu-video.sh
- run-qwen35-9B-8xklx-openr1mm-sync.sh
- run-qwen36-35B-A3B-8xgpu-image.sh
- run-qwen36-35B-A3B-fp8-8xgpu-image.sh
- run-qwen36-35B-A3B-int4-8xgpu-image.sh
- eval_pokemon_sft.py
- run-qwen3-vl-4B-math-8xgpu.sh
- run-qwen3-vl-4B-pokemon-1xgpu.sh
- run-qwen3-vl-4B-pokemon-8xgpu.sh
- run-qwen3.5-35B-A3B-mtp-sft-16xgpu.sh
- run-qwen3.5-35B-A3B-mtp-sft-8xklx.sh
- run-qwen3.5-35B-A3B-pokemon-lora-mtp-8xgpu.sh
- run-qwen3.5-35B-A3B-pokemon-mtp-8xgpu.sh
- run-qwen3.5-397B-A17B-mtp-sft-128k-128xgpu.sh
- run-qwen3.5-9B-8xklx.sh
- run-qwen3.5-9B-math-8xgpu.sh
- run-qwen3.5-9B-math-dynamic-cp-8xgpu.sh
- run-glm5-744B-A40B-128xgpu.sh
- run-kimi-k2.6-256xgpu-int4.sh
- run-qwen3-30B-A3B-16xgpu-async.sh
- run-qwen3-30B-A3B-16xgpu.sh
- run-qwen3-30B-A3B-8xgpu.sh
- run-qwen3-30B-A3B-fp8-8xgpu.sh
- run-qwen3-30B-A3B-int4-8xgpu.sh
- run-qwen3-4B-16xgpu.sh
- run-qwen3-4B-4xgpu-async-npu.sh
- run-qwen3-4B-4xgpu-async.sh
- run-qwen3-4B-4xnpu-colocate.sh
- run-qwen3-4B-8xgpu-async-npu.sh
- run-qwen3-4B-8xgpu-async.sh
- run-qwen3-4B-8xgpu-hybrid-async.sh
- run-qwen3-4B-8xgpu-ppo.sh
- run-qwen3-4B-8xgpu.sh
- run-qwen3-4B-8xklx.sh
- run-qwen3-4B-fp16-8xgpu.sh
- run-qwen3-4B-lora-adapter-8xgpu-async.sh
- run-qwen3-4B-lora-merge-8xgpu.sh
- run-qwen35-35B-A3B-16xgpu-async.sh
- run-qwen35-35B-A3B-16xgpu.sh
- run-qwen35-35B-A3B-16xklx.sh
- run-qwen35-35B-A3B-16xnpu-async.sh
- run-qwen35-35B-A3B-16xnpu-colocate.sh
- run-qwen35-35B-A3B-8xgpu-colocate.sh
- run-qwen35-35B-A3B-8xnpu-colocate.sh
- run-qwen35-35B-A3B-mtp-16xgpu.sh
- run-qwen35-35B-A3B-mtp-8xgpu.sh
- run-qwen35-397B-A17B-128xgpu.sh
- run-qwen35-9B-4xnpu-colocate.sh
- run-qwen35-9B-8xgpu-async.sh
- run-qwen35-9B-8xgpu-ppo.sh
- run-qwen35-9B-8xgpu.sh
- run-qwen35-9B-8xklx.sh
- run-qwen35-9B-8xnpu-async.sh
- run-qwen35-9B-mtp-8xgpu.sh
- run-qwen36-27B-8xklx.sh
- run-qwen36-35B-A3B-8xgpu-vpp.sh
- run-qwen36-35B-A3B-8xgpu.sh
- run-qwen36-35B-A3B-8xklx.sh
- __init__.py
- code-quality-checklist.md
- python-ml-checklist.md
- removal-plan.md
- security-checklist.md
- solid-checklist.md
- SKILL.md
- official_best_practices.md
- skill_examples.md
- SKILL.md
- case-rollout-eval-onload-hang.md
- SKILL.md
- SKILL.md
- content-verification-guide.md
- doc-template.md
- SKILL.md
- LICENSE
- SKILL.md
- examples.md
- SKILL.md
- SKILL.md
- baselines.md
- rules.md
- SKILL.md
- migration_mapping.md
- SKILL.md
- classify_patch.sh
- SKILL.md
- SKILL.md
- gitleaks.md
- prompt-a-main-to-dev.md
- prompt-b-dev-to-main.md
- check_duplicate_defs.py
- plan_github_to_dev.py
- SKILL.md
- migration_mapping.md
- SKILL.md
- __init__.py
- test_quantizer_fp8.py
- __init__.py
- test_broadcast_converted.py
- test_dtype_codes.py
- test_lora_weight_sync.py
- __init__.py
- test_actor_agree_drained.py
- test_actor_http_timeout.py
- test_critic_value_head_ci.py
- test_data_vpp.py
- test_fp16_optimizer_config.py
- test_frozen_weight_dgrad.py
- test_gdn_cp_reassembly.py
- test_model_forward_only_dynamic_cp.py
- test_model_provider_vpp.py
- test_mtp_rl_forward_kwargs.py
- test_mtp_sft_forward_kwargs.py
- test_opd_loss_aggregation.py
- test_ppo_gae_parity.py
- test_reinforce_plus_plus_loss.py
- test_reinforce_plus_plus_wiring.py
- test_rloo_cp_reduction.py
- test_rloo_policy_loss_dispatch.py
- test_save_hf_fp8.py
- test_save_hf_post_hook.py
- test_sft_chunked_ce.py
- test_sft_data.py
- test_sft_train_actor_eval.py
- test_sft_train_data_fields.py
- __init__.py
- test_arguments.py
- test_genrm_offload_drain.py
- test_router_registration.py
- __init__.py
- __init__.py
- test_actor_sft_partition.py
- test_genrm_engine_pick.py
- test_rollout_weight_update_handshake.py
- test_sft.py
- __init__.py
- test_controller_s3_model_cleanup.py
- test_controller_sft_colocate.py
- test_controller_sft_loop.py
- test_optional_roles.py
- test_registry_reinforce_plus_plus.py
- test_registry_rloo.py
- test_registry_sft.py
- test_service_affinity.py
- predict_prompts.jsonl
- sft_eval_messages.jsonl
- sft_eval_prompts.jsonl
- sft_train_messages.jsonl
- test_dcs_weight_conversion.py
- test_rollout_engine_recovery.py
- __init__.py
- conftest.py
- test_coordination.py
- test_genrm_recovery.py
- test_opd_teacher_controller_colocate.py
- test_scale_in.py
- test_scale_out.py
- test_state_machine.py
- test_teacher_manager.py
- test_teacher_manager_factory.py
- test_teacher_sglang_overrides.py
- test_utils.py
- test_weight_sync.py
- __init__.py
- __init__.py
- custom_reward_fixtures.py
- test_custom_reward_worker.py
- test_math_dapo_utils.py
- test_reward_registry.py
- test_reward_router.py
- test_reward_worker.py
- __init__.py
- test_base_types.py
- test_data_source.py
- test_on_policy_distillation_payload.py
- test_on_policy_distillation_teacher_failures.py
- test_request_permit.py
- test_sglang_rollout_diagnostics.py
- __init__.py
- test_chat_template.py
- test_chat_template_patch.py
- test_multimodal.py
- test_qwen_chat_template_patch.py
- test_streaming.py
- __init__.py
- test_ppl.py
- __init__.py
- test_run_predict_loop.py
- __init__.py
- __init__.py
- __init__.py
- test_train_affinity.py
- __init__.py
- test_agent_messages.py
- test_env_extractors.py
- __init__.py
- conftest.py
- test_client.py
- test_convert_dataset.py
- test_gateway_app.py
- test_gateway_config.py
- test_gateway_service.py
- test_prepare_r2e_gym.py
- test_protocol.py
- test_recipe_paths.py
- test_result.py
- test_run_adapter.py
- test_run_agent_app.py
- test_verbose_logging.py
- __init__.py
- test_deepeyes_processor_patch.py
- test_mini_swe_agent_server.py
- test_sft_smoke.py
- test_chat_template.py
- __init__.py
- test_convert_hf_to_fp8.py
- test_convert_hf_to_int4.py
- test_model_upload_hook.py
- test_process_tool_chat.py
- __init__.py
- test_autoscaler_service.py
- test_metrics_collector.py
- test_monitor.py
- test_scaling_decision.py
- __init__.py
- test_check_sample_length_multimodal.py
- test_data_utils.py
- test_encode_executor.py
- test_processing_utils.py
- test_seqlen_balancing.py
- test_streaming_dataset.py
- test_streaming_tq_iterator.py
- test_audio_utils_tar.py
- test_image_utils.py
- test_ppo_utils_grpo.py
- test_reinforce_plus_plus.py
- test_rloo_advantages.py
- __init__.py
- test_arguments_opd_teacher_colocate.py
- test_arguments_reinforce_plus_plus.py
- test_arguments_rloo.py
- test_distributed_masked_normalize.py
- test_flops_counter.py
- test_health_monitor.py
- test_hf_export.py
- test_http_utils.py
- test_logging_utils.py
- test_megatron_peft_utils.py
- test_metric_utils_rloo.py
- test_metrics_service.py
- test_multimodal_rollout_stats.py
- test_reinforce_reward_processing.py
- test_train_metric_utils.py
- test_visualize_tui_sorting.py
- __init__.py
- test_agentic_rollout.py
- test_model_source.py
- test_s3_model_loader.py
- test_s3_shm_consumer_sites.py
- .dockerignore
- .gitignore
- .gitleaks.toml
- .pre-commit-config.yaml
- AGENTS.md
- CLAUDE.md
- CONTRIBUTING.md
- LICENSE
- Makefile
- MANIFEST.in
- package.json
- pyproject.toml
- README.md
- README_zh.md
- requirements.txt
- setup.py
- versioneer.py
# 설치 가이드
1. 코드 내려받기
git clone https://github.com/redai-infra/Relax
깃허브에서 프로젝트 코드 전체를 내 컴퓨터로 내려받습니다.
cd Relax
방금 내려받은 프로젝트 폴더 안으로 이동합니다.
2. Docker
쉬움 추천사전 준비물
- Git GitHub에서 프로젝트 코드를 내려받으려면 필요합니다.
- Docker Desktop 컨테이너를 빌드하고 실행하려면 필요합니다. 설치 후 실행해서 백그라운드에 켜두세요.
docker pull ghcr.io/redai-infra/relaxrl:latest
이 명령어를 터미널에 그대로 입력해 실행하세요.
docker run -it --gpus all --ipc=host --network=host \
빌드된 이미지를 실제 컨테이너로 실행합니다.
<img src="https://img.shields.io/badge/Docker-Image-blue?logo=docker" alt="Docker Image">
이 명령어를 터미널에 그대로 입력해 실행하세요.
터미널에 docker compose ps 를 입력해 컨테이너들이 Up 상태인지 확인하세요. README에 포트 번호가 적혀있다면 브라우저에서 http://localhost:포트번호 로 접속해보세요.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
3. Node.js
쉬움사전 준비물
npm install
package.json에 명시된 라이브러리들을 내려받아 설치합니다.
npm start
개발/실행 서버를 켭니다.
명령어 실행 후 터미널에 나타나는 주소(보통 http://localhost:3000 형태)를 브라우저에서 열어보세요.
4. Python
쉬움사전 준비물
cd /root/Relax && pip install -e .
방금 내려받은 프로젝트 폴더 안으로 이동합니다.
에러 메시지 없이 실행되고 터미널에 안내 문구가 출력되면 정상입니다.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
5. Make
보통사전 준비물
- Git GitHub에서 프로젝트 코드를 내려받으려면 필요합니다.
- Make Linux/macOS는 보통 기본 설치되어 있습니다. Windows는 별도 설치(예: MSYS2, WSL)가 필요합니다.
make
생성된 빌드 설정을 바탕으로 실제 컴파일을 진행해 실행 파일을 만듭니다.
에러 없이 끝나면 성공입니다. 생성된 실행 파일을 직접 실행해보세요.
// repository documentation
Was this content helpful?
(0 ratings)
