sglang-jax
JAX backend for SGL
파일 탐색기
최종 버전 다운로드 (.zip)- accuracy_benchmark.sh
- report-template.md
- SKILL.md
- speed_benchmark.sh
- report-template.md
- SKILL.md
- capture-recipes.md
- pallas-kernel-profiling.md
- SKILL.md
- config.yaml
- action.yml
- action.yml
- 1-bug-report.yml
- 2-feature-request.yml
- 3-documentation.yml
- coordinator.yml
- lint.yml
- multi-host-test.yml
- nightly-slack-notify.yml
- nightly-test-daily.yml
- nightly-test-summize.yml
- postmerge-pr-comment.yml
- pr-test.yml
- publish.yml
- release-docker.yml
- release-pypi.yml
- release.yml
- rerun-test.yml
- runner-utilization.yml
- slash-command.yml
- tag.yml
- weekly-test.yml
- pull_request_template.md
- __init__.py
- bench_sglang_jax.py
- __init__.py
- bench_recurrent_reuse_sweep.py
- bench_unified_radix_ab.py
- bench_biased_topk.py
- tune_biased_topk_bt.py
- __init__.py
- bench_flashattention.py
- get_block_spec_config.py
- get_block_spec_config_v3.py
- tune_target_verify_matrix_v3.py
- tune_target_verify_v3.py
- utils.py
- __init__.py
- bench_gdn.py
- __init__.py
- bench_grouped_topk.py
- tune_grouped_topk_bt.py
- bench_megablox_gmm.py
- tune_gmm_v1_block_sizes.py
- __init__.py
- bench_mla.py
- get_block_spec_config.py
- get_block_spec_config_mla.py
- utils.py
- __init__.py
- example_shapes_glm51_tp64.txt
- tune_blockwise_block_sizes.py
- bench_speculative_kernels.py
- bench_update_kv_cache.py
- utils.py
- __init__.py
- __init__.py
- bench_ep_moe.py
- bench_fused_moe.py
- bench_gmm.py
- run_fused_moe_ablation.sh
- utils.py
- __init__.py
- utils.py
- architecture.svg
- chunked_prefill_architecture.svg
- lora_architecture.jpg
- lora_dependency_graph.svg
- scheduler_workflow.jpg
- structured_output_arch.svg
- 01-request-flow.svg
- 01-system-overview.excalidraw
- 01-system-overview.svg
- 02-request-routing.excalidraw
- 02-request-routing.svg
- 02-tokenization-flow.svg
- 03-batch-building-flow.svg
- 03-event-loop.excalidraw
- 03-event-loop.svg
- 03-request-state-machine.excalidraw
- 03-request-state-machine.svg
- 04-batch-conversion-flow.svg
- 04-model-loading-pipeline.svg
- 05-model-class-hierarchy.excalidraw
- 05-model-class-hierarchy.svg
- 06-moe-dataflow.svg
- 07-kv-cache-allocation-flow.svg
- 07-kv-cache-layers.excalidraw
- 07-kv-cache-layers.svg
- 08-kernel-catalog.excalidraw
- 08-kernel-catalog.svg
- 09-draft-verify-cycle.excalidraw
- 09-draft-verify-cycle.svg
- 09-speculative-decoding-flow.svg
- 10-lora-integration.excalidraw
- 10-lora-integration.svg
- 11-quantization-flow.svg
- 11-quantization-pipeline.excalidraw
- 11-quantization-pipeline.svg
- 12-multimodal-request-flow.svg
- 01-architecture-overview.md
- 02-entrypoints-and-tokenization.md
- 03-scheduler.md
- 04-model-executor.md
- 05-models.md
- 06-layers-and-attention.md
- 07-kv-cache.md
- 08-pallas-kernels.md
- 09-speculative-decoding.md
- 10-lora.md
- 11-quantization.md
- 12-multimodal.md
- 13-configuration-reference.md
- 14-pd-disaggregation.md
- index.md
- index.md
- DeepSeek-R1.md
- DeepSeek-V2.md
- DeepSeek-V3.md
- GLM-4.5.md
- Gemma2.md
- Grok2.md
- Ling-2.6.md
- Llama3.1.md
- Llama3.3-70B.md
- Kimi-Linear.md
- Qwen.md
- Qwen2.5-VL.md
- Qwen3-MoE.md
- Qwen3.md
- MiMo-7B.md
- MiMo-V2-Flash.md
- MiMo-V2.5-Pro.md
- index.md
- basic-api-usage.md
- launch-flags-reference.md
- tpu-topology-reference.md
- gke-indexed-job.md
- index.md
- single-host-docker.md
- skypilot.md
- troubleshooting.md
- tpu_resources_guide.md
- Wan2.1.md
- Wan2.2.md
- index.md
- install.md
- add-recipe.md
- docs.json
- index.md
- gke-indexed-job.md
- index.md
- single-host-docker.md
- skypilot.md
- troubleshooting.md
- benchmark_and_profiling.md
- ci_architecture.md
- contribution_guide.md
- how_to_join_community.md
- index.md
- jax_tutorial.md
- release_process.md
- tpu_resources_guide.md
- attention_backend.md
- chunked_prefill.md
- global_jit_compile.md
- index.md
- quantization.md
- radix_cache.md
- run_in_pathways.md
- server_arguments.md
- speculative_decoding.md
- structured_output.md
- gke-tpu-install.md
- install.md
- sglang-jax-gpu-install.md
- index.md
- multimodal_usage.md
- conf.py
- cookbook_overview.md
- index.rst
- Makefile
- README.md
- requirements.txt
- __init__.py
- bailing_hybrid.py
- dtype_config.py
- gemma4.py
- kimi_linear.py
- load_config.py
- model_config.py
- quantization_config.py
- qwen3_5.py
- __init__.py
- base_grammar_backend.py
- bitmask_ops.py
- llguidance_backend.py
- __init__.py
- kv_manager.py
- transfer.py
- __init__.py
- capacity.py
- core.py
- metrics.py
- multihost_sync.py
- zmq_notifier.py
- __init__.py
- conn.py
- wrapper.py
- __init__.py
- conn.py
- wrapper.py
- __init__.py
- bootstrap.py
- debug_utils.py
- decode.py
- decode_watchdog.py
- factory.py
- host_ip.py
- launch_router.py
- mini_lb.py
- mini_lb_helpers.py
- pathways_pd.py
- pathways_scheduler.py
- pd_auth.py
- prefill.py
- req_time_stats.py
- router.py
- router_args.py
- run_bootstrap.py
- runtime.py
- __init__.py
- protocol.py
- serving_base.py
- serving_chat.py
- serving_completions.py
- serving_embedding.py
- serving_rerank.py
- serving_score.py
- usage_processor.py
- utils.py
- engine.py
- EngineBase.py
- http_server.py
- __init__.py
- deepseek.py
- __init__.py
- expert_location.py
- base_format_detector.py
- core_types.py
- ebnf_composer.py
- function_call_parser.py
- glm47_moe_detector.py
- glm4_moe_detector.py
- mimo_detector.py
- qwen25_detector.py
- qwen3_coder_detector.py
- utils.py
- __init__.py
- kernel.py
- __init__.py
- tuned_block_sizes.py
- __init__.py
- ref.py
- sparse_mla.py
- sparse_mla_prefill.py
- streamindex_topk.py
- __init__.py
- kernel.py
- tuned_block_configs.py
- __init__.py
- bench_compare.py
- bench_se_mlp.py
- bench_v2.py
- kernel.py
- tuned_block_configs.py
- __init__.py
- __init__.py
- compute_conv1d.py
- compute_gdn.py
- config.py
- memory_ref.py
- metadata.py
- PROVENANCE.md
- vmem_ldst.py
- wrapper.py
- __init__.py
- fused_chunk_parallel_adapter.py
- gated_delta.py
- common.py
- gmm.py
- gmm_v2.py
- tuned_block_sizes.py
- megablox_gmm_backend.py
- __init__.py
- kernel.py
- __init__.py
- tuned_block_sizes.py
- __init__.py
- kda.py
- naive.py
- __init__.py
- ref.py
- __init__.py
- kernel.py
- tuned_block_sizes.py
- __init__.py
- paged_attention.py
- __init__.py
- blockwise_kernel.py
- kernel.py
- tuned_block_sizes.py
- util.py
- blockwise_utils.py
- kernel.py
- ragged_paged_attention.py
- ragged_paged_attention_v3.py
- tuned_block_sizes.py
- tuned_block_sizes_v3.py
- util.py
- __init__.py
- native.py
- simple_gla.py
- simple_gla_fused.py
- build_eagle_tree_structure_kernel.py
- kernel.py
- tree_speculative_sampling_target_only_kernel.py
- verify_tree_greedy_kernel.py
- tuned_block_sizes.py
- update_kv_cache.py
- perf.py
- _pathways_compat.py
- fused_mlp.py
- gated_rmsnorm.py
- group_rmsnorm.py
- __init__.py
- gdn_backend.py
- kda_backend.py
- lightning_backend.py
- short_convolution.py
- base_attn_backend.py
- dsa_sparse_backend.py
- flashattention_backend.py
- hybrid_linear_attn_backend.py
- mla_backend.py
- native_backend.py
- utils.py
- activation.py
- binary_search.py
- embeddings.py
- fused_moe.py
- gate.py
- layernorm.py
- linear.py
- logits_processor.py
- moe.py
- radix_attention.py
- radix_lightning_attention.py
- radix_linear_attention.py
- routed_experts_capturer.py
- sampler.py
- base_backend.py
- bgmv_backend.py
- context_manager.py
- layers.py
- lora.py
- lora_config.py
- lora_manager.py
- lora_memory_pool.py
- lora_registry.py
- utils.py
- __init__.py
- communication.py
- detokenizer_manager.py
- dp_rank_assignment.py
- dp_schedule_policy.py
- io_struct.py
- schedule_batch.py
- schedule_policy.py
- scheduler.py
- scheduler_metrics_mixin.py
- scheduler_output_processor_mixin.py
- scheduler_profiler_mixing.py
- template_manager.py
- tiktoken_tokenizer.py
- tokenizer_manager.py
- tp_worker.py
- tp_worker_overlap_thread.py
- utils.py
- __init__.py
- full_component.py
- recurrent_component.py
- tree_component.py
- __init__.py
- allocator.py
- base_prefix_cache.py
- cache_init_params.py
- chunk_cache.py
- common.py
- hicache_controller.py
- host_kv_pool.py
- kv_cache_builder.py
- memory_pool.py
- radix_cache.py
- recurrent_state_pool.py
- registry.py
- swa_radix_cache.py
- unified_radix_cache.py
- aot_dispatch.py
- base_model_runner.py
- compilation_manager.py
- forward_batch_info.py
- model_runner.py
- model_runner_kv_cache_mixin.py
- __init__.py
- arch.py
- loader.py
- bailing_moe.py
- bailing_moe_linear.py
- deepseek_v3.py
- dflash.py
- gemma2.py
- gemma4.py
- glm4_moe.py
- glm5_moe.py
- grok.py
- kimi_k25_vl_generation.py
- kimi_linear.py
- llama.py
- llama_eagle3.py
- mimo.py
- mimo_mtp.py
- mimo_v2_flash.py
- mimo_v2_nextn.py
- mimo_v2_pro.py
- minimax_m2.py
- qwen.py
- qwen2.py
- qwen2_moe.py
- qwen3.py
- qwen3_5.py
- qwen3_moe.py
- registry.py
- umt5.py
- modality_enum.py
- ServerArgs.py
- flux_model_config.py
- wan_model_config.py
- kimi_k25_config.py
- __init__.py
- mimo_audio_backbone_config.py
- mimo_audio_config.py
- qwen_2_5_vl_config.py
- flux_vae_config.py
- vae_base_config.py
- wan_vae_config.py
- config_registry.py
- multimodal_base_config.py
- http_server.py
- flash_attention.py
- get_block_spec_config.py
- tuned_block_sizes.py
- varlen_attention.py
- varlen_tuned_block_sizes.py
- flash_attention_backend.py
- layer.py
- adalayernorm.py
- image_hash.py
- layernorm.py
- mlp.py
- rotary_embedding.py
- visual_embedding.py
- audio_backbone_scheduler.py
- audio_scheduler.py
- diffusion_scheduler.py
- embed_scheduler.py
- encoder_scheduler.py
- vae_scheduler.py
- vit_scheduler.py
- device_manager.py
- global_scheduler.py
- io_struct.py
- mrope_utils.py
- multimodal_detokenizer.py
- multimodal_tokenizer.py
- prompt_builder.py
- schedule_batch.py
- stage.py
- utils.py
- __init__.py
- audio_backbone_model_runner.py
- audio_backbone_model_worker.py
- audio_model_runner.py
- audio_model_worker.py
- diffusion_model_runner.py
- diffusion_model_worker.py
- embed_model_runner.py
- embed_model_worker.py
- encoder_model_runner.py
- encoder_model_worker.py
- vae_model_runner.py
- vae_model_worker.py
- vit_model_runner.py
- vit_model_worker.py
- flow_match_euler_discrete_scheduler.py
- flow_unipc_multistep_scheduler.py
- flux.py
- flux_dit_weights_mapping.py
- base.py
- clip.py
- t5.py
- kimi_k25_vit.py
- kimi_k25_vl_generation.py
- __init__.py
- mimo_audio_backbone.py
- mimo_audio_backbone_weights_mapping.py
- mimo_audio_tokenizer.py
- mimo_audio_tokenizer_weights_mapping.py
- qwen2_5_vit.py
- qwen2_5_vl_generation.py
- audio_encoder.py
- qwen3_omni_thinker.py
- qwen3_omni_thinker_embedding.py
- vision_encoder.py
- __init__.py
- flux1_dev_stage_config.yaml
- mimo_audio_stage_config.yaml
- qwen2_5_vl_stage_config.yaml
- qwen2_5_vl_stage_config_tp4.yaml
- qwen3_omni_stage_config.yaml
- wan2_1_stage_config.yaml
- wan2_2_stage_config.yaml
- yaml_registry.py
- autoencoder.py
- common.py
- flux_vae_weight_mappings.py
- wan_dit.py
- wan_dit_weights_mapping.py
- commons.py
- vae_weights_mappings.py
- wanvae.py
- tokenizer_utils.py
- __init__.py
- frequency_penalty.py
- min_new_tokens.py
- orchestrator.py
- presence_penalty.py
- __init__.py
- sampling_batch_info.py
- sampling_params.py
- base_worker.py
- dflash_info.py
- dflash_util.py
- dflash_worker.py
- draft_extend_fused.py
- eagle_draft_worker.py
- eagle_info.py
- eagle_util.py
- eagle_worker.py
- multi_layer_draft_worker.py
- multi_layer_eagle_worker.py
- overlap_utils.py
- relay_buffer.py
- spec_info.py
- spec_utils.py
- fp8.yaml
- fp8_bailing.yaml
- fp8_block_128_dynamic.yaml
- fp8_deepseek_v3.yaml
- fp8_grok.yaml
- fp8_qwen3_30b_a3b.yaml
- fp8_w8a8.yaml
- int8.yaml
- int8_block_128_dynamic.yaml
- int8_moe_block_128_linear_channel_dynamic.yaml
- int8_w8a8.yaml
- debug_utils.py
- quantization_utils.py
- __init__.py
- common_utils.py
- debug_utils.py
- jax_utils.py
- mesh_utils.py
- parallel_utils.py
- profiling_utils.py
- tunix_utils.py
- weight_utils.py
- __init__.py
- conversation.py
- hf_transformers_utils.py
- jinja_template_utils.py
- memory_profiler.py
- precision_tracer.py
- reasoning_parser.py
- server_args.py
- test_bitmask_ops.py
- __init__.py
- biased_topk_test.py
- fused_moe_v1_test.py
- fused_moe_v2_test.py
- gmm_test.py
- grouped_topk_test.py
- kda_test.py
- moe_block_quant_test.py
- quantized_linear_test.py
- simple_gla_fused_test.py
- test_gdn_fused_chunk_parallel_provenance.py
- __init__.py
- cache_hit_kit.py
- __init__.py
- mock_recurrent_state_pool.py
- test_group_rmsnorm.py
- test_lightning_backend.py
- test_lightning_backend_dp.py
- test_merged_column_parallel_linear.py
- test_rotary_embedding.py
- test_sequence_parallel.py
- test_future_token_map.py
- conftest.py
- run_multi_process_radix_cache_test.sh
- test_hicache_controller.py
- test_hicache_e2e.py
- test_hicache_e2e_tpu.py
- test_host_kv_pool.py
- test_host_kv_pool_tpu.py
- test_host_kv_pool_tpu_dp.py
- test_hybrid_req_to_token_pool.py
- test_kv_cache.py
- test_paged_allocator_multi_dp.py
- test_radix_cache.py
- test_req_to_token_pool.py
- test_swa_allocator.py
- test_swa_radix_cache.py
- test_unified_radix_cache.py
- test_unified_radix_tree_flag.py
- test_dflash.py
- test_mimo_v2_nextn.py
- test_minimax_m2.py
- test_qwen3_5.py
- qwen3_omni_moe_thinker_text_prefill.txt
- wan_vae_diffusers_decode_output.npy
- wan_vae_diffusers_encode_output.npy
- test_clip_text_model.py
- test_t5_encoder.py
- __init__.py
- test_diffusion_precision.py
- test_diffusion_scheduler.py
- test_flash_attention_kernel.py
- test_flow_match_euler_discrete_scheduler.py
- test_flux_autoencoder.py
- test_flux_transformer_2d_weight_loading.py
- test_kimi_k25_weight_mapping.py
- test_mimo_audio_tokenizer.py
- test_mimo_audio_tokenizer_model.py
- test_qwen3_omni_moe_encoder.py
- test_qwen3_omni_vision_encoder.py
- test_qwen_omni_thinker.py
- test_stage_config_routing.py
- test_vae.py
- test_vae_scheduler.py
- test_varlen_attention_kernel.py
- test_wan2_1_dit.py
- test_wan2_1_dit_weight_loading.py
- test_wan_vae_precision.py
- test_dflash_info.py
- test_dflash_server_args.py
- test_dflash_worker.py
- test_eagle_tree_build.py
- test_eagle_utils.py
- test_spec_dp_shapes.py
- test_spec_info.py
- __init__.py
- flashattention_common.py
- long_prompt.txt
- runners.py
- test_bailing_moe_linear.py
- test_compilation_manager.py
- test_flashattention_dp.py
- test_flashattention_gqa.py
- test_flashattention_mha.py
- test_flashattention_misc.py
- test_gdn_attention.py
- test_gdn_attention_dp.py
- test_gdn_fused_chunk_parallel_prefill.py
- test_gdn_fused_chunk_parallel_prefill_dp.py
- test_gdn_fused_chunk_parallel_state_contract.py
- test_gdn_prefill_dispatch.py
- test_hf_transformers_fastokens.py
- test_kda_attention.py
- test_kda_attention_dp.py
- test_kernel_utils.py
- test_kv_cache_builder.py
- test_linear_tp.py
- test_mesh.py
- test_mixed_chunk_dp.py
- test_mla_attention.py
- test_model_runner_kv_cache_mixin.py
- test_moe_topk.py
- test_moe_weight_loader_sharding.py
- test_pagedattention.py
- test_profiler_tracer_levels.py
- test_sampler.py
- test_sampler_deterministic_cond.py
- test_scheduler_chunked_ownership.py
- test_scheduler_idle_check.py
- test_short_conv.py
- test_swa_schedule_budget.py
- test_tp_worker_overlap_thread.py
- test_tuned_block_sizes_mla.py
- test_tuned_block_sizes_v3.py
- test_utils.py
- test_weight_loader_qkv_split.py
- tool_parser_test_config.py
- moe_parity_repro.py
- trace_diff.py
- __init__.py
- __main__.py
- bench_offline_throughput.py
- bench_one_batch.py
- bench_one_batch_server.py
- bench_serving.py
- check_env.py
- global_config.py
- launch_server.py
- profiler.py
- raiden.py
- utils.py
- version.py
- pyproject.toml
- test_ci_scripts.py
- test_notify.py
- ci_common.py
- failure_classifier.py
- github_output.py
- runner_utilization_report.py
- bisect_preflight.py
- check_trend_threshold.py
- ci_failure_issues.py
- coordinator_decide.py
- finish_check.py
- parse_bisect_comment.py
- plot_acc.py
- plot_perf.py
- publish.py
- publish_trend.py
- regression_notify.py
- release_gate.py
- release_metadata.py
- slack_notify.py
- slash_command_dispatch.py
- slash_command_parse.py
- slash_command_stage_dispatch.py
- trend_alert_issues.py
- deploy.sh
- pd_raiden_smoke_job.yaml.in
- pd_singlehost_eval_job.yaml
- build_raiden_wheel.sh
- run_raiden_smoke.sh
- .zshrc
- analyze_hot_experts.py
- cleanup_cluster.sh
- get_cluster_name.sh
- inspect_expert_dist.py
- killall_sglang.sh
- launch_tpu.sh
- plot_expert_balance.py
- tpu_resource.sky.yaml
- test_pathways_pd.py
- test_pd_auth.py
- test_pd_bootstrap.py
- test_pd_decode.py
- test_pd_lifecycle.py
- test_pd_metrics.py
- test_pd_prefill.py
- test_pd_raiden.py
- test_pd_router.py
- test_pd_transfer.py
- sglang_mmlu.py
- simple_eval_aime25.py
- simple_eval_aime26.py
- simple_eval_common.py
- simple_eval_csimpleqa.py
- simple_eval_gpqa.py
- simple_eval_gsm8k.py
- simple_eval_humaneval.py
- simple_eval_math.py
- simple_eval_mgsm.py
- simple_eval_mmlu.py
- test_simple_eval_common.py
- test_glm47_detector.py
- test_glm4_moe_detector.py
- test_mimo_detector.py
- test_qwen25_detector.py
- test_qwen3_coder_detector.py
- __init__.py
- test_ref.py
- test_sparse_mla_prefill_parity.py
- test_streamindex_topk.py
- multi_prompts_decode_output_cpu.json
- single_prompt_decode_output_cpu.json
- single_prompt_prefill_output_cpu.json
- hf_runner.py
- dump_hf_lora_output.py
- hf_lora_logprobs.json
- test_align_lora_accuracy.py
- test_bgmv_backend.py
- test_dynamic_lora.py
- test_lora_manager_optimization.py
- test_static_lora.py
- __init__.py
- test_aot_dispatch.py
- pd_transfer_matrix.py
- pd_transfer_probe.py
- pd_transfer_smoke.py
- README.md
- test_flux1_dev_models.py
- test_wan2_1_models.py
- ds-v2-lite-fa-mha-v6e-4.yaml
- ds-v2-lite-mla-v6e-4.yaml
- kimi-linear-v6e-4.yaml
- mimo-flash-v6e-4x4.yaml
- qwen3-32b-fa-v6e-4.yaml
- qwen3-32b-perf-v6e-4.yaml
- qwen3-8b-fa-v6e-4.yaml
- qwen3-moe-epmoe-perf-v6e-4.yaml
- qwen3-moe-epmoe-v6e-4.yaml
- qwen3-moe-fp8-epmoe-v6e-4.yaml
- qwen3-moe-fp8-fused-v6e-4.yaml
- qwen3-moe-fused-perf-v6e-4.yaml
- qwen3-moe-fused-v6e-4.yaml
- accuracy_case_runner.py
- launch_from_profile.py
- multi_host_suite.py
- perf_case_runner.py
- profile_loader.py
- README.md
- suite_runner.py
- accuracy_result.v1.yaml
- accuracy_case_runner.py
- perf_case_runner.py
- README.md
- suite_runner.py
- cases.py
- drivers.py
- perf_baseline.py
- profiles.py
- results.py
- __init__.py
- test_openai_server.py
- test_protocol.py
- test_serving_chat.py
- test_serving_completions.py
- test_tool_calls.py
- test_ebnf.py
- test_json_mode.py
- test_structural_tag.py
- __init__.py
- test_openai_api_surface.py
- test_openai_server_params_validation.py
- __init__.py
- test_w8_block_dynamic_quantization.py
- test_w8_moe_block_linear_channel_quantization.py
- test_w8_quantization.py
- test_multi_engines_in_one_process.py
- test_return_routed_experts.py
- run_curl.py
- run_eval.py
- run_suite.py
- test_bench_one_batch.py
- test_bench_recurrent_reuse_sweep.py
- test_bench_serving.py
- test_bench_serving_dense.py
- test_bench_serving_dense_tp_4.py
- test_bench_serving_moe.py
- test_bench_serving_prompt_types.py
- test_bench_unified_radix_ab_args.py
- test_chunked_prefill_size.py
- test_deepseek_v2_lite_models.py
- test_dp_rank_assignment.py
- test_dp_schedule_policy.py
- test_dp_schedule_shape_aware.py
- test_dtype_config_consistency.py
- test_dtype_config_llama.py
- test_engine_determine_generation.py
- test_engine_flush_cache.py
- test_engine_pause_continue.py
- test_eplb.py
- test_eval_accuracy_large.py
- test_eval_accuracy_large_unified_radix.py
- test_features.py
- test_gemma4_models.py
- test_jax_utils.py
- test_logprobs.py
- test_logprobs_dp.py
- test_merge_cache_loc.py
- test_model_loader.py
- test_moe_block_quant_e2e.py
- test_moe_eval_accuracy_large.py
- test_nightly_infra_smoke.py
- test_parallel_utils.py
- test_penalty.py
- test_prepare_for_extend_protected_len.py
- test_qwen1_5_models_dummy.py
- test_qwen2_5_models.py
- test_qwen3_5_models.py
- test_qwen3_models.py
- test_qwen3_moe_models.py
- test_qwen3_omni_vision_alignment.py
- test_reasoning_parser.py
- test_recurrent_boundary_split.py
- test_recurrent_cow_metadata.py
- test_recurrent_split_equivalence.py
- test_recurrent_state_sizing.py
- test_recurrent_track_metadata.py
- test_recurrent_track_scatter.py
- test_retract_decode.py
- test_schedule_batch_dp.py
- test_server_info.py
- test_server_pause_continue.py
- test_sliding_window_attention.py
- test_speculative_decoding.py
- test_srt_engine.py
- test_tokenizer_manager_event.py
- test_umt5_models_alignment.py
- test_unified_radix_cache_serving.py
- README.md
- .codespellrc
- .gitignore
- .isort.cfg
- .pre-commit-config.yaml
- Dockerfile
- LICENSE
- Makefile
- README.md
# 설치 가이드
1. 코드 내려받기
git clone https://github.com/sgl-project/sglang-jax
깃허브에서 프로젝트 코드 전체를 내 컴퓨터로 내려받습니다.
cd sglang-jax
방금 내려받은 프로젝트 폴더 안으로 이동합니다.
2. Docker
쉬움 추천사전 준비물
- Git GitHub에서 프로젝트 코드를 내려받으려면 필요합니다.
- Docker Desktop 컨테이너를 빌드하고 실행하려면 필요합니다. 설치 후 실행해서 백그라운드에 켜두세요.
docker build -t sglang-jax .
Dockerfile을 기반으로 실행 가능한 이미지를 빌드합니다.
docker run -p 8080:80 sglang-jax
빌드된 이미지를 실제 컨테이너로 실행합니다.
터미널에 docker compose ps 를 입력해 컨테이너들이 Up 상태인지 확인하세요. README에 포트 번호가 적혀있다면 브라우저에서 http://localhost:포트번호 로 접속해보세요.
3. Python
쉬움사전 준비물
pip install .
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
python <실행할 파일명>.py # README에서 정확한 실행 파일명을 확인하세요
파이썬 스크립트(또는 모듈)를 실행합니다.
에러 메시지 없이 실행되고 터미널에 안내 문구가 출력되면 정상입니다.
4. Make
보통사전 준비물
- Git GitHub에서 프로젝트 코드를 내려받으려면 필요합니다.
- Make Linux/macOS는 보통 기본 설치되어 있습니다. Windows는 별도 설치(예: MSYS2, WSL)가 필요합니다.
make
생성된 빌드 설정을 바탕으로 실제 컴파일을 진행해 실행 파일을 만듭니다.
에러 없이 끝나면 성공입니다. 생성된 실행 파일을 직접 실행해보세요.
// repository documentation
Was this content helpful?
(0 ratings)
