omniserve
[MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
파일 탐색기
최종 버전 다운로드 (.zip)- lserve-acc.png
- lserve-speed.png
- teaser.png
- abs_speed.png
- accuracy.png
- efficiency.png
- perplexity.png
- qoq.png
- teaser.png
- vlm_cap_example.png
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- dataset2maxlen.json
- dataset2prompt.json
- model2maxlen.json
- model2path.json
- attn_patterns
- eval.py
- metrics.py
- models
- pred.py
- pred_test.py
- prompt.txt
- utils.py
- addiction.txt
- aord.txt
- apple.txt
- avg.txt
- before.txt
- bias.txt
- boss.txt
- copy.txt
- corpdev.txt
- desres.txt
- diff.txt
- ecw.txt
- founders.txt
- foundervisa.txt
- gap.txt
- gba.txt
- gh.txt
- goodtaste.txt
- hubs.txt
- iflisp.txt
- island.txt
- know.txt
- langdes.txt
- laundry.txt
- love.txt
- mod.txt
- newideas.txt
- nft.txt
- philosophy.txt
- popular.txt
- pow.txt
- rootsoflisp.txt
- rss.txt
- siliconvalley.txt
- startuplessons.txt
- submarine.txt
- sun.txt
- superangels.txt
- todo.txt
- unions.txt
- useful.txt
- vb.txt
- vcsqueeze.txt
- vw.txt
- want.txt
- web20.txt
- weird.txt
- wisdom.txt
- worked.txt
- __init__.py
- attn_patterns
- models
- needle_in_haystack.py
- utils.py
- visualize.py
- longbench.sh
- submit_longbench.sh
- niah_test.sh
- submit_niah.sh
- cudaBf16Fallbacks.cuh
- cudaFp8Utils.h
- cudaTypeUtils.cuh
- decoderMaskedMultiheadAttentionUtils.h
- gptKernels.h
- input_metadata_helper.cu
- input_metadata_helper.h
- kvCacheUtils.h
- memoryUtils.h
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- fused_attention.cpp
- fused_attention.h
- applyBiasRopeUpdateKVCache.h
- update_kv_cache.cu
- update_kv_cache.h
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- fused_attention.cpp
- fused_attention.h
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- fused_attention.cpp
- fused_attention.h
- applyBiasRopeUpdateKVCache.h
- update_kv_cache.cu
- update_kv_cache.h
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- fused_attention.cpp
- fused_attention.h
- applyBiasRopeUpdateKVCache.h
- cudaBf16Fallbacks.cuh
- cudaFp8Utils.h
- cudaTypeUtils.cuh
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- decoderMaskedMultiheadAttentionUtils.h
- fused_attention.cpp
- fused_attention.h
- gptKernels.h
- input_metadata_helper.cu
- input_metadata_helper.h
- kvCacheUtils.h
- memoryUtils.h
- README.md
- update_kv_cache.cu
- update_kv_cache.h
- block_info.h
- context_pool_kernel.cu
- context_pool_kernel.h
- context_pool_utils.h
- pybind.cpp
- static_switch.h
- fused_kv_page_selector.cpp
- fused_kv_page_selector.h
- KVPageSelector.cu
- KVPageSelector.h
- KVPageSelectorTemplate.hpp
- gemm_cuda.cu
- gemm_cuda.h
- pybind.cpp
- gemm_cuda.cu
- gemm_cuda.h
- pybind.cpp
- pybind.cpp
- w8a8_gemm_cuda.cu
- w8a8_gemm_cuda.h
- activation.cpp
- activation_kernels.cu
- dispatch_utils.h
- fused.cpp
- fused_kernels.cu
- layernorm.cpp
- layernorm_kernels.cu
- reduction_utils.cuh
- utils.cuh
- setup.py
- __init__.py
- block_manager.py
- policy.py
- scheduler.py
- arg_utils.py
- llm_engine.py
- block_table_utils.py
- ctx_attn_func.py
- ctx_attn_init.py
- __init__.py
- w4a8_linear.py
- w4a8_moe_linear.py
- w8a8_linear.py
- __init__.py
- activation.py
- ctx_update_kv.py
- decoding_attention.py
- layernorm.py
- sampler.py
- llama_w16a16_unpad.py
- llama_w4a8_unpad.py
- llama_w8a8_unpad.py
- mixtral_w4a8_unpad.py
- transformers_utils.py
- __init__.py
- constants.py
- input_metadata.py
- quant_config.py
- tokenizer.py
- utils.py
- weight_utils.py
- __init__.py
- cache_engine.py
- model_runner.py
- worker.py
- __init__.py
- attn_config.py
- block.py
- config.py
- conversation.py
- logger.py
- prefix.py
- sampling_params.py
- sequence.py
- checkpoint_converter.py
- convert.sh
- quant_utils.py
- benchmark.sh
- launch.sh
- benchmark_a100.sh
- benchmark_l40s.sh
- lserve_e2e.sh
- qserve_e2e.sh
- test_vllm_full_net_sampling.py
- .gitignore
- LICENSE
- lserve_benchmark.py
- lserve_e2e_generation.py
- pyproject.toml
- qserve_benchmark.py
- qserve_e2e_generation.py
- README.md
# 설치 가이드
1. 코드 내려받기
git clone https://github.com/mit-han-lab/omniserve
깃허브에서 프로젝트 코드 전체를 내 컴퓨터로 내려받습니다.
cd omniserve
방금 내려받은 프로젝트 폴더 안으로 이동합니다.
2. 공식 설치 스크립트
쉬움 추천사전 준비물
- Python 3 pip 명령어를 쓰려면 Python이 필요합니다.
pip install --upgrade pip # enable PEP 660 support
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install flash-attn --no-build-isolation
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install packaging
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install ninja
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
설치 후 새 터미널을 열고, 프로그램의 버전 확인 명령(예: --version)으로 정상 설치됐는지 확인하세요.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
3. Python
쉬움사전 준비물
pip install --upgrade pip # enable PEP 660 support
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install -e .
requirements.txt 등에 명시된 파이썬 라이브러리를 설치합니다.
pip install flash-attn --no-build-isolation
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install packaging
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install ninja
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
에러 메시지 없이 실행되고 터미널에 안내 문구가 출력되면 정상입니다.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
// repository documentation
Was this content helpful?
(0 ratings)
