omniserve
[MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
File Explorer
Download Latest Version (.zip)- lserve-acc.png
- lserve-speed.png
- teaser.png
- abs_speed.png
- accuracy.png
- efficiency.png
- perplexity.png
- qoq.png
- teaser.png
- vlm_cap_example.png
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- config.json
- full_attention_heads.tsv
- dataset2maxlen.json
- dataset2prompt.json
- model2maxlen.json
- model2path.json
- attn_patterns
- eval.py
- metrics.py
- models
- pred.py
- pred_test.py
- prompt.txt
- utils.py
- addiction.txt
- aord.txt
- apple.txt
- avg.txt
- before.txt
- bias.txt
- boss.txt
- copy.txt
- corpdev.txt
- desres.txt
- diff.txt
- ecw.txt
- founders.txt
- foundervisa.txt
- gap.txt
- gba.txt
- gh.txt
- goodtaste.txt
- hubs.txt
- iflisp.txt
- island.txt
- know.txt
- langdes.txt
- laundry.txt
- love.txt
- mod.txt
- newideas.txt
- nft.txt
- philosophy.txt
- popular.txt
- pow.txt
- rootsoflisp.txt
- rss.txt
- siliconvalley.txt
- startuplessons.txt
- submarine.txt
- sun.txt
- superangels.txt
- todo.txt
- unions.txt
- useful.txt
- vb.txt
- vcsqueeze.txt
- vw.txt
- want.txt
- web20.txt
- weird.txt
- wisdom.txt
- worked.txt
- __init__.py
- attn_patterns
- models
- needle_in_haystack.py
- utils.py
- visualize.py
- longbench.sh
- submit_longbench.sh
- niah_test.sh
- submit_niah.sh
- cudaBf16Fallbacks.cuh
- cudaFp8Utils.h
- cudaTypeUtils.cuh
- decoderMaskedMultiheadAttentionUtils.h
- gptKernels.h
- input_metadata_helper.cu
- input_metadata_helper.h
- kvCacheUtils.h
- memoryUtils.h
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- fused_attention.cpp
- fused_attention.h
- applyBiasRopeUpdateKVCache.h
- update_kv_cache.cu
- update_kv_cache.h
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- fused_attention.cpp
- fused_attention.h
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- fused_attention.cpp
- fused_attention.h
- applyBiasRopeUpdateKVCache.h
- update_kv_cache.cu
- update_kv_cache.h
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- fused_attention.cpp
- fused_attention.h
- applyBiasRopeUpdateKVCache.h
- cudaBf16Fallbacks.cuh
- cudaFp8Utils.h
- cudaTypeUtils.cuh
- decoderMaskedMultiheadAttention.cu
- decoderMaskedMultiheadAttention.h
- decoderMaskedMultiheadAttentionTemplate.hpp
- decoderMaskedMultiheadAttentionUtils.h
- fused_attention.cpp
- fused_attention.h
- gptKernels.h
- input_metadata_helper.cu
- input_metadata_helper.h
- kvCacheUtils.h
- memoryUtils.h
- README.md
- update_kv_cache.cu
- update_kv_cache.h
- block_info.h
- context_pool_kernel.cu
- context_pool_kernel.h
- context_pool_utils.h
- pybind.cpp
- static_switch.h
- fused_kv_page_selector.cpp
- fused_kv_page_selector.h
- KVPageSelector.cu
- KVPageSelector.h
- KVPageSelectorTemplate.hpp
- gemm_cuda.cu
- gemm_cuda.h
- pybind.cpp
- gemm_cuda.cu
- gemm_cuda.h
- pybind.cpp
- pybind.cpp
- w8a8_gemm_cuda.cu
- w8a8_gemm_cuda.h
- activation.cpp
- activation_kernels.cu
- dispatch_utils.h
- fused.cpp
- fused_kernels.cu
- layernorm.cpp
- layernorm_kernels.cu
- reduction_utils.cuh
- utils.cuh
- setup.py
- __init__.py
- block_manager.py
- policy.py
- scheduler.py
- arg_utils.py
- llm_engine.py
- block_table_utils.py
- ctx_attn_func.py
- ctx_attn_init.py
- __init__.py
- w4a8_linear.py
- w4a8_moe_linear.py
- w8a8_linear.py
- __init__.py
- activation.py
- ctx_update_kv.py
- decoding_attention.py
- layernorm.py
- sampler.py
- llama_w16a16_unpad.py
- llama_w4a8_unpad.py
- llama_w8a8_unpad.py
- mixtral_w4a8_unpad.py
- transformers_utils.py
- __init__.py
- constants.py
- input_metadata.py
- quant_config.py
- tokenizer.py
- utils.py
- weight_utils.py
- __init__.py
- cache_engine.py
- model_runner.py
- worker.py
- __init__.py
- attn_config.py
- block.py
- config.py
- conversation.py
- logger.py
- prefix.py
- sampling_params.py
- sequence.py
- checkpoint_converter.py
- convert.sh
- quant_utils.py
- benchmark.sh
- launch.sh
- benchmark_a100.sh
- benchmark_l40s.sh
- lserve_e2e.sh
- qserve_e2e.sh
- test_vllm_full_net_sampling.py
- .gitignore
- LICENSE
- lserve_benchmark.py
- lserve_e2e_generation.py
- pyproject.toml
- qserve_benchmark.py
- qserve_e2e_generation.py
- README.md
// repository documentation
Was this content helpful?
(0 ratings)
