omniserve

(★ 854)

[MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention

  • .gitignore
  • LICENSE
  • lserve_benchmark.py
  • lserve_e2e_generation.py
  • pyproject.toml
  • qserve_benchmark.py
  • qserve_e2e_generation.py
  • README.md
// repository documentation