SentenceKV

(★ 15)

Official implementation of "SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching" (COLM 2025). A novel KV cache compression method that organizes cache at sentence level using semantic similarity.

SentenceKV 최신버젼 다운로드

최종 버전 다운로드 (.zip)
// repository documentation