quant.cpp
LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.
// repository documentation
Was this content helpful?
(0 ratings)