llama.cpp-turboq-mtp

(★ 90)

Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090

llama.cpp-turboq-mtp 최신버젼 다운로드

최종 버전 다운로드 (.zip)
// repository documentation