llama.cpp-turboq-mtp
Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090
// repository documentation
Was this content helpful?
(0 ratings)