qwentin

(★ 11)

27B at 256k context on one RTX 5090 — FP6 + 4-bit KV + MTP spec-decode, hand-written SM120 tensor-core kernels, OpenAI API.

qwentin Latest Version Download

Download Latest Version (.zip)
// repository documentation