lazy-moe

(★ 23)

The GPU-free LLM inference engine. Combines lazy expert loading + TurboQuant KV compression to run models that shouldn't fit on your hardware. Built from scratch, fully local, zero cloud.

lazy-moe Latest Version Download

Download Latest Version (.zip)
// repository documentation