reame

(★ 107)

CPU-first LLM inference server on llama.cpp. Runs useful models on free-tier ARM boxes; rewriting the input made it ~6x faster and more accurate than tuning the engine. MIT, benchmarks and failures included.

reame 최신버젼 다운로드

최종 버전 다운로드 (.zip)
// repository documentation