reame

reame

Reame — CPU-first LLM inference server on llama.cpp: disk KV cache, self-regulating speculation, generation archive, interleaved multi-user, the Conclave. Your hardware, your realm.

102 959 HN points
MIT
last commit 2026-07-12
Source
Share:

About

CPU-first LLM inference server on llama.cpp. Runs useful models on free-tier ARM boxes; rewriting the input made it ~6x faster and more accurate than tuning the engine. MIT, benchmarks and failures included.

Languages

Contributors1

No features listed.

Comments Theme
slug: reame