
rai
CPU-only LLM inference engine in pure Rust — 4-bit quantized models, hand-written AVX2 kernels, speculative decoding, and a local HTTP/MCP server. No GPU, no Python runtime.
About
Languages
Contributors1
No features listed.
Comments Theme
Platforms
Hosting
Self-hosted
Install
Docker
Cargo





