rai

rai

CPU-only LLM inference engine in pure Rust — 4-bit quantized models, hand-written AVX2 kernels, speculative decoding, and a local HTTP/MCP server. No GPU, no Python runtime.

10 0
Apache-2.0
last commit 2026-09-19
Website Source
Share:

About

CPU-only LLM inference engine in pure Rust — 4-bit quantized models, hand-written AVX2 kernels, speculative decoding, and a local HTTP/MCP server. No GPU, no Python runtime.

rai website preview

Languages

Contributors1

No features listed.

Comments Theme
slug: rai