bw24

bw24

From-scratch Rust+CUDA inference engine, bit-exact by construction — NVFP4, MoE, MTP speculative decoding, tuned against measured limits of one RTX 5090 Laptop (sm_120a).

291 351 HN points
MIT
last commit 2026-07-09
Website Source
Share:

About

from-scratch LLM inference for RTX 5090 (sm_120a) and H100 (sm_90a)

bw24 website preview

Languages

Contributors3

No features listed.

Comments Theme
slug: bw24