Single-binary GGUF model runtime in C — CPU/CUDA/Metal, OpenAI-compatible. Serves, scores, and trains LoRA directly through the quantized weights it deploys, with byte-reproducible adapters. Tool calls survive the token limit; sparse MoE and schema-constrained decoding included.