About
llama.cpp backend that runs GGUF models directly on the Axera AX8850 NPU — no model conversion, no per-model compile. 24-30 t/s decode on a Raspberry Pi 5 with the CPU idle.
Languages
Contributors1
No features listed.
Comments Theme
llama.cpp backend that runs GGUF models directly on the Axera AX8850 NPU — no model conversion, no per-model compile. 24-30 t/s decode on a Raspberry Pi 5 with the CPU idle.
No features listed.