
hotpin-llm
HotPin: routing-guided expert pinning for lossless low-RAM MoE inference on consumer CPU+NVMe - 120B model at 3.84 tok/s in 19 GB RAM, no GPU
About
HotPin: routing-guided expert pinning for lossless low-RAM MoE inference on consumer CPU+NVMe - 120B model at 3.84 tok/s in 19 GB RAM, no GPU
Languages
Contributors1
No features listed.
Comments Theme




