vllm

vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

90k 21k107 HN points
Apache-2.0
last commit 2026-08-30
Website Source
Share:

About

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm website preview

Languages

Contributors30

No features listed.

Comments Theme
slug: vllm