vLLM
Runtimes · Self-hosted
FreeOpen sourceGood for
- Serve language models to many simultaneous users from shared GPU hardware.
- Use paged attention, batching, caching, and quantization to manage inference.
- Expose OpenAI-compatible, Anthropic, or gRPC interfaces for applications.
Watch out
- It is GPU-first and operationally oriented toward concurrent service workloads.