Skip to content

vllm

Model Serving & MLOps

A high-throughput and memory-efficient inference and serving engine for LLMs

Python Latest v0.26.0 · 1d ago Security brief →

Features

  • State-of-the-art LLM inference throughput with PagedAttention memory management
  • Support for diverse quantization formats (FP8, MXFP4, INT4, GPTQ, etc.) and optimized kernels
  • Flexible serving options: OpenAI‑compatible API, Anthropic Messages API, gRPC, streaming outputs

Recent releases

View all 22 releases →
No immediate action
v0.26.0 Breaking risk

Inkling models, DeepSeek‑V4 boost, KV offloading, security fixes, Rust multimodal

No immediate action
v0.25.1 Bug fix

FFmpeg import fix & RMSNorm dtype guard

Review required
v0.25.0 Breaking risk
Auth RBAC Dependencies

MRv2 default, PagedAttention removal, new models, parsers, performance

Review required
v0.24.0 Breaking risk
Auth Breaking upgrade

MiniMax‑M3, DeepSeek‑V4, MRv2, Streaming Parser, DiffusionGemma

No immediate action
v0.23.0 Breaking risk

DeepSeek‑V4, MRv2, Rust streaming, KV offload, Unified parser

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

About

Stars
86,572
Forks
19,567
Languages
Python Rust Cuda

Install & Platforms

Install via
pip

Community & Support

Beta — feedback welcome: [email protected]