Skip to content
vllm
Model Serving & MLOps
A high-throughput and memory-efficient inference and serving engine for LLMs
Python
·
Latest v0.26.0 · 1d ago
Security brief →
Features
-
State-of-the-art LLM inference throughput with PagedAttention memory management
-
Support for diverse quantization formats (FP8, MXFP4, INT4, GPTQ, etc.) and optimized kernels
-
Flexible serving options: OpenAI‑compatible API, Anthropic Messages API, gRPC, streaming outputs
No immediate action
v0.26.0
Breaking risk
·
Inkling models, DeepSeek‑V4 boost, KV offloading, security fixes, Rust multimodal
No immediate action
v0.25.1
Bug fix
·
FFmpeg import fix & RMSNorm dtype guard
Review required
v0.25.0
Breaking risk
·
Auth
RBAC
Dependencies
MRv2 default, PagedAttention removal, new models, parsers, performance
Review required
v0.24.0
Breaking risk
·
Auth
Breaking upgrade
MiniMax‑M3, DeepSeek‑V4, MRv2, Streaming Parser, DiffusionGemma
No immediate action
v0.23.0
Breaking risk
·
DeepSeek‑V4, MRv2, Rust streaming, KV offload, Unified parser
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
About
Languages
Python
·
Rust
·
Cuda
View on GitHub
Homepage
Documentation
Search tools, categories, lists, and users
Use ↑↓ to navigate, Enter to open, Esc to close
No results for ""
⌘K to open
↑↓ navigate
⏎ open