LocalAI
Model Serving & MLOpsOpen‑source AI engine that lets you run any LLM, vision, voice, image or video model on CPU‑only or GPU hardware with a single API.
Features
- Composable core: backends (llama.cpp, vLLM, whisper.cpp, stable‑diffusion, etc.) are pulled on demand
- Drop‑in API compatibility with OpenAI, Anthropic and ElevenLabs endpoints
- Supports every modality – LLMs, vision, voice, image and video behind one API
- Runs on any hardware: NVIDIA/AMD/Intel GPUs, Apple Silicon, Vulkan or pure CPU
- Multi‑user security features (API key auth, quotas, role‑based access)
Recent releases
View all 40 releases →
Upgrade now
v4.6.0
Mixed
RCE / SSRF
ROCm speed, realtime warm-up, chat forking, load resilience, PII metrics,
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Install & Platforms
Install via
docker
Platforms
macos
linux
arm64