Skip to content

LocalAI

Model Serving & MLOps

Open‑source AI engine that lets you run any LLM, vision, voice, image or video model on CPU‑only or GPU hardware with a single API.

Go Latest v4.7.1 · 12d ago Security brief →

Features

  • Composable core: backends (llama.cpp, vLLM, whisper.cpp, stable‑diffusion, etc.) are pulled on demand
  • Drop‑in API compatibility with OpenAI, Anthropic and ElevenLabs endpoints
  • Supports every modality – LLMs, vision, voice, image and video behind one API
  • Runs on any hardware: NVIDIA/AMD/Intel GPUs, Apple Silicon, Vulkan or pure CPU
  • Multi‑user security features (API key auth, quotas, role‑based access)

Recent releases

View all 40 releases →
No immediate action
v4.7.1 Bug fix

Config injection fix

Review required
v4.7.0 Breaking risk
Auth

Voice cloning + video avatars + streaming TTS

No immediate action
v4.6.2 Mixed

Code refactor + CI fix + MiniCPM models

No immediate action
v4.6.1 New feature

Capabilities API + Agent metrics + Auth logs

Upgrade now
v4.6.0 Mixed
RCE / SSRF

ROCm speed, realtime warm-up, chat forking, load resilience, PII metrics,

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

About

Stars
47,661
Forks
4,257
Languages
Go JavaScript Python

Install & Platforms

Install via
docker
Platforms
macos linux arm64

Community & Support

Beta — feedback welcome: [email protected]