Skip to content

LLMKube

Containers & Orchestration

A Kubernetes operator that simplifies self‑hosted LLM inference across NVIDIA, Apple Silicon, and AMD hardware with a two‑line YAML configuration

Go Latest llmkube-0.9.12 · 7h ago Security brief →

Features

  • Runs LLMs on heterogeneous GPU hardware (NVIDIA, Apple Silicon, AMD) via declarative Kubernetes resources
  • Automates model download, caching, health‑checks, scaling and OpenAI‑compatible API exposure
  • Provides ModelRouter for policy‑aware routing between local models and external providers (Anthropic, OpenAI, etc.)

Recent releases

View all 199 releases →
No immediate action
foreman-0.9.12 Feature

Opt-in Foreman add-on

No immediate action
v0.9.12 Mixed

Bug fixes + new features

No immediate action
foreman-0.9.11 Feature

Foreman install sequence

No immediate action
v0.9.11 Mixed

Cache prep + federation + foreman matching

No immediate action
foreman-0.9.10 Feature

Agentic workload scheduling

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

About

Stars
174
Forks
25
Languages
Go Shell Makefile

Install & Platforms

Install via
brew helm
Platforms
linux macos arm64

Community & Support

Beta — feedback welcome: [email protected]