This release adds 3 notable features for engineering teams evaluating rollout.
Published 2d
Containers & Orchestration
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
autoscaling
edge-computing
gguf
gpu
+12 more
self-hosted
inference
kubernetes
llama-cpp
llm
local-llm
metal
mlx
multi-gpu
nvidia
tgi
vllm
Summary
AI summaryUpdates 0.9.11, Bug Fixes, and 2026-07-24 across a mixed release.
Full changelog
0.9.11 (2026-07-24)
Features
- cache prep subcommand with direct chown/chmod syscalls (#890 item 3) (#1258) (bde1a87)
- charts/foreman: expose the agent --accelerator flag (#1242) (#1261) (a39fd71)
- declarative bindAddress on InferenceService (#1240) (#1247) (bf01598)
- federation: FederatedCluster registry + fleet status rollup (#1234, #1235) (#1241) (1b5bf08)
- foreman: tiered str_replace matching (exact, trailing-ws, uniform-indent) (#1231) (0372e24)
- grafana: chart the InferenceService metrics with no panel (#1245) (a5ce923)
- Kueue prerequisites: spec.suspend, GPUQuota deferral, pkg/apiutil GPU mapping (#1254) (c95e06f)
- metrics: publish operator state metrics from observed state (#1229) (5db0915)
- prefetch Model artifacts into the shared cache without an InferenceService (#1218) (7c1161b)
Bug Fixes
- chart: make PrometheusRule select the chart's own scrape targets (#1239) (ffc697a)
- delete GPUQuota metric series when the CR is deleted (#1230) (#1246) (1f58c3c)
- foreman: edits disarm a nudged RepeatedToolCall hash (#1215) (#1216) (60df2ad)
- gate bite check reverts to the upstream merge-base, never the fork tip (#1260) (8a6928f)
- grafana: delete dashboard panels for metrics nothing emits (#1244) (9c586d9)
- metrics: report GPU queue depth per namespace (#1243) (5779c05)
- recycle inference pods without invalid deadlines (#1232) (0b716d7)
- waitForIdle proceeds past crashlooping pods; truthful deferral reason on mixed state (#1250) (#1262) (88c87f9)
Documentation
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About LLMKube
Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.
Beta — feedback welcome: [email protected]