This release adds 3 notable features for engineering teams evaluating rollout.
Published 8h
Containers & Orchestration
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
autoscaling
edge-computing
gguf
gpu
+12 more
self-hosted
inference
kubernetes
llama-cpp
llm
local-llm
metal
mlx
multi-gpu
nvidia
tgi
vllm
Summary
AI summaryUpdates Bug Fixes, 0.9.12, and 2026-07-26 across a mixed release.
Full changelog
0.9.12 (2026-07-26)
Features
- add --prompt-depth-sweep to benchmark command (#1280) (eb511f9)
- foreman: support chat_template_kwargs on agent chat requests (#1285) (1028335)
- grafana: add SGLang and vLLM runtime dashboards (#1266) (07f340a)
- grafana: union llama.cpp, sglang and vLLM metrics on the inference dashboard (#1248) (7cb226b)
- hack: spine probes for models given repo access (#1274) (b7b3d9b)
Bug Fixes
- Available condition reflects Stopped and Suspended phases (#1273) (208bd79)
- disable llama.cpp prompt cache in benchmark requests (#1272) (eaa61ad)
- foreman: apply the empty-assistant guard to the preserved truncated turn (#1282) (b1f51cc)
- foreman: hold the session compaction drop point steady across turns (#1287) (febcb2a)
- foreman: three first-run failures on a clean cluster (#1295) (5ca3e64)
- gate: check non-Go changes against the shared merge-base, not HEAD~1 (#1284) (6940033)
- spine-probes: restore --max-tokens and the 'none' ablation arm (#1278) (9493201)
Documentation
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About LLMKube
Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.
Beta — feedback welcome: [email protected]