This release adds 3 notable features for engineering teams evaluating rollout.
Published 4d
Containers & Orchestration
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
autoscaling
edge-computing
gguf
gpu
+12 more
self-hosted
inference
kubernetes
llama-cpp
llm
local-llm
metal
mlx
multi-gpu
nvidia
tgi
vllm
Affected surfaces
deps
Summary
AI summaryUpdates Bug Fixes, 0.9.9, and https://github.com/defilantech/LLMKube/issues/1196 across a mixed release.
Full changelog
0.9.9 (2026-07-22)
Features
- api: vendor-neutral gpuSharing tiers on InferenceService (#1196 stories 1+2) (#1199) (6cbd1fd)
- chart: gpuSharing pool + VRAM values wired to operator flags (#1196 story 3) (#1205) (95eed30)
- controller: bound InferenceService pod lifetime via maxPodLifetimeSeconds (#1182) (3b39ee9)
- foreman: make the verify gate optional in the issue-batch decomposition (#1166) (e4e49aa)
- metrics: emit GPUQuota usage + admission-denial metrics (#416) (#1193) (dad2687)
- quota: VRAM-based GPUQuota accounting from the gpuSharing tier (#1196 story 4) (#1200) (2c9000c)
- runtime: Blackwell-ready image pins, NVIDIA llama.cpp divert, runtimeImages overrides, platform floors (#1197) (#1204) (9d40b36)
- webhook: admission-time gpuSharing validation (#1196 story 5) (#1206) (a9de8f7)
Bug Fixes
- agent: inject current date into issue-fix prompt and anchor research queries (#1202) (#1207) (5067095)
- chart: nil-safe gpuSharing navigation in deployment args (#1209) (bbbcdcf)
- codegen: regenerate stale committed codegen and widen the sync check (#1190) (aedcb80)
- deps: bump golang.org/x/text to v0.39.0 (GO-2026-5970) (#1192) (af5c78e)
- foreman: bound reasoning-model coder turns with a per-turn token budget (#1175) (5100c11)
- foreman: count identical read_file re-reads as no-progress under the edit-free nudge (#1179) (2fad7ec)
- foreman: inject Workload.spec.intent into coder prompts (#1201) (#1203) (9a8371c)
- foreman: warn when an MCP allowlist entry matches no server tool (#1184) (0664814)
- foreman: widen coder scope-overlap gate to a bottom-quartile floor (#1181) (f07cba4)
- webhook: serve GPUQuota webhook at its configured path, with e2e coverage (#416) (#1195) (b377f52)
Documentation
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About LLMKube
Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.
Beta — feedback welcome: [email protected]