This release adds 3 notable features for engineering teams evaluating rollout.
✓ No known CVEs patched in this version
Topics
+12 more
Summary
AI summaryUpdates 0.9.5, Bug Fixes, and 2026-07-13 across a mixed release.
Changes in this release
| Type | Severity | Summary | CVE |
|---|---|---|---|
| Feature | Medium |
add --planner-token flag to CLI for gateway‑routed planner add --planner-token flag to CLI for gateway‑routed planner Source: llm_adapter@2026-07-13 Confidence: high |
— |
| Feature | Medium |
add extraVolumes/extraVolumeMounts passthrough on InferenceService add extraVolumes/extraVolumeMounts passthrough on InferenceService Source: llm_adapter@2026-07-13 Confidence: high |
— |
| Feature | Medium |
extend drain‑before‑roll idle checks to vLLM, TGI, SGLang, and multi‑replica services extend drain‑before‑roll idle checks to vLLM, TGI, SGLang, and multi‑replica services Source: llm_adapter@2026-07-13 Confidence: high |
— |
| Feature | Medium |
expose Foreman CRD status as Prometheus metrics via CRS expose Foreman CRD status as Prometheus metrics via CRS Source: llm_adapter@2026-07-13 Confidence: high |
— |
| Feature | Medium |
implement honest‑verdict slice 1 with claim evidence and work‑class policy in Foreman implement honest‑verdict slice 1 with claim evidence and work‑class policy in Foreman Source: llm_adapter@2026-07-13 Confidence: high |
— |
| Bugfix | Medium |
honor x-kubernetes-preserve-unknown-fields in CI sample validation honor x-kubernetes-preserve-unknown-fields in CI sample validation Source: llm_adapter@2026-07-13 Confidence: high |
— |
| Bugfix | Medium |
scope slicer branch names by run ID in CLI scope slicer branch names by run ID in CLI Source: llm_adapter@2026-07-13 Confidence: high |
— |
| Bugfix | Medium |
default Model accelerator from gpu.runtime setting default Model accelerator from gpu.runtime setting Source: llm_adapter@2026-07-13 Confidence: high |
— |
| Bugfix | Medium |
reject trailing‑underscore pins at plan validation in slicer reject trailing‑underscore pins at plan validation in slicer Source: llm_adapter@2026-07-13 Confidence: high |
— |
| Bugfix | Medium |
surface pinned‑prefix drift in slicer output surface pinned‑prefix drift in slicer output Source: llm_adapter@2026-07-13 Confidence: high |
— |
Full changelog
0.9.5 (2026-07-13)
Features
- ci: validate config/samples against CRD schemas (Fixes #1021) (#1083) (b3ae847)
- cli: add --planner-token for gateway-routed planner (Fixes #1053) (#1090) (f71d2cc)
- controller: add extraVolumes/extraVolumeMounts passthrough on InferenceService (#1079) (80e004c)
- controller: extend drain-before-roll idle checks to vLLM, TGI, SGLang, and multi-replica services (#1088) (1d884df)
- foreman: expose CRD status as Prometheus metrics via CRS (Fixes #1001) (#1086) (a459748)
- foreman: honest-verdict slice 1 (claim evidence + work-class policy) (#1078) (6f2c216)
Bug Fixes
- ci: honor x-kubernetes-preserve-unknown-fields in validate-samples (Fixes #1085) (#1091) (5447e88)
- cli: scope slicer branch names by run id (Fixes #1054) (#1081) (9288401)
- controller: default Model accelerator from gpu.runtime (Fixes #1074) (#1087) (0390c0f)
- slicer: reject trailing-underscore pins at plan validation (Fixes #1058) (#1082) (d4118d3)
- slicer: surface pinned-prefix drift (Fixes #1084) (#1089) (0f6edd2)
Documentation
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About LLMKube
Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.
Beta — feedback welcome: [email protected]