This release adds 3 notable features for engineering teams evaluating rollout.
Published 20d
Containers & Orchestration
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
autoscaling
edge-computing
gguf
gpu
+12 more
self-hosted
inference
kubernetes
llama-cpp
llm
local-llm
metal
mlx
multi-gpu
nvidia
tgi
vllm
Affected surfaces
breaking_upgrade
Summary
AI summaryUpdates Bug Fixes, 0.9.1, and 2026-07-07 across a mixed release.
Full changelog
0.9.1 (2026-07-07)
Features
- chart: gate CRD installation behind crds.enabled toggle (#998) (acf2e6a)
- foreman: append actionable steers to structural-lint gate feedback (#984) (4d496c6)
- foreman: closest-line fallback when str_replace has no unique anchor (#1000) (02ec677)
- foreman: grounded-finding rail for reviewer NO-GO verdicts (#988) (cadbd30)
- foreman: verdict-from-findings rail (promote GO with a grounded blocker) (#992) (b82fcca)
Bug Fixes
- cachekey: unify model cache-key derivation across controller and CLI (#985) (1fb122c)
- foreman: grant coder Job pods get on secrets for cloud-proxy auth (#987) (ffe45d7)
- foreman: key test-presence gate on net-new functions only (#983) (0236630)
- foreman: tolerate self-committed model work and detect git apply edits (#995) (730d1cd)
- foreman: truncate over-length submit_result summaries instead of rejecting (#999) (8c36edd)
Documentation
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About LLMKube
Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.
Beta — feedback welcome: [email protected]