This release adds 3 notable features for engineering teams evaluating rollout.
Published 1mo
Containers & Orchestration
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
autoscaling
edge-computing
gguf
gpu
+12 more
self-hosted
inference
kubernetes
llama-cpp
llm
local-llm
metal
mlx
multi-gpu
nvidia
tgi
vllm
Summary
AI summaryUpdates 0.8.9, Bug Fixes, and 2026-06-19 across a mixed release.
Full changelog
0.8.9 (2026-06-19)
Features
- add vulkan accelerator enum and make readiness-check Vulkan-aware (#735) (76cf370)
- cli: add --node-port to pin a stable NodePort on InferenceService (#737) (a6c1a03)
- cli: add
llmkube scalesubcommand to scale InferenceService replicas (#736) (bdb89f8) - controller: vendor-neutral DRA (resource.k8s.io/v1) scheduling for InferenceService (#750) (6eb0f27)
- foreman: deterministic coder verification gate with feedback loop (#749) (8cf3295)
- foreman: loop convergence forcing (EditFreeStreak + final-turns submit) (#741) (cd3f068)
Bug Fixes
- cli: correct --node-port help text and add NodePort test coverage (#742) (ef55900)
- controller: honor GPU resourceName override in checkAcceleratorAvailability (#747) (5aa3152)
- controller: honor RouterRule.Timeout in gateway-mode AIGatewayRoute generation (#748) (1978a72)
- foreman: make install-foreman-agent produce a working plist (#743) (b946fef)
- foreman: scope-overlap rail rescues honest paraphrase from false NO-GO (#746) (00ae36e)
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About LLMKube
Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.
Beta — feedback welcome: [email protected]