This release adds 12 notable features for engineering teams evaluating rollout.
Published 1mo
Containers & Orchestration
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
autoscaling
edge-computing
gguf
gpu
+12 more
self-hosted
inference
kubernetes
llama-cpp
llm
local-llm
metal
mlx
multi-gpu
nvidia
tgi
vllm
Affected surfaces
auth
rbac
breaking_upgrade
Summary
AI summaryBroad release touches 0.8.8, Bug Fixes, 2026-06-17, and hardware.gpu.runtime.
Full changelog
0.8.8 (2026-06-17)
Features
- AMD/Vulkan runtime image selection (hardware.gpu.runtime) (#727) (1a4544f)
- crd: make GPU resource name configurable to support AMD/Vulkan/Intel scheduling (#709) (c88becf)
- gateway: active HTTP health checks on the ModelRouter BTP for fast backend ejection (#662) (#704) (ba99060)
- gateway: event-driven route-level ejection of unhealthy backends (#662) (#706) (815f2bf)
- gateway: gateway-scoped audit access log + fail-loud auditLog in Gateway mode (2c) (#703) (b874b5e)
- gateway: header-only data-classification routing + fail-closed sensitive guard (2e-core) (#707) (0249665)
- gateway: InferenceService Envoy AI Gateway exposure (MVP) (#692) (3b095dc)
- gateway: ModelRouter dataPlane Gateway mode with cross-tier failover (2a) (#693) (2842634)
- gateway: ModelRouter JWT authentication via SecurityPolicy (2d-core) (#695) (73a2ea9)
- gateway: ModelRouter per-team model allowlists via SecurityPolicy authorization (2d.2) (#702) (94428b4)
- gateway: ModelRouter token budgets and 429 enforcement (2b) (#694) (627e85a)
- metal-agent: withdraw endpoint when runtime is unhealthy (#662) (#705) (5ed9395)
- selfupdate: bound download size + GC old agent versions (#690) (5205a62)
- webhook: ModelRouter validating webhook for apply-time honest-boundary rejection (#708) (13d9321)
Bug Fixes
- cache: restore shared model cache as the default (perService becomes opt-in) (#732) (44ab7dc)
- per-node model cache so GPU on a second node can schedule (#728) (#729) (79bccce)
Documentation
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About LLMKube
Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.
Beta — feedback welcome: [email protected]