This release adds 3 notable features for engineering teams evaluating rollout.
Published 23d
Containers & Orchestration
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
autoscaling
edge-computing
gguf
gpu
+12 more
self-hosted
inference
kubernetes
llama-cpp
llm
local-llm
metal
mlx
multi-gpu
nvidia
tgi
vllm
Summary
AI summaryUpdates Bug Fixes, 0.8.28, and 2026-07-04 across a mixed release.
Full changelog
0.8.28 (2026-07-04)
Features
- foreman: bounded fix iteration on reviewer NO-GO instead of terminal failure (#959) (d820fff)
- foreman: coder-tier escalation, re-dispatch a failed coder to a larger model (#964) (ce8f655)
- foreman: executor-owned revise-from-branch restore for revision tasks (#967) (b76051c)
- foreman: open the pull request on review GO (#956) (fd852e1)
- inference: add spec.modelCache.claimName for user-owned cache PVCs (#960) (aab5a58)
Bug Fixes
- foreman: accept workspace-internal absolute paths in resolveInside (#957) (34b126c)
- foreman: defer generic self-gate when runtime is missing from coder image (#958) (df185ec)
- foreman: reject no-op str_replace where old_string equals new_string (#969) (c71f38b)
- foreman: scope-overlap check catches Go files in new directories (#962) (486a944)
- inference: warn when modelCache.claimName is silently ignored (#966) (d49cd22)
Documentation
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About LLMKube
Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.
Beta — feedback welcome: [email protected]