Skip to content

LLMKube

v0.9.2 Feature

This release adds 3 notable features for engineering teams evaluating rollout.

✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

ai apple-silicon autoscaling edge-computing gguf gpu
+12 more
self-hosted inference kubernetes llama-cpp llm local-llm metal mlx multi-gpu nvidia tgi vllm

Summary

AI summary

Updates 0.9.2, Bug Fixes, and 2026-07-09 across a mixed release.

Changes in this release

Feature Medium

Adds coder grounding rail to flag doc‑contradicting metric writes in foreman.

Adds coder grounding rail to flag doc‑contradicting metric writes in foreman.

Source: llm_adapter@2026-07-16

Confidence: high

Feature Medium

Adds distinct ALREADY-RESOLVED coder outcome handling in foreman.

Adds distinct ALREADY-RESOLVED coder outcome handling in foreman.

Source: llm_adapter@2026-07-16

Confidence: high

Feature Medium

Adds flag for GO changes that modify no functional production code in foreman.

Adds flag for GO changes that modify no functional production code in foreman.

Source: llm_adapter@2026-07-16

Confidence: high

Feature Medium

Adds MCP client for agents (Phase 1: HTTP transport, context7) in foreman.

Adds MCP client for agents (Phase 1: HTTP transport, context7) in foreman.

Source: llm_adapter@2026-07-16

Confidence: high

Feature Medium

Adds semantic issueAsk verification instead of verbatim substring matching in foreman.

Adds semantic issueAsk verification instead of verbatim substring matching in foreman.

Source: llm_adapter@2026-07-16

Confidence: high

Feature Medium

Adds escalation from str_replace to write_file after repeated failures in foreman.

Adds escalation from str_replace to write_file after repeated failures in foreman.

Source: llm_adapter@2026-07-16

Confidence: high

Bugfix Medium

Exempts REJECT from grounded‑finding demotion in foreman.

Exempts REJECT from grounded‑finding demotion in foreman.

Source: llm_adapter@2026-07-16

Confidence: high

Bugfix Medium

Gates Job downloads of the go.mod toolchain using GOTOOLCHAIN=auto in foreman.

Gates Job downloads of the go.mod toolchain using GOTOOLCHAIN=auto in foreman.

Source: llm_adapter@2026-07-16

Confidence: high

Bugfix Medium

Guards file‑less findings and resolves bare/absolute paths in foreman’s grounding rail.

Guards file‑less findings and resolves bare/absolute paths in foreman’s grounding rail.

Source: llm_adapter@2026-07-16

Confidence: high

Bugfix Medium

Reaps audit ConfigMaps older than the retention period in foreman.

Reaps audit ConfigMaps older than the retention period in foreman.

Source: llm_adapter@2026-07-16

Confidence: high

Refactor Low

Removes dead metrics and a phantom PrometheusRule error‑rate rule in metrics.

Removes dead metrics and a phantom PrometheusRule error‑rate rule in metrics.

Source: granite4.1:30b@2026-07-16-audit

Confidence: low

Full changelog

0.9.2 (2026-07-09)

Features

  • foreman: coder grounding rail (flag doc-contradicting metric writes) (#1017) (aa969ed)
  • foreman: distinct ALREADY-RESOLVED coder outcome (#970) (#1011) (8e041e0)
  • foreman: flag a GO that changes no functional production code (#1024) (3e1a6fb)
  • foreman: MCP client for agents (Phase 1: HTTP transport, context7) (#1014) (d34343f)
  • foreman: semantic issueAsk verification instead of verbatim substring (#1020) (85d92ee)
  • foreman: str_replace escalates to write_file after repeated failures (#1026) (de1b8cb)

Bug Fixes

  • foreman: exempt REJECT from grounded-finding demotion (#1008) (18098cb)
  • foreman: gate Job downloads the go.mod toolchain (GOTOOLCHAIN=auto) (#1019) (5f4c51c)
  • foreman: guard file-less findings + resolve bare/absolute paths in grounding rail (#1009) (304a2b8)
  • foreman: reap audit ConfigMaps older than retention (#990) (#1027) (08d06d6)
  • foreman: reviewer diffs against the upstream base, not stale local fork main (#1006) (6ef7278)
  • metrics: remove dead metrics + phantom PrometheusRule error-rate rule (#1010) (7c92f3f)

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track LLMKube

Get notified when new releases ship.

Sign up free

About LLMKube

Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.

All releases →

Related context

Earlier breaking changes

  • v0.8.1 foreman: requestTimeoutSeconds now sets loop-wide budget, default changes from 600 to 3600.

Beta — feedback welcome: [email protected]