Skip to content

LLMKube

v0.9.5 Feature

This release adds 3 notable features for engineering teams evaluating rollout.

✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

ai apple-silicon autoscaling edge-computing gguf gpu
+12 more
self-hosted inference kubernetes llama-cpp llm local-llm metal mlx multi-gpu nvidia tgi vllm

Summary

AI summary

Updates 0.9.5, Bug Fixes, and 2026-07-13 across a mixed release.

Changes in this release

Feature Medium

add --planner-token flag to CLI for gateway‑routed planner

add --planner-token flag to CLI for gateway‑routed planner

Source: llm_adapter@2026-07-13

Confidence: high

Feature Medium

add extraVolumes/extraVolumeMounts passthrough on InferenceService

add extraVolumes/extraVolumeMounts passthrough on InferenceService

Source: llm_adapter@2026-07-13

Confidence: high

Feature Medium

extend drain‑before‑roll idle checks to vLLM, TGI, SGLang, and multi‑replica services

extend drain‑before‑roll idle checks to vLLM, TGI, SGLang, and multi‑replica services

Source: llm_adapter@2026-07-13

Confidence: high

Feature Medium

expose Foreman CRD status as Prometheus metrics via CRS

expose Foreman CRD status as Prometheus metrics via CRS

Source: llm_adapter@2026-07-13

Confidence: high

Feature Medium

implement honest‑verdict slice 1 with claim evidence and work‑class policy in Foreman

implement honest‑verdict slice 1 with claim evidence and work‑class policy in Foreman

Source: llm_adapter@2026-07-13

Confidence: high

Bugfix Medium

honor x-kubernetes-preserve-unknown-fields in CI sample validation

honor x-kubernetes-preserve-unknown-fields in CI sample validation

Source: llm_adapter@2026-07-13

Confidence: high

Bugfix Medium

scope slicer branch names by run ID in CLI

scope slicer branch names by run ID in CLI

Source: llm_adapter@2026-07-13

Confidence: high

Bugfix Medium

default Model accelerator from gpu.runtime setting

default Model accelerator from gpu.runtime setting

Source: llm_adapter@2026-07-13

Confidence: high

Bugfix Medium

reject trailing‑underscore pins at plan validation in slicer

reject trailing‑underscore pins at plan validation in slicer

Source: llm_adapter@2026-07-13

Confidence: high

Bugfix Medium

surface pinned‑prefix drift in slicer output

surface pinned‑prefix drift in slicer output

Source: llm_adapter@2026-07-13

Confidence: high

Full changelog

0.9.5 (2026-07-13)

Features

  • ci: validate config/samples against CRD schemas (Fixes #1021) (#1083) (b3ae847)
  • cli: add --planner-token for gateway-routed planner (Fixes #1053) (#1090) (f71d2cc)
  • controller: add extraVolumes/extraVolumeMounts passthrough on InferenceService (#1079) (80e004c)
  • controller: extend drain-before-roll idle checks to vLLM, TGI, SGLang, and multi-replica services (#1088) (1d884df)
  • foreman: expose CRD status as Prometheus metrics via CRS (Fixes #1001) (#1086) (a459748)
  • foreman: honest-verdict slice 1 (claim evidence + work-class policy) (#1078) (6f2c216)

Bug Fixes

Documentation

  • proposals: honest-verdict harness design (declare-then-verify coder gates) (#1076) (630de33)

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track LLMKube

Get notified when new releases ship.

Sign up free

About LLMKube

Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.

All releases →

Related context

Earlier breaking changes

  • v0.8.1 foreman: requestTimeoutSeconds now sets loop-wide budget, default changes from 600 to 3600.

Beta — feedback welcome: [email protected]