Skip to content

LLMKube

v0.9.9 Feature

This release adds 3 notable features for engineering teams evaluating rollout.

✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

ai apple-silicon autoscaling edge-computing gguf gpu
+12 more
self-hosted inference kubernetes llama-cpp llm local-llm metal mlx multi-gpu nvidia tgi vllm

Affected surfaces

deps

Summary

AI summary

Updates Bug Fixes, 0.9.9, and https://github.com/defilantech/LLMKube/issues/1196 across a mixed release.

Full changelog

0.9.9 (2026-07-22)

Features

  • api: vendor-neutral gpuSharing tiers on InferenceService (#1196 stories 1+2) (#1199) (6cbd1fd)
  • chart: gpuSharing pool + VRAM values wired to operator flags (#1196 story 3) (#1205) (95eed30)
  • controller: bound InferenceService pod lifetime via maxPodLifetimeSeconds (#1182) (3b39ee9)
  • foreman: make the verify gate optional in the issue-batch decomposition (#1166) (e4e49aa)
  • metrics: emit GPUQuota usage + admission-denial metrics (#416) (#1193) (dad2687)
  • quota: VRAM-based GPUQuota accounting from the gpuSharing tier (#1196 story 4) (#1200) (2c9000c)
  • runtime: Blackwell-ready image pins, NVIDIA llama.cpp divert, runtimeImages overrides, platform floors (#1197) (#1204) (9d40b36)
  • webhook: admission-time gpuSharing validation (#1196 story 5) (#1206) (a9de8f7)

Bug Fixes

  • agent: inject current date into issue-fix prompt and anchor research queries (#1202) (#1207) (5067095)
  • chart: nil-safe gpuSharing navigation in deployment args (#1209) (bbbcdcf)
  • codegen: regenerate stale committed codegen and widen the sync check (#1190) (aedcb80)
  • deps: bump golang.org/x/text to v0.39.0 (GO-2026-5970) (#1192) (af5c78e)
  • foreman: bound reasoning-model coder turns with a per-turn token budget (#1175) (5100c11)
  • foreman: count identical read_file re-reads as no-progress under the edit-free nudge (#1179) (2fad7ec)
  • foreman: inject Workload.spec.intent into coder prompts (#1201) (#1203) (9a8371c)
  • foreman: warn when an MCP allowlist entry matches no server tool (#1184) (0664814)
  • foreman: widen coder scope-overlap gate to a bottom-quartile floor (#1181) (f07cba4)
  • webhook: serve GPUQuota webhook at its configured path, with e2e coverage (#416) (#1195) (b377f52)

Documentation

  • grafana: AMD (Strix/gfx1151) GPU observability dashboard (#1188) (3da1f2b)
  • observability: multi-tenancy guide + GPUQuota Grafana dashboard (#1097) (#1194) (aedbd52)
  • operations: GPU sharing guide (#1196 story 6) (#1208) (c9007a3)

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track LLMKube

Get notified when new releases ship.

Sign up free

About LLMKube

Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.

All releases →

Related context

Earlier breaking changes

  • v0.8.1 foreman: requestTimeoutSeconds now sets loop-wide budget, default changes from 600 to 3600.

Beta — feedback welcome: [email protected]