Skip to content

LLMKube

v0.9.6 Feature

This release adds 3 notable features for engineering teams evaluating rollout.

✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

ai apple-silicon autoscaling edge-computing gguf gpu
+12 more
self-hosted inference kubernetes llama-cpp llm local-llm metal mlx multi-gpu nvidia tgi vllm

Summary

AI summary

Updates 0.9.6, Bug Fixes, and 2026-07-14 across a mixed release.

Changes in this release

Feature Medium

Adds GPUQuota CRD types for multi-tenant GPU governance.

Adds GPUQuota CRD types for multi-tenant GPU governance.

Source: llm_adapter@2026-07-15

Confidence: high

Feature Medium

Adds GPUQuota status reconciler in controller.

Adds GPUQuota status reconciler in controller.

Source: llm_adapter@2026-07-15

Confidence: high

Feature Medium

Adds GPUQuota validating webhook for InferenceService.

Adds GPUQuota validating webhook for InferenceService.

Source: llm_adapter@2026-07-15

Confidence: high

Feature Medium

Adds s3:// model source support via curl --aws-sigv4.

Adds s3:// model source support via curl --aws-sigv4.

Source: llm_adapter@2026-07-15

Confidence: high

Feature Medium

Adds GPUQuota admission decision function.

Adds GPUQuota admission decision function.

Source: llm_adapter@2026-07-15

Confidence: high

Feature Low

Closes out SGLang kitchen‑sink: adds minor flags, accept thresholds, typed LoRA adapters, and LoRAAdapter CRD.

Closes out SGLang kitchen‑sink: adds minor flags, accept thresholds, typed LoRA adapters, and LoRAAdapter CRD.

Source: llm_adapter@2026-07-15

Confidence: high

Feature Low

Gates InferenceService quota webhook and tenant RBAC behind multitenancy toggle in Helm charts.

Gates InferenceService quota webhook and tenant RBAC behind multitenancy toggle in Helm charts.

Source: llm_adapter@2026-07-15

Confidence: high

Feature Low

Preserves a coder's gate‑failed branch instead of discarding it in foreman.

Preserves a coder's gate‑failed branch instead of discarding it in foreman.

Source: llm_adapter@2026-07-15

Confidence: high

Bugfix Medium

Fixes foreman to honor GateProfile source extensions in scope‑overlap issue‑ref extraction.

Fixes foreman to honor GateProfile source extensions in scope‑overlap issue‑ref extraction.

Source: llm_adapter@2026-07-15

Confidence: high

Other Low

Adds air‑gapped local file‑path model source example in samples.

Adds air‑gapped local file‑path model source example in samples.

Source: llm_adapter@2026-07-15

Confidence: low

Full changelog

0.9.6 (2026-07-14)

Features

  • api: add GPUQuota CRD types for multi-tenant GPU governance (#1101) (5dd867a)
  • controller: add GPUQuota status reconciler (#1117) (1d583e4)
  • controller: add GPUQuota validating webhook for InferenceService (#1118) (26804bd)
  • controller: add s3:// model source via curl --aws-sigv4 (#1098) (#1125) (ed35142)
  • foreman: preserve a coder's gate-failed branch instead of discarding it (#1115) (bc39c77)
  • helm: gate InferenceService quota webhook and tenant RBAC behind multitenancy toggle (#1122) (26a994e)
  • quota: add GPUQuota admission decision function (#1107) (4897093)
  • runtime: close out the SGLang kitchen-sink (#1060): minor flags, accept thresholds, typed LoRA adapters, LoRAAdapter CRD (#1103) (8edf8bd)

Bug Fixes

  • foreman: honor GateProfile source extensions in scope-overlap issue-ref extraction (#1120) (0c20431)

Documentation

  • samples: add air-gapped local file-path model source example (#1099) (#1124) (913ff3d)

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track LLMKube

Get notified when new releases ship.

Sign up free

About LLMKube

Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.

All releases →

Related context

Earlier breaking changes

  • v0.8.1 foreman: requestTimeoutSeconds now sets loop-wide budget, default changes from 600 to 3600.

Beta — feedback welcome: [email protected]