Skip to content

LLMKube

v0.9.8 Feature

This release adds 3 notable features for engineering teams evaluating rollout.

✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

ai apple-silicon autoscaling edge-computing gguf gpu
+12 more
self-hosted inference kubernetes llama-cpp llm local-llm metal mlx multi-gpu nvidia tgi vllm

Summary

AI summary

Updates 0.9.8, Bug Fixes, and 2026-07-19 across a mixed release.

Changes in this release

Feature Low

Adds InferenceService hardware-labels info metric to controller.

Adds InferenceService hardware-labels info metric to controller.

Source: llm_adapter@2026-07-20

Confidence: medium

Feature Low

Adds llamacpp-router runtime for multi-model serving in controller.

Adds llamacpp-router runtime for multi-model serving in controller.

Source: llm_adapter@2026-07-20

Confidence: medium

Feature Low

Adds explicit vllm serve entrypoint for image‑agnostic vLLM in controller.

Adds explicit vllm serve entrypoint for image‑agnostic vLLM in controller.

Source: llm_adapter@2026-07-20

Confidence: medium

Feature Low

Passes staged model directory to vLLM/SGLang for multi‑file models in controller.

Passes staged model directory to vLLM/SGLang for multi‑file models in controller.

Source: llm_adapter@2026-07-20

Confidence: medium

Feature Low

Introduces ROCm/HIP runtime tier for AMD nodes in controller.

Introduces ROCm/HIP runtime tier for AMD nodes in controller.

Source: llm_adapter@2026-07-20

Confidence: medium

Feature Low

Gates foreman‑finalize on Foreman verify verdict in finalize component.

Gates foreman‑finalize on Foreman verify verdict in finalize component.

Source: llm_adapter@2026-07-20

Confidence: medium

Feature Low

Adds in‑executor envtest gate feedback loop to foreman.

Adds in‑executor envtest gate feedback loop to foreman.

Source: llm_adapter@2026-07-20

Confidence: medium

Feature Low

Introduces provider‑neutral CodeHost/WorkItems/ChangePolicy seams in foreman.

Introduces provider‑neutral CodeHost/WorkItems/ChangePolicy seams in foreman.

Source: llm_adapter@2026-07-20

Confidence: medium

Feature Low

Enables per‑InferenceService SLO declaration via Pyrra in observability.

Enables per‑InferenceService SLO declaration via Pyrra in observability.

Source: llm_adapter@2026-07-20

Confidence: medium

Bugfix Low

Updates boilerplate copyright year from 2025 to 2026.

Updates boilerplate copyright year from 2025 to 2026.

Source: llm_adapter@2026-07-20

Confidence: medium

Full changelog

0.9.8 (2026-07-19)

Features

  • controller: add InferenceService hardware-labels info metric (#1121) (#1160) (039adaa)
  • controller: add llamacpp-router runtime for multi-model serving (#1152) (fae5983)
  • controller: explicit vllm serve entrypoint for image-agnostic vLLM (#1164) (#1165) (3eeb580)
  • controller: pass staged model directory to vLLM/SGLang for multi-file models (#1157) (#1159) (a24fa94)
  • controller: ROCm/HIP runtime tier for AMD nodes (#701) (#1154) (3704d47)
  • finalize: gate foreman-finalize on the Foreman verify verdict (#1150) (#1153) (5c7b2cd)
  • foreman: in-executor envtest gate feedback loop (#768) (#1135) (074544a)
  • foreman: provider-neutral CodeHost/WorkItems/ChangePolicy seams (#1158) (#1161) (d3a8ad6)
  • observability: per-InferenceService SLO declaration via Pyrra (#415) (#1149) (4ae022c)

Bug Fixes

Documentation

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track LLMKube

Get notified when new releases ship.

Sign up free

About LLMKube

Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.

All releases →

Related context

Earlier breaking changes

  • v0.8.1 foreman: requestTimeoutSeconds now sets loop-wide budget, default changes from 600 to 3600.

Beta — feedback welcome: [email protected]