Skip to content

LLMKube

v0.8.17 Feature

This release adds 3 notable features for engineering teams evaluating rollout.

✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

ai apple-silicon autoscaling edge-computing gguf gpu
+12 more
self-hosted inference kubernetes llama-cpp llm local-llm metal mlx multi-gpu nvidia tgi vllm

Summary

AI summary

Updates 0.8.17, https://github.com/defilantech/LLMKube/issues/839, and Bug Fixes across a mixed release.

Full changelog

0.8.17 (2026-06-25)

Features

  • cli: explicit model-cache control flags on deploy/delete (Fixes #722) (#830) (9fc4187)
  • foreman: durable audit-log record + llmkube audit export (#837) (#838) (57d2ad0)
  • foreman: GateProfile API type + language presets + resolver (Addresses #839) (#840) (165a375)
  • foreman: run the fast gate from the resolved GateProfile (Addresses #839) (#841) (4225da4)
  • foreman: scope guard reads source extensions from GateProfile (Addresses #839) (#842) (c36ac87)
  • foreman: verify-gate Job image from GateProfile (Addresses #839) (#843) (327dd88)
  • foreman: verify-gate Job runs resolved commands for non-Go profiles (Addresses #839) (#844) (84994fd)
  • inferenceservice: speculativeDecoding for llama.cpp MTP/draft (Fixes #502) (#827) (564726a)
  • router: optional backend displayName for /v1/models id (Fixes #792) (#826) (7090f83)
  • router: Prometheus metrics + OTel spans for routing decisions (Addresses #433) (#834) (b8fa96d)

Bug Fixes

  • foreman: scope guard skips check on zero-Go-file diffs; NUL-safe git status (Fixes #800) (#823) (92e5501)

Documentation

  • examples: spot-capacity GPU NodePool for Foreman gate Jobs (Fixes #659) (#824) (9fa4846)
  • Karpenter GPU autoscaling guide (Fixes #658) (#825) (2d586ef)

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track LLMKube

Get notified when new releases ship.

Sign up free

About LLMKube

Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.

All releases →

Related context

Earlier breaking changes

  • v0.8.1 foreman: requestTimeoutSeconds now sets loop-wide budget, default changes from 600 to 3600.

Beta — feedback welcome: [email protected]