Skip to content

This release adds 2 notable features for engineering teams evaluating rollout.

Published 28d LLM Frameworks
✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

ai apple-silicon benchmarks cli gguf gpu
+7 more
huggingface inference llm local-llm ollama python vram

Summary

AI summary

Improved VRAM estimates for sliding-window attention models and added detection for Intel Arc Pro B70 / Battlemage G31.

Full changelog

Added

  • Sliding-window attention metadata is now used in model fetching and KV cache estimation, improving VRAM estimates for models that use SWA. (#124)
  • Intel Arc Pro B70 / Battlemage G31 now has curated detection and simulation defaults, including PCI device 0xe223, 32 GB VRAM, and 608 GB/s bandwidth. (#93, #136)

Fixed

  • Model and benchmark metadata fetches now request gzip, deflate instead of brotli, avoiding broken br responses from mirrors or intermediate servers. (#128, #136)

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track Find the best local LLM for your hardware, ranked by benchmarks

Get notified when new releases ship.

Sign up free

About Find the best local LLM for your hardware, ranked by benchmarks

All releases →

Beta — feedback welcome: [email protected]