This release adds 2 notable features for engineering teams evaluating rollout.
Published 28d
LLM Frameworks
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
benchmarks
cli
gguf
gpu
+7 more
huggingface
inference
llm
local-llm
ollama
python
vram
Summary
AI summaryImproved VRAM estimates for sliding-window attention models and added detection for Intel Arc Pro B70 / Battlemage G31.
Full changelog
Added
- Sliding-window attention metadata is now used in model fetching and KV cache estimation, improving VRAM estimates for models that use SWA. (#124)
- Intel Arc Pro B70 / Battlemage G31 now has curated detection and simulation defaults, including PCI device
0xe223, 32 GB VRAM, and 608 GB/s bandwidth. (#93, #136)
Fixed
- Model and benchmark metadata fetches now request
gzip, deflateinstead of brotli, avoiding brokenbrresponses from mirrors or intermediate servers. (#128, #136)
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
Track Find the best local LLM for your hardware, ranked by benchmarks
Get notified when new releases ship.
Sign up freeAbout Find the best local LLM for your hardware, ranked by benchmarks
All releases →Related context
Related tools
Beta — feedback welcome: [email protected]