This release adds 2 notable features for engineering teams evaluating rollout.
Published 1mo
LLM Frameworks
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
benchmarks
cli
gguf
gpu
+7 more
huggingface
inference
llm
local-llm
ollama
python
vram
Summary
AI summarySearch terms like "7B" match model parameter size accurately.
Full changelog
Added
HF_ENDPOINTsupport for Hugging Face model metadata fetches, so users behind a mirror can point whichllm at a compatible Hub endpoint. (#128, #131)- Manual detected-GPU overrides for usable VRAM and bandwidth, useful for iGPU and unified-memory systems where automatic detection is too conservative. (#132, #133)
- README guidance for safer first-run flags when users want full-GPU, usable-speed recommendations with extra VRAM headroom.
Fixed
- Search terms such as
7B,0.5B, and500Mnow match model parameter size instead of plain substrings, soqwen 7bno longer returns1.7Bor30B-A3Bby accident. (#107, #126) - GGUF sizing now treats FP16 and ternary
TQ1_0/TQ2_0quant types correctly, avoiding underestimates for those files. (#125)
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
Track Find the best local LLM for your hardware, ranked by benchmarks
Get notified when new releases ship.
Sign up freeAbout Find the best local LLM for your hardware, ranked by benchmarks
All releases →Related context
Related tools
Beta — feedback welcome: [email protected]