This release adds 4 notable features for engineering teams evaluating rollout.
Published 1mo
LLM Frameworks
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
ai
apple-silicon
benchmarks
cli
gguf
gpu
+7 more
huggingface
inference
llm
local-llm
ollama
python
vram
Summary
AI summaryAdded multi‑GPU support, T5 lineage handling, and new CLI invocations.
Full changelog
Added
- Multi-GPU simulation for repeated
--gpuflags, comma-separated GPU specs, and count shorthand like2x RTX 4090. python -m whichllmnow runs the CLI.--gpu-onlyand--fit full-gpufilter recommendations to models that fit fully in GPU VRAM.- T5 lineage support for version-aware benchmark handling.
Fixed
- Cached model and benchmark data are read as UTF-8.
- GTX 1650 simulation distinguishes GDDR5 and GDDR6 variants by memory clock.
- RAM reserve logic now uses a bounded reserve formula instead of a fixed 80% usable-RAM cap.
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
Track Find the best local LLM for your hardware, ranked by benchmarks
Get notified when new releases ship.
Sign up freeAbout Find the best local LLM for your hardware, ranked by benchmarks
All releases →Related context
Related tools
Beta — feedback welcome: [email protected]