This release adds 2 notable features for engineering teams evaluating rollout.
Published 2d
Model Serving & MLOps
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
deepseek
gemma
gemma3
glm
go
gpt-oss
+8 more
llama
llama3
llm
llms
minimax
mistral
ollama
qwen
Summary
AI summaryFixed Qwen3 MoE decoding for differently‑quantized experts and sped up packed gate/up projection by ~4–9% on Apple M5 Max.
Full changelog
What's Changed
- Support Laguna on Apple GPUs via the MLX engine
- Quantize draft-model output heads at the requested type when creating speculative-decoding drafts.
- Fixed Qwen3 MoE decoding for differently-quantized experts, plus faster packed gate/up projection (~4–9% on M5 Max).
Full Changelog: https://github.com/ollama/ollama/compare/v0.32.3...v0.32.4
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About ollama
Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Beta — feedback welcome: [email protected]