Skip to content

ollama

v0.32.4 Feature

This release adds 2 notable features for engineering teams evaluating rollout.

✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

deepseek gemma gemma3 glm go gpt-oss
+8 more
llama llama3 llm llms minimax mistral ollama qwen

Summary

AI summary

Fixed Qwen3 MoE decoding for differently‑quantized experts and sped up packed gate/up projection by ~4–9% on Apple M5 Max.

Full changelog

What's Changed

  • Support Laguna on Apple GPUs via the MLX engine
  • Quantize draft-model output heads at the requested type when creating speculative-decoding drafts.
  • Fixed Qwen3 MoE decoding for differently-quantized experts, plus faster packed gate/up projection (~4–9% on M5 Max).

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.3...v0.32.4

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track ollama

Get notified when new releases ship.

Sign up free

About ollama

Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

All releases →

Related context

Related tools

Beta — feedback welcome: [email protected]