Skip to content

ollama

v0.32.1 Breaking

This release includes breaking changes for platform teams planning a safe upgrade.

Published 11d Model Serving & MLOps
✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

deepseek gemma gemma3 glm go gpt-oss
+8 more
llama llama3 llm llms minimax mistral ollama qwen

ReleasePort's take

Light signal
editorial:auto 10d

Ollama v0.32.1 resolves a recurrent MLX model cache leak that could cause memory usage to grow across requests and enhances cache snapshot performance.

Why it matters: Fixes a bug (severity 40) that caused unbounded memory growth in MLX caches; improves stability for all deployments using Ollama's MLX features.

Summary

AI summary

Fixed recurrent MLX cache leak and improved memory usage across requests.

Changes in this release

Feature Low

Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations

Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations

Source: llm_adapter@2026-07-16

Confidence: high

Feature Low

MLX text model loading now respects `OLLAMA_LOAD_TIMEOUT` environment variable

MLX text model loading now respects `OLLAMA_LOAD_TIMEOUT` environment variable

Source: llm_adapter@2026-07-16

Confidence: high

Feature Low

Agent web search and fetch now prompt users to run `ollama signin` when authentication is required

Agent web search and fetch now prompt users to run `ollama signin` when authentication is required

Source: llm_adapter@2026-07-16

Confidence: high

Feature Low

Interactive agent now receives the current working directory for better project context

Interactive agent now receives the current working directory for better project context

Source: llm_adapter@2026-07-16

Confidence: high

Bugfix Medium

Fixed recurrent MLX model cache leak that could increase memory use across requests and improved cache snapshot performance

Fixed recurrent MLX model cache leak that could increase memory use across requests and improved cache snapshot performance

Source: llm_adapter@2026-07-16

Confidence: high

Bugfix Medium

Fixed `ollama launch` so selecting **Pick another model** for a deprecated model passed with `--model` correctly opens the model picker

Fixed `ollama launch` so selecting **Pick another model** for a deprecated model passed with `--model` correctly opens the model picker

Source: llm_adapter@2026-07-16

Confidence: high

Full changelog

What's Changed

  • Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations
  • Fixed a recurrent MLX model cache leak that could increase memory use across requests, and improved cache snapshot performance
  • MLX text model loading now respects OLLAMA_LOAD_TIMEOUT
  • Agent web search and fetch now tell users to run ollama signin when authentication is required
  • The interactive agent now receives the current working directory for better project context
  • Fixed ollama launch so choosing Pick another model for a deprecated model passed with --model opens the model picker
  • Updated VS Code setup documentation for the official Ollama extension

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.0...v0.32.1-rc0

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track ollama

Get notified when new releases ship.

Sign up free

About ollama

Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

All releases →

Related context

Related tools

Beta — feedback welcome: [email protected]