This release adds 2 notable features for engineering teams evaluating rollout.
Published 4d
Model Serving & MLOps
✓ No known CVEs patched
✓ No known CVEs patched in this version
Topics
deepseek
gemma
gemma3
glm
go
gpt-oss
+8 more
llama
llama3
llm
llms
minimax
mistral
ollama
qwen
Summary
AI summaryFixed model downloads that stall before sending data.
Full changelog
What's Changed
- Fixed model downloads that stall before sending data.
- Improved integrations: restored Claude Code Channels, fixed Anthropic thinking streams, and made Hermes Desktop respect
--force-build. - Expanded GPU support with CUDA on Windows ARM64, B200 support through CUDA 12, and lower memory use on Linux CUDA/ROCm iGPUs.
- Added chat, thinking, and tool calling support for Laguna 2.1 models, including a Metal inference fix.
- Fixed GLM tool calls being silently dropped at the end of generation.
- Updated the MLX and llama.cpp engines.
Full Changelog: https://github.com/ollama/ollama/compare/v0.32.1...v0.32.3
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About ollama
Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Beta — feedback welcome: [email protected]