Ollama v0.33.1 Adds Qwen3.8 Flash Next Support and Claude Desktop Integration
1 min readOllama's latest release marks a significant milestone for local LLM deployment by introducing native support for Qwen3.8 Flash Next through its MLX runner. This development enables users to run cutting-edge models directly on Apple Silicon and other supported hardware without cloud dependencies. The integration with Claude Desktop as a third-party gateway provider represents a major shift in how local models can be consumed by popular AI applications.
Beyond model support, v0.33.1 introduces critical stability improvements including better prefill caching and fixes for agent clients that previously hung during long context processing. These engineering improvements address real-world pain points for practitioners running local inference in production environments. Combined with structured output support in the MLX runner, this release makes Ollama a more robust platform for edge deployment scenarios.
The timing is crucial as Qwen3.8 Flash Next represents a new generation of efficient models designed for on-device inference. With Ollama's streamlined integration, practitioners can now leverage this model with minimal setup complexity while maintaining full local control over data and inference.
Read the full article on Ollama Release.
Source: Ollama Release · Relevance: 9/10