Ollama v0.33.1 Adds Qwen3.8-Flash-Next Support via MLX Backend
1 min readOllama v0.33.1 extends its model support with native Qwen3.8-Flash-Next compatibility through the MLX backend, directly benefiting Apple Silicon users. The release also introduces structured output support in the MLX runner and addresses GPU timeout issues when loading models from slower storage, improving the user experience for resource-constrained environments.
This update is significant for macOS developers and enthusiasts running local LLMs on Apple Silicon hardware. Qwen3.8-Flash-Next's efficiency profile makes it particularly suitable for M-series Macs, where MLX backend acceleration provides near-native performance. The structured output feature enables more reliable programmatic usage of local models for agents and tool-calling scenarios, expanding the practical applications for on-device inference.
Read the full article on Ollama release.
Source: Ollama release · Relevance: 9/10