Self-hosting
Ollama v0.40.0-rc0: Local Models Now Run on MLX by Default on Apple Silicon
Ollama has been the "it just works" option for running local models on a Mac for a long time, and it has quietly shipped everything through the llama.cpp runtime. Release candidate v0.40.0-rc0, published 25 Sep 2026, changes that default for Apple Silicon: supported model architectures now run on MLX automatically. Here is what actually changes, what the performance numbers do and do not say, and why you might want to wait for the stable release.
The short answer
- What changed. In v0.40.0-rc0, on Apple Silicon devices, model architectures supported by the MLX runtime run on MLX by default — no flags, no config. Ollama's own example:
ollama pull qwen3.8thenollama run qwen3.8. - Who benefits. Anyone running supported architectures on M-series Macs, especially at long context. Community numbers put MLX ahead of llama.cpp's Metal backend in most tests, but the margin depends heavily on the model and context length — there is no honest single multiplier.
- Should you switch now? If you live on the stable channel, wait for v0.40.0 stable. It is a release candidate: Ollama says it will "be testing and enabling additional models" during the pre-release, and the official docs have not been updated yet. Nothing here is wrong — it is just early.
What changed in v0.40.0-rc0
The release notes are short, and the substance is one line: on Apple Silicon, model architectures supported by the MLX runtime will automatically run on MLX. Before this release, Ollama served every model through llama.cpp (with Metal acceleration on macOS). The engine choice is now made for you, per architecture, at run time. The same release line (v0.34.4, published 23 Sep 2026) also landed a fix that makes structured outputs on thinking models apply in a single pass, plus faster prompt processing for Qwen 3.8 on Apple Silicon — so the MLX work and the Apple Silicon tuning are clearly arriving together.
Note the version jump: v0.40.0-rc0 sits on top of the v0.34 stable line. That is a large minor bump for one headline feature, which suggests the MLX runtime integration has been in progress for a while.
What MLX default means in practice
For most Mac users, the change is invisible in the best way: pull a model, run a model, and if its architecture is covered by the MLX runtime, you get MLX without touching anything. The catch is the phrase "supported by the MLX runtime." The release notes name one example (Qwen 3.8) and promise more models during the pre-release, but do not publish the architecture list. The official docs (docs directory on GitHub, checked 26 Sep 2026) contain no MLX page at all — gpu.mdx and faq.mdx do not mention MLX once. So for now, the practical way to know whether your model is running on MLX is to watch the load output in the terminal.
There is also no documented switch to force a model back onto llama.cpp. That matters if you rely on llama.cpp-specific behavior — GGUF quantization choices, custom flags, or specific sampler behavior. If that describes you, stay on v0.34 stable until the docs catch up.
How much faster is MLX, really?
Honest answer: it depends, and most published numbers are thin. Two community data points we could actually verify as of 26 Sep 2026:
- Paperhallway's MLX-vs-llama.cpp piece (25 Jun 2026) measured MLX 15–25% faster than llama.cpp's Metal backend on the same hardware at 32K-token contexts.
- Julien Simon's Arcee comparison (10B and 32B models) got roughly 51 tok/s on MLX vs 47 tok/s on llama.cpp — about 10%.
Those two do not contradict each other; they measure different models at different context lengths. The pattern across both: the longer the context, the bigger MLX's edge, because MLX's memory management on unified memory handles long sequences better than the Metal path. At short contexts, expect single-digit gains. Treat any claim of "2x faster" with suspicion — we found none from a source we would stake a page on.
Caveats before you switch
- It is an rc, labeled as one. The stable line is v0.34.4. If your local setup is load-bearing for work, rc software is how weekend plans die.
- Docs lag the feature. No MLX page, no architecture list, no fallback switch documented (checked 26 Sep 2026). Expect all three to land with or shortly after the stable release.
- Model coverage is partial. Unsupported architectures keep running on llama.cpp — the default change is per-architecture, not global.
How to try it
Grab the v0.40.0-rc0 pre-release from the Ollama GitHub releases page, then the official two-liner applies: ollama pull qwen3.8, ollama run qwen3.8. Watch the startup logs to see which runtime serves your model. If you are picking hardware for local inference rather than tuning a runtime, our upcoming self-hosting hardware breakdown covers the DGX Spark vs Mac side of that question.
Update anchor: this page will be revised when Ollama ships v0.40.0 stable, with the confirmed architecture list and any official MLX documentation.
Sources
- Ollama release v0.40.0-rc0 notes, GitHub, 25 Sep 2026 — MLX-by-default behavior, Qwen 3.8 example, pre-release model expansion.
- Ollama release v0.34.4 notes, GitHub, 23 Sep 2026 — single-pass structured outputs, Qwen 3.8 prompt-processing fix on Apple Silicon.
- Ollama GitHub docs directory (gpu.mdx, faq.mdx), checked 26 Sep 2026 — no MLX documentation present.
- Paperhallway, "Ollama MLX Apple Silicon Performance 2026", 25 Jun 2026 — 15–25% MLX advantage at 32K context (community measurement).
- Julien Simon, llama.cpp vs MLX comparison on 10B/32B Arcee models — ~51 vs ~47 tok/s (community measurement).