AI coding tools

MiniMax M3.1-Flash-Preview Review

·

MiniMax has a problem that most Chinese AI labs would love to have: its newest coding model is getting praised in the wrong language. M3.1-Flash-Preview launched quietly on September 27, 2026, and the people who have actually stress-tested it are writing in Chinese. Two independent rounds — one from the tech outlet Zimubang on October 1, another from APPSO at ifanr on October 2 — ran the same prompts through both M3.1-Flash and Claude Opus 5.5, and both walked away saying roughly the same thing: on frontend and creative coding work, the cheap model holds. The English-speaking decision layer — the "should I use this" coverage — is a benchmark card on daily.dev and not much else. So here is that missing piece: what the Chinese tests actually did, where the model taps out, and what it costs you. One of those numbers expires on October 14.

Short answer

  • MiniMax M3.1-Flash-Preview (model ID MiniMax-M3.1-Flash-Preview) is a 428B-parameter MoE with ~23B active parameters, a 1M-token context window, and ~150 tokens/s output, per MiniMax's own agent-tools documentation. It takes text, image, and video input.
  • Two independent Chinese hands-on rounds (Zimubang, Oct 1; APPSO/ifanr, Oct 2) ran identical prompts against Claude Opus 5.5 — UI motion design, pixel animation, a full music video, WebGL painting tools, 3D scenes — and concluded the Flash model was not the weak side of the pairing on frontend work.
  • The part that stings: on the KingBench 3 coding suite it scores 66.25% (53/80) — a 35-point jump over M3's 31.25%, but still 27.5 points behind Opus 5.5's 93.75%. Chinese video reviewers who tried actual generation tasks report it "scores high but breaks in practice" on things like elevator simulations and archery games.
  • Access is the catch for API buyers: M3.1-Flash-Preview has no published pay-as-you-go price — it's not on MiniMax's API pricing page and not on OpenRouter. You get it through the M Plan subscription: $22/$55/$132 per month (Go/Explore/Build), first month 50% off through October 14, usable in MiniMax Code or via API key in Claude Code, Cursor, Codex, and friends.
  • Want to pay per token on MiniMax today? That's still MiniMax-M3 at $0.30/$1.20 per million tokens (≤512K input tier, permanent discount pricing), with the same 1M context.

What the tests actually did

Both rounds used the same blunt methodology: take prompts that had already produced good results with Opus 5.5, feed them to M3.1-Flash-Preview unchanged, and compare what comes back. No tuning, no cherry-picking, no retry tournaments. Zimubang's three paired tasks are the most telling:

  • A one-shot motion design piece — animated charts on a dark canvas. Zimubang's observation: Opus 5.5 rendered the charts in segments, while M3.1-Flash delivered complete charts with hover states that showed the underlying values.
  • A 160×90 retro pixel-art runner game — both models produced playable results; M3.1-Flash's version fired lasers in multiple directions and drew a larger main character.
  • A K-pop style music video from a song prompt ("Claude-Pop", for those keeping score at home) — here Zimubang's verdict favored the challenger outright: Opus 5.5's camera stayed nearly static with no real shot breakdown, while M3.1-Flash's 2D animation carried actual visual expressiveness.

APPSO's round at ifanr went broader instead of head-to-head: a 15-second motion portfolio, a hand-drawn short, an ink-wash recreation of the classic Chinese animation Tadpoles Looking for Mother, a New York skyline piece where the time of day responds to a draggable timeline, an arcade kart racer, a submarine exploration game, and a little ghost story game set on an autumn farm. Their bottom line, translated and paraphrased: "the Flash positioning did not become a visible disadvantage in finished-work quality." That's a reviewer who ran seven deliberately different genres of frontend work and found the model could organize a coherent visual idea in each — which tracks with the model's training story: MiniMax says the M3 series was trained multimodal from pretraining onward, so visual understanding and code aren't bolted together after the fact.

One number circulating in the Chinese coverage deserves its own caveat: a community case relayed (not run) by Zimubang claims M3.1-Flash generated a complete blog system in one pass at roughly 160 tokens/s, consuming about 70 million tokens over the run. Treat that as an anecdote with a receipt problem — we could not find the original source, and 70M tokens on one task is a workload profile most people will never have. The official ~150 tokens/s figure from MiniMax's docs is in the same neighborhood, which is what makes the anecdote plausible rather than proven.

There's also an interesting provenance thread: before the official name dropped, the model was rumored to be the anonymous "Space Bunny" topping frontend leaderboards, and OpenCode co-founder Dax publicly rated its frontend design as aligned with Opus 5.5's output. MiniMax never confirmed the alias. Both Chinese test rounds mention it; neither depends on it.

The numbers that don't flatter it

Here is the part the hype threads skip. On KingBench 3 — the coding benchmark whose numbers were aggregated by daily.dev from AGI Hunt's testing — the release-over-release jump is real and large:

ModelKingBench 3
Claude Opus 5.593.75%
SWE283.75%
GPT6 Soul82.5%
MiniMax M3.1-Flash-Preview66.25% (53/80)
MiniMax M331.25%

A 35-point climb in one release is a serious improvement. It also leaves the model fourth in this particular table, 27.5 points behind the flagship it keeps getting compared to. And the Chinese testing culture did not let that slide: Bilibili hands-on reviewers who took the KingBench 3 score at face value and then ran real generation tasks — an elevator simulator, an archery game — reported the score-to-output gap showing up as visible breakage. One video is literally titled "scores high, but is the actual generation stable or not" and answers: not yet, entirely. That matches the shape of the model: MoE with ~23B active parameters is a routing-efficient design, not a brute-force one, and it shows when a task needs sustained multi-step rigor instead of a strong single-pass aesthetic judgment.

So the two Chinese test rounds and the benchmark crowd are not contradicting each other. They're measuring different things. Frontend creative work — where the model's visual pretraining and its ~150 tokens/s speed both pay off — is its sweet spot. Hard agentic coding with strict acceptance criteria is where the 66% pass rate means roughly one task in three comes back broken. If your pipeline can absorb a retry, that's a cost question. If it can't, that's a disqualifier.

What it costs you

This is where the story gets weird for anyone used to buying tokens. M3.1-Flash-Preview is not on MiniMax's API pricing page. It's not on OpenRouter. As of October 9, 2026, there is no published per-token price anywhere official. The model lives inside subscriptions:

M Plan tierMonthlyRelative usageVideo model
Go$221×—
Explore$553×H3
Build$1327.5×H3

All three tiers use M3.1-Flash-Preview as the text model — there is no better tier to buy your way into. You can use it in the MiniMax Code desktop app, or grab an API key and point Claude Code, Codex, Cursor, or OpenCode at it; MiniMax supports both OpenAI-compatible and Anthropic-compatible endpoints. One structural detail: the model cannot turn thinking off. It ships with five effort levels (Low through Max, default Max), and MiniMax's own migration guide warns that assuming "just swap the model name" will bite you on how thinking content is read and output budgets are set.

The launch offer: first month 50% off any monthly tier, through October 14, 2026. That makes the entry ticket effectively $11 for a month of trying it — less than a single Opus 5.5 subscription coffee run, if you've read our Claude API heavy-user cost breakdown and know what the flagship habit actually runs.

If you want to pay per token on MiniMax, the model you can actually buy that way is still MiniMax-M3 — same 1M context, same API family, the model M3.1-Flash improved on:

MiniMax-M3 (pay-as-you-go)Input / MOutput / MCache read / M
Prompts ≤ 512K$0.30$1.20$0.06
Prompts > 512K$0.60$2.40$0.12

Those are the "permanent 50% off" prices from MiniMax's own pricing page, and they match OpenRouter's live listing to the cent ($0.30/$1.20, context 1,048,576). For scale: against Claude Opus 5.5's $4/$20 per million, M3's input is 7.5% of the price and output is 6%. That's the same budget tier as DeepSeek V4.1 Flash — our Flash-tier comparison covers that neighborhood, and Zhipu's official pricing is in the GLM-5.3-Flash price breakdown.

Should you switch

Our read, as a buyer and not a benchmark enjoyer:

  • Use it for: frontend prototypes, marketing-site demos, creative one-shots, "make this pretty and interactive" tasks where a strong single pass matters more than sustained correctness. Two independent reviewers who actually ran it against Opus 5.5 came to the same conclusion on this class of work. At $11 for the first month (through October 14), the trial math is trivial.
  • Don't use it for: the load-bearing parts of an agent pipeline. A 66% KingBench 3 score against Opus's 93.75% is not a rounding difference, and the "scores high, breaks in practice" reports from Chinese reviewers land exactly where you'd expect — long tasks with state and physics. This is not your agent harness backbone today.
  • Read the name: it says Preview, and the pricing model itself is mid-migration — the older Token Plan is being folded into M Plan. A model whose billing structure is being reorganized is a model whose contract you shouldn't build a quarter of roadmap on. The budget-model field is also getting crowded fast — Claude Haiku 5.5 just landed at $0.10/$0.50 — so holding a month of M Plan and re-evaluating is a legitimate strategy, not indecision.

Sources and caveats

Everything on this page was checked on October 9, 2026. Prices, plan tiers, and the October 14 offer deadline are a snapshot — MiniMax's own docs carry a "no long-term public price confirmed" disclaimer on the preview model, so re-verify before you commit budget. The hands-on findings are aggregated and paraphrased from two Chinese-language test rounds — Zimubang's M3.1 hands-on (Oct 1) and APPSO's same-prompt comparison (Oct 2) — plus KingBench 3 figures as aggregated by daily.dev. We did not run these tests ourselves, we can't read Chinese benchmark forums the way a native reviewer can, and the 70-million-token blog system anecdote is relayed, not verified. Official numbers come from MiniMax's pricing page, agent tools documentation, and M Plan offer page; Opus 5.5 pricing from Anthropic's model documentation. No affiliate links, no sponsored placements, no free compute — nobody paid for this page.