AI news

Qwen3.8 Flash vs DeepSeek 4.1 Flash: Two 'Flash' Tiers, One Million Tokens Each

·

Short answer

  • Qwen3.8 Flash and DeepSeek V4.1 Flash both serve a 1M-token context window by default. Qwen's list price is the simpler read: $0.15 per 1M input tokens, about $0.47 per 1M output. DeepSeek answers with structure — cache-hit discounts and an off-peak window — instead of one headline number.
  • 14 September was the deadline DeepSeek's own release notes set for retiring V4 Pro. Its changelog says Pro continues. As of tonight, neither page has been updated. If you run deepseek-v4-pro, check your served-model logs before you trust either page.

Late August quietly produced a second budget tier worth taking seriously. On 26 August Qwen published Qwen3.8-Flash-Next, an architecture preview, then followed it with Qwen3.8 Flash — the production version, now served on QwenCloud with a 1M context window by default, official built-in tools, and an input price listed at $0.15 per million tokens. The problem for Qwen: DeepSeek's V4.1 Flash, live since 10 September, already promises the same million-token window, with a 552B-parameter architecture and open weights behind it. Two budget tiers, both flashing a 1M badge. This page puts them side by side the way we put everything on this site: by what you can verify today — and it lands on an inconvenient date for DeepSeek, which we get to at the end.

What Qwen3.8 Flash actually shipped

The release is a two-step. Qwen3.8-Flash-Next (26 August) is the architecture preview, published with weights on Hugging Face; its model card describes Qwen3.8 Flash as the official production version built on it, "with more production features, e.g., 1M context length by default, official built-in tools." Qwen's own homepage listing, as indexed on 14 September, serves the production model on QwenCloud "priced at 0.15 USD per..." million input tokens — the page is a JavaScript app we can't quote beyond the indexed text, so treat that figure as Qwen's own headline number, re-verify against their price page before you budget.

Third-party listings agree on the shape and disagree on one number. Input is consistently $0.15–0.16 per 1M. Output is $0.47 per 1M on two aggregators (one rounds to $0.50). The cached-input price is the outlier: one aggregator lists $0.016 per 1M, another $0.03 per 1M — a near-2x disagreement nobody has settled in print.

Two million-token windows, two pricing philosophies

On specs you can check today, the two models line up like this:

Qwen3.8 Flash DeepSeek V4.1 Flash
Context window 1M by default 1M
Max output tokens not listed at check time 384K
Open weights preview weights (Flash-Next) on Hugging Face full 552B weights on Hugging Face
Multimodal not stated at check time native vision in the main model
Architecture detail preview-stage only MoE, 8B-in / 16B-out active per token

The pricing philosophies diverge more than the specs. Qwen posts a flat, quotable rate card: $0.15/M in, ~$0.47/M out, cached input TBD pending the disagreement above. DeepSeek doesn't lead with a single number — its Flash tier layers a deep cache-hit discount and a 50%-off off-peak window (UTC 01:00–04:00 and 06:00–10:00, Mon–Fri) on top of rates it cut at launch, which makes the real comparison a function of your traffic pattern, not of their headline. We keep DeepSeek's numbers on our cheapest-provider page and let you run your own volumes through the cost calculator rather than pasting rates that rot; our deployment guide covers the behavioral side (thinking mode on by default, silent model-name reroutes) that spec sheets never show. One honest gap: we haven't yet run Qwen3.8 Flash through the same trap-hunting — when we do, this page gets a dated amendment.

If both Flash tiers are on your shortlist, the deciding input isn't either vendor's rate card — it's your own token mix. A cache-heavy agent workload and a cold-start batch job will pick different winners.

The V4-Pro deadline came and went — with both pages still standing

Today is the date DeepSeek set for its own model's retirement, and its documentation still can't agree on it. Both passages below are live on the official site as of 14 September, ~14:00 UTC:

"We're phasing out V4-Pro. Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches." — release notes, Sept 10

"In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." — changelog, Sept 10 (still the latest entry)

04:00 UTC has passed. Neither page was updated, and the changelog shows no new entry since 10 September. Our read: the reroute announced in the release notes did not visibly happen, the changelog's "user demand" clause is the operative policy — but the correct move for anyone billing against deepseek-v4-pro today is to verify the model string in your own response logs and the rate on your invoice, not to trust either page. (Background on the tiers themselves: V4.1 Flash vs V4 Pro, and how the launch day unfolded in our Sept-10 report.)

Everything on this page was checked on 14 September 2026; the DeepSeek situation in particular may resolve in either direction the day after publication, and this page will carry a dated correction if it does.

Sources

  • Qwen homepage listing (Qwen3.8 Flash production entry, 1M context / built-in tools / $0.15 per 1M input), as indexed 14 September 2026
  • Hugging Face model card — Qwen/Qwen3.8-Flash-Next (26 August 2026), incl. production-version description
  • Aggregator listings for Qwen3.8 Flash pricing: llm-stats.com, datacamp.com, kie.ai, mindstudio.ai (26–29 August entries; cached-input figures conflict, both cited)
  • api-docs.deepseek.com — /updates/ (changelog) and /news/news260910 (release notes), read live 14 September 2026, quoted verbatim above
  • Novita public model list — both Flash models listed as of 14 September (plain citation, no affiliate link)