DeepSeek

DeepSeek V4.1 Flash vs V4 Pro: We Checked the “Comprehensive Win” Claim

·

DeepSeek's September 10 release notes say V4.1 Flash has "comprehensively surpassed V4 Pro in performance, cost, speed, and total time." The company's own benchmark table mostly agrees: of the 16 benchmarks where both models have published scores, V4.1 Flash wins 14. But not all 16. On GPQA Diamond, the retiring flagship still leads 92.4 to 90.9. On HLE (text-only subset), it leads 42.7 to 39.1. Both holdouts are pure knowledge tests — every agentic, coding, terminal, and security benchmark goes to the new model, which also costs 70–86% less per token.

The short version

  • V4.1 Flash beats V4-Pro-0813 on 14 of 16 shared benchmarks — and undercuts it on every pricing line (cheapest: $0.003/1M cached input, off-peak)
  • The two exceptions are knowledge tests: GPQA Diamond (90.9 vs 92.4) and HLE text-only (39.1 vs 42.7)
  • Switch the model name to deepseek-flash; the old Flash names are retired but temporarily routed
  • From 04:00 UTC, September 14, all deepseek-v4-pro requests route to V4.1 Flash at Flash prices
Official benchmark table comparing DeepSeek V4.1 Flash with V4-Pro-0813, V4-Flash-0731, GLM 5.3, Kimi K3, GPT-5.6-Sol, and Claude Opus 5
DeepSeek's published benchmark table, September 10, 2026

1. What DeepSeek claimed, exactly

The launch post makes two claims worth separating:

  • "Benchmark results ahead of flagship models, including DeepSeek-V4-Pro" — from new pre-training methods plus larger-scale RL post-training.
  • "V4.1 Flash has comprehensively surpassed V4 Pro in performance, cost, speed, and total time" — the justification for retiring V4 Pro entirely.

"Comprehensively" is doing some work in that second sentence. DeepSeek's own table (linked above) compares V4.1 Flash against V4-Pro-0813, the previous V4-Flash-0731, GLM 5.3, Kimi K3, GPT-5.6-Sol, and Claude Opus 5 across 19 benchmarks. Against V4 Pro specifically, 16 rows have scores for both models. Flash wins 14.

2. The two benchmarks V4 Pro still wins

Benchmark V4.1-Flash V4-Pro-0813
GPQA Diamond 90.9 92.4
HLE (text-only subset) 39.1 42.7

Source: DeepSeek announcement, Sept 10, 2026.

HLE scores marked with an asterisk in DeepSeek's table are text-only subsets. V4-Pro-0813 has no full multimodal HLE score because V4 Pro doesn't support vision input — V4.1 Flash does, scoring 36.8 on the full HLE. So the fair reading: on text-only HLE, the old flagship wins 42.7 to 39.1; on vision benchmarks, there's no comparison to make.

Why the holdouts are knowledge tests

Both losses fit a pattern. GPQA Diamond and HLE measure dense scientific and academic knowledge recall. V4.1 Flash is built differently on purpose: a new Causal Encoder–Decoder architecture with just 8B active parameters for input and 16B for output (552B total, MoE) — an asymmetric design optimized for agent workloads that read far more than they write. Knowledge-heavy benchmarks reward something the new architecture visibly trades away. Whether that gap closes in the upcoming V4.1 Pro is the open question.

3. The 14 it wins — and by how much

Benchmark V4.1-Flash V4-Pro-0813 Δ
Terminal-Bench 2.1 90.6 87.9 +2.7
Terminal-Bench 3.0 30.0 11.8 +18.2
Terminal-Bench 4.0 31.2 12.4 +18.8
DeepSWE v1.1 74.2 62.7 +11.5
CyberGym 88.1 83.3 +4.8
SEC-Bench Pro 62.8 56.4 +6.4
ExploitGym 15.3 5.4 +9.9
Automation-Bench 54.8 43.2 +11.6
Agents' Last Exam 31.8 25.7 +6.1
NL2Repo-Bench 65.4 61.5 +3.9
ProgramBench 20.3 15.5 +4.8
HLE (w/ tools) 63.9 60.0 +3.9
Codeforces (rating) 3471 3348 +123
MathArena Apex 65.6 65.3 +0.3

The Terminal-Bench jumps are the headline: on the two newest versions, Flash scores roughly 2.5× the retiring flagship. It also beats its own predecessor (V4-Flash-0731) on all 16 shared rows.

Against the Western frontier models in the same table, the picture is mixed rather than dominant: Flash tops Terminal-Bench 2.1 (90.6 vs 89.1 for Claude Opus 5 and 88.8 for GPT-5.6-Sol) and edges both on DeepSWE v1.1 (74.2 vs 74.0 and 73.0). But Claude Opus 5 leads on Terminal-Bench 3.0/4.0 (43.3/51.8), NL2Repo-Bench (75.3), and ProgramBench (37.0), and GPT-5.6-Sol leads on ExploitGym (33.7) and HLE (44.5). At 3–10× the price, they should.

4. Pricing (USD card, effective 10 Sep 2026): Flash is cheaper on every line

Official USD pricing (source):

Per 1M tokens deepseek-flash (cheapest) deepseek-v4-pro
Input, cache hit — off-peak $0.003 $0.022
Input, cache hit — peak $0.006 $0.044
Input, cache miss — off-peak $0.15 $0.66
Input, cache miss — peak $0.30 $1.32
Output — off-peak $0.60 $1.98
Output — peak $1.20 $3.96

That's 86% cheaper on cache hits, 77% on cache misses, and 70% on output. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday; off-peak rates are 50% of peak. Cache-heavy agent pipelines — the ones that re-read long context every turn — capture most of the saving, which is exactly the workload the architecture targets.

deepseek-flash vs deepseek-v4-pro pricing page screenshot
Official pricing, effective 04:00 UTC September 10, 2026

Everything else the card tells you

Both models run 1M context with 384K max output, but only Flash supports vision, and Flash's concurrency limit is 2,500 versus Pro's 500. The cost cuts trace back to architecture: V4.1 Flash's KV cache needs just 1/4 the HBM and 1/8 the SSD storage of the previous generation (announcement).

5. Migration: what to change, and by when

Switch to deepseek-flash now. The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are already retired and route to V4.1 Flash at Flash prices.

deepseek-v4-pro is effectively dead at 04:00 UTC on September 14. Until V4.1 Pro ships, requests to it silently route to V4.1 Flash and bill at Flash rates.

After September 14, your "V4 Pro" responses will come from a different model. If any production flow relies on the old flagship's stronger knowledge recall (the GPQA/HLE gap above), re-run your evals this week — don't discover the regression in user reports.

Official partners WorkBuddy (including CodeBuddy) and OpenCode have fully integrated V4.1 Flash. Open weights are on Hugging Face, with the technical report (PDF) alongside.

6. FAQ

Is V4.1 Flash actually better than V4 Pro?

On 14 of the 16 benchmarks both models have published scores for, yes — often by a lot. It loses exactly two, both knowledge tests: GPQA Diamond (90.9 vs 92.4) and text-only HLE (39.1 vs 42.7). If your workload is graduate-level science recall rather than agentic coding, the retiring flagship still scores higher.

How much cheaper is it than V4 Pro?

On every line of the official USD card: 86% cheaper on cached input ($0.003 vs $0.022 per 1M off-peak), 77% on uncached input ($0.15 vs $0.66), 70% on output ($0.60 vs $1.98). Peak rates are double the off-peak rates on both models.

What happens to my app on September 14?

Nothing breaks. Requests to deepseek-v4-pro are routed to V4.1 Flash and billed at Flash rates until a V4.1 Pro ships. But your outputs will change, so re-run evals before the date — especially if you chose Pro for its knowledge recall.

Sources

Caveats, stated plainly. Benchmark figures are DeepSeek's own, read from its launch-day table; they have not been independently reproduced. Pricing is the official USD card as of 10 September 2026; prices can change, so re-check the console before committing a budget. This page will be updated if figures move — check the verified date at the top.

We have no affiliation with DeepSeek and carry no affiliate links on this page.