DeepSeek V4 Pro Retirement Canceled — It Stays Live Past Sept 14
Short answer
DeepSeek V4 Pro is not retiring. On September 14, 2026 — the day it was supposed to be phased out — official docs confirmed the model continues to be served, billing unchanged. If you use deepseek-v4-pro, do nothing.
The retirement is officially called off
DeepSeek's original plan, stated in their docs since the V4.1 Flash launch, was blunt: API service for V4 Pro would end after September 14, 2026. Multiple outlets (including our own coverage last week) reported the retirement as scheduled.
That changed. The current official pricing page now reads:
"In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. We will provide further notice should there be any changes." — DeepSeek API Docs, Models & Pricing, retrieved September 15, 2026
No new pricing tier, no migration step, no deprecation shim. The model name deepseek-v4-pro (currently serving DeepSeek-V4-Pro-0813) keeps working exactly as before.
Our earlier report — V4.1 Flash Launches — and V4 Pro Retires on Sept 14 — accurately reflected the official plan at the time. That plan has since been reversed. This post supersedes it on the retirement question; the Flash launch details there still stand.
V4 Pro pricing, verified today
Numbers below are straight from the official pricing page as of September 15, 2026 — note that some third-party pricing roundups are already out of date on these figures.
| Per 1M tokens | Flash off-peak | Flash peak | V4 Pro off-peak | V4 Pro peak |
|---|---|---|---|---|
| Input (cache hit) | $0.003 | $0.006 | $0.022 | $0.044 |
| Input (cache miss) | $0.15 | $0.30 | $0.66 | $1.32 |
| Output | $0.60 | $1.20 | $1.98 | $3.96 |
Both models share a 1M-token context window. Off-peak covers all hours except 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday — so most of your week runs at the cheaper rate. For a deeper dive on the discount window, see our off-peak pricing explainer.
What you give up with V4 Pro, compared to Flash: vision input is not supported, and the concurrency limit is 500 versus Flash's 2,500. If you need a head-to-head breakdown of capability differences, we covered that in V4.1 Flash vs V4 Pro.
The quiet footnote: legacy Flash names got rerouted
The same docs update contains a second detail worth knowing. The legacy model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted — but the models behind them have been retired, and requests are silently served by DeepSeek-V4.1-Flash and billed at the Flash price.
So the one-way street DeepSeek originally planned for V4 Pro did happen for older Flash models: old name in, new model out, new (cheaper) price on the bill. If you're still sending deepseek-v4-flash in production, your invoices already reflect V4.1 Flash rates.
Should you change anything?
If you're on V4 Pro: nothing to do. The official message is explicitly "billing method remaining unchanged," with a promise to give notice before any future change. High-concurrency production workloads (500 max concurrent connections) and text-only pipelines are exactly what this model is for.
If you're price-sensitive: V4 Pro still costs roughly 4–5× Flash at peak rates ($1.32 vs $0.30 per 1M cache-miss input; $3.96 vs $1.20 output). For most workloads Flash is the rational default — our cheapest DeepSeek API providers roundup covers the budget angle across hosts.
This situation has already flipped once — the retirement was scheduled, then canceled on the deadline day. DeepSeek's own footnote says they "will provide further notice should there be any changes." All figures here were verified against the official pricing page on September 15, 2026; check the source before making procurement decisions based on them.