DeepSeek V4 Pro Retirement Canceled — It Stays Live Past Sept 14

·

Short answer

DeepSeek V4 Pro is not retiring. On September 14, 2026 — the day it was supposed to be phased out — official docs confirmed the model continues to be served, billing unchanged. If you use deepseek-v4-pro, do nothing.

The retirement is officially called off

DeepSeek's original plan, stated in their docs since the V4.1 Flash launch, was blunt: API service for V4 Pro would end after September 14, 2026. Multiple outlets (including our own coverage last week) reported the retirement as scheduled.

That changed. The current official pricing page now reads:

"In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. We will provide further notice should there be any changes." — DeepSeek API Docs, Models & Pricing, retrieved September 15, 2026

No new pricing tier, no migration step, no deprecation shim. The model name deepseek-v4-pro (currently serving DeepSeek-V4-Pro-0813) keeps working exactly as before.

Our earlier report — V4.1 Flash Launches — and V4 Pro Retires on Sept 14 — accurately reflected the official plan at the time. That plan has since been reversed. This post supersedes it on the retirement question; the Flash launch details there still stand.

V4 Pro pricing, verified today

Numbers below are straight from the official pricing page as of September 15, 2026 — note that some third-party pricing roundups are already out of date on these figures.

Per 1M tokens Flash off-peak Flash peak V4 Pro off-peak V4 Pro peak
Input (cache hit) $0.003 $0.006 $0.022 $0.044
Input (cache miss) $0.15 $0.30 $0.66 $1.32
Output $0.60 $1.20 $1.98 $3.96

Both models share a 1M-token context window. Off-peak covers all hours except 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday — so most of your week runs at the cheaper rate. For a deeper dive on the discount window, see our off-peak pricing explainer.

What you give up with V4 Pro, compared to Flash: vision input is not supported, and the concurrency limit is 500 versus Flash's 2,500. If you need a head-to-head breakdown of capability differences, we covered that in V4.1 Flash vs V4 Pro.

The quiet footnote: legacy Flash names got rerouted

The same docs update contains a second detail worth knowing. The legacy model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted — but the models behind them have been retired, and requests are silently served by DeepSeek-V4.1-Flash and billed at the Flash price.

So the one-way street DeepSeek originally planned for V4 Pro did happen for older Flash models: old name in, new model out, new (cheaper) price on the bill. If you're still sending deepseek-v4-flash in production, your invoices already reflect V4.1 Flash rates.

Should you change anything?

If you're on V4 Pro: nothing to do. The official message is explicitly "billing method remaining unchanged," with a promise to give notice before any future change. High-concurrency production workloads (500 max concurrent connections) and text-only pipelines are exactly what this model is for.

If you're price-sensitive: V4 Pro still costs roughly 4–5× Flash at peak rates ($1.32 vs $0.30 per 1M cache-miss input; $3.96 vs $1.20 output). For most workloads Flash is the rational default — our cheapest DeepSeek API providers roundup covers the budget angle across hosts.

This situation has already flipped once — the retirement was scheduled, then canceled on the deadline day. DeepSeek's own footnote says they "will provide further notice should there be any changes." All figures here were verified against the official pricing page on September 15, 2026; check the source before making procurement decisions based on them.