DeepSeek posted a platform notice that it plans to officially release V4.1 Flash around September 10, 2026 (Beijing Time). The same banner says that after V4.1 Flash launches — and before V4.1 Pro ships — requests to the Pro model will be routed to V4.1 Flash and billed at Flash pricing.

DeepSeek platform notice: V4.1 Flash around September 10, 2026; Pro traffic routes to Flash at Flash price

DeepSeek also claims extensive internal and external testing shows V4.1 Flash has comprehensively surpassed V4 Pro on performance, cost, speed, and task completion time. That remains a vendor claim until independent runs catch up. The notice asks developers to send feedback from comparative testing between V4 Pro and V4.1 Flash.

That routing promise matters against the current list prices. DeepSeek’s API docs still price production deepseek-v4-flash and deepseek-v4-pro on a peak / off-peak schedule (USD per 1M tokens). Off-peak rates are half of peak; peak windows are 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday. The table below is the off-peak slice — the gap the Pro→Flash remap would close on paper:

Model Concurrency Input · cache hit Input · cache miss Output
V4.1 Flash (post-launch billing per platform notice) $0.007 $0.22 $0.66
deepseek-v4-flash (V4-Flash-0731) 2,500 $0.007 $0.22 $0.66
deepseek-v4-flash-vision-exp 2,500 $0.007 $0.22 $0.66
deepseek-v4-pro (V4-Pro-0813) 500 $0.022 $0.66 $1.98

V4.1 Flash is not yet a separate priced row on the public pricing page; the first row follows DeepSeek’s banner (and the expired beta’s Flash-matched rates). All listed models share 1M context and up to 384K max output on the docs page. Product prices can change — check DeepSeek’s pricing page for the live schedule.

Earlier Brocker coverage of the V4 Flash family sits at DeepSeek V4 Flash and the experimental vision drop at V4-Flash-Vision-Exp.

Confirmed

  • Platform banner (observed September 9, 2026): official V4.1 Flash release targeted around September 10, 2026 (Beijing Time).
  • Same notice: after V4.1 Flash launches and before V4.1 Pro, Pro-model requests route to V4.1 Flash and bill at Flash’s price.
  • Same notice: DeepSeek says internal and external testing shows V4.1 Flash surpassed V4 Pro on performance, cost, speed, and task completion time.
  • Same notice: invites feedback from comparative testing between V4 Pro and V4.1 Flash.
  • DeepSeek API docs (Models & Pricing): off-peak Flash input $0.007 cache hit / $0.22 cache miss / $0.66 output; Pro $0.022 / $0.66 / $1.98; peak = 2×; concurrency 2,500 (Flash / Vision Exp) and 500 (Pro).

Unknown

  • Independent replication that V4.1 Flash beats V4 Pro on the metrics DeepSeek lists.
  • Exact UTC cutover, region coverage, concurrency/SLA when Pro traffic is remapped to production V4.1 Flash.
  • V4.1 Pro ship date beyond “prior to the release of V4.1 Pro.”
  • Whether short-lived preview model IDs (for example expires-on-0910 builds reported elsewhere) map 1:1 to the production V4.1 Flash name after launch.
  • Whether native multimodal / architecture claims circulating in secondary coverage are part of the official production drop.

Our take

The story worth covering is not another “Flash beats Pro” slogan — it is DeepSeek temporarily collapsing Pro traffic onto Flash pricing. For agent and batch buyers, that is a bill-shape change: same Pro endpoint name may no longer mean Pro unit economics. Treat the quality claim as marketing until third-party evals land; treat the routing policy as the operational hook.

Sources