DeepSeek’s ultra-low API era now has a date and a rate card. On 6 August 2026 the Hangzhou lab warned developers of a “significant” price increase without saying how much or when. On 13 August, alongside V4-Pro coverage, DeepSeek published the missing numbers: from 16:00 UTC on 16 August 2026, V4-Flash and V4-Pro move to peak / off-peak billing — and both tiers sit above today’s flat promotional list.

That closes the open question in Brocker’s earlier report. The strategy shift is no longer a soft notice; it is a scheduled reset on the official pricing page, and the multiples match what Reuters, Fortune, and Engadget summarized as roughly fourfold on peak output for the V4 line.

Confirmed

  • Effective: 16:00 UTC, 16 August 2026 (per DeepSeek API pricing docs).
  • Structure: Peak hours 01:00–04:00 and 06:00–10:00 UTC; all other hours off-peak. Off-peak rates are half of peak.
  • Old flat list (still billed until the cutover, per docs): V4-Flash cache-miss input / output $0.14 / $0.28; V4-Pro $0.435 / $0.87 per million tokens. Cache-hit input was far lower ($0.0028 Flash / $0.003625 Pro).
  • New V4-Flash (per 1M tokens): off-peak cache hit / miss / output $0.007 / $0.22 / $0.66; peak $0.014 / $0.44 / $1.32.
  • New V4-Pro: off-peak $0.022 / $0.66 / $1.98; peak $0.044 / $1.32 / $3.96.
  • Versus the old flat output rates, peak output is about 4.7× on Flash ($0.28 → $1.32) and about 4.6× on Pro ($0.87 → $3.96). Cache-hit input rises by a larger multiple (especially Pro), which hits agent workloads that reuse long prefixes.
  • Independent press (Reuters via CNA; Fortune; Engadget) frames the move as DeepSeek ending the rock-bottom promo era while remaining cheaper than many Western flagships on headline output — with the caveat that OpenAI’s GPT-5.6 Luna (~$1.20/M output in Engadget’s comparison) can undercut Flash at peak.

Analysis

August 6 was the warning; August 13–16 is the invoice. DeepSeek is not abandoning cheap inference so much as repricing the floor and teaching customers to schedule around UTC peak windows. Off-peak Flash/Pro still look aggressive against Opus/Sol-class lists; peak Flash output competing with Luna is the uncomfortable row for teams that assumed “DeepSeek = always cheapest.”

The larger story matches the earlier Brocker thesis: ultra-low rates bought share and strained capacity; revenue and load-shaping now matter more than shocking the West on a price chart. Whether reliability and queue behavior improve after 16 August will decide if this is a mature commercial pivot or a self-inflicted migration event toward Luna, Qwen, GLM, and self-hosted open weights.

Unknown

  • How much traffic shifts to off-peak windows vs churns to other providers in the first weeks after cutover.
  • Whether DeepSeek keeps a quieter promo channel (credits, enterprise deals) that softens the public list.
  • How the new rates interact with third-party hosts that resell DeepSeek-compatible endpoints.

Our take

Treat 16 August as the real end of “DeepSeek as free-ish compute,” not as DeepSeek leaving the cheap tier entirely. Bake budgets on peak numbers if your agents run during UTC morning load; use off-peak only when you can actually schedule it. The Aug 6 piece asked how big the hike would be — the answer is now on DeepSeek’s own pricing table, and it is large enough that price-performance routing finally matters again.

FAQ

When do new prices start?
16:00 UTC on 16 August 2026, per DeepSeek’s pricing docs.

What are peak hours?
01:00–04:00 and 06:00–10:00 UTC daily; other hours are off-peak at half the peak rate.

Is it still cheaper than Claude/GPT flagships?
On typical flagship list rates, yes — especially off-peak. Against Western budget SKUs such as GPT-5.6 Luna, peak Flash is no longer an automatic win.

Sources