DeepSeek's next flagship model, DeepSeek V4, appears to be on track for an official launch on August 3, 2026, according to multiple converging signals from infrastructure partners and developer testing. The most concrete indicator comes from SiliconFlow, a major Chinese AI inference platform, which announced a tenfold increase in the cache-hit price for the DeepSeek V4 Pro model effective that date — moving from 0.1 yuan to 1.0 yuan per million tokens. Such a sharp, pre-emptive price adjustment strongly suggests the platform is preparing for the commercial availability of the official model weights, as there would be little commercial rationale for raising prices on a pre-release or test version.
At the same time, developers probing DeepSeek's API endpoints report that certain V4-related interfaces can now correctly answer questions about the release dates of competing models such as Anthropic's Claude Opus 4.5 and Sonnet 4.5 — information that previously required live internet access. These capabilities appear inconsistently across accounts, a pattern consistent with phased, gradual rollouts where only a subset of traffic is routed to newer model builds. Combined with recent service instability across DeepSeek's platform, the evidence points to an active, iterative testing cycle that has been running since mid-July.
What's New / Specs
- Target launch date: August 3, 2026 (tentative, inferred from SiliconFlow pricing change)
- SiliconFlow price adjustment: DeepSeek V4 Pro cache-hit pricing rises from 0.1 yuan to 1.0 yuan per million tokens effective August 3
- API capability expansion: Select V4 endpoints now answer factual queries about competitor model release dates (Claude Opus 4.5, Sonnet 4.5) without internet access
- Rollout pattern: Gradual, phased testing with inconsistent reproducibility across developer accounts
- Service stability: Frequent fluctuations reported since mid-July, consistent with iterative deployment and rollback cycles
- Performance expectations: Community benchmarks anticipate official V4 performance between Zhipu's GLM 5.2 and Anthropic's Claude Opus 4.8
The SiliconFlow pricing move is the strongest commercial signal to date. Cache-hit pricing typically applies to repeated or prefix-cached requests, a feature that becomes economically significant only when a model is in sustained production use. Raising this rate tenfold ahead of a public release suggests SiliconFlow expects a sharp increase in V4 Pro traffic and wants to align its margins before the official weights become widely available. The platform has not publicly confirmed that the price change is tied to the V4 launch, but the timing is difficult to explain otherwise.
On the capability side, the ability to answer specific, recent competitor release dates — information that post-dates typical training cutoffs — implies either an updated knowledge cutoff in the V4 build or the integration of a retrieval mechanism accessible through the API. The fact that not all developers can reproduce the behavior reinforces the interpretation that DeepSeek is routing a fraction of API traffic to newer model variants for live evaluation. This approach mirrors the gradual rollout strategies used by other major labs to stress-test inference stacks and monitor quality regressions before a full release.
Why It Matters
The competitive landscape DeepSeek V4 enters is markedly different from the one its predecessor faced. On the same day the SiliconFlow price change was noted, OpenAI announced an aggressive 80% price reduction across its GPT-5.6 series, dramatically improving the cost-performance ratio of its flagship models. With multimodal capabilities included at that price point, OpenAI has effectively claimed the "most cost-effective frontier model" position that DeepSeek V3 and V4 Pro previously contested in the Chinese market.
Meanwhile, Zhipu AI has moved in the opposite direction, raising subscription prices for its GLM series by two to three times in China. This divergence — OpenAI cutting prices to drive volume and lock in developers, Zhipu raising prices to extract more revenue from its installed base — creates a pincer movement on DeepSeek's positioning. If V4's performance lands between GLM 5.2 and Claude Opus 4.8 as expected, it will need to justify its pricing against a suddenly cheaper GPT-5.6 and a more expensive but domestically entrenched GLM alternative.
For enterprise buyers and application developers in China, the August 3 date represents a decision point. Many have been holding off on model migrations pending V4's official release, especially given the API improvements already visible in testing. The SiliconFlow price hike also signals that inference costs for V4 Pro will be higher than the deeply discounted rates seen during the preview period. Organizations budgeting for AI workloads in Q3 and Q4 will need to factor in both the performance gains and the revised economics.
Our Take
DeepSeek's gradual testing approach since mid-July reflects a maturity in its release engineering that was less evident in earlier launches. The combination of phased API rollouts, infrastructure partner coordination (as evidenced by SiliconFlow's pricing move), and public service fluctuations suggests a lab that is learning to manage the operational complexity of deploying frontier models at scale. This is a positive signal for reliability, even if the short-term instability frustrates developers.
However, the competitive pressure is real and immediate. OpenAI's 80% price cut on GPT-5.6 is not a marginal adjustment — it reshapes the cost calculus for any application where multimodal input, long context, or tool use matter. DeepSeek V4 will need to demonstrate not just raw reasoning parity with Claude Opus 4.8, but also superior efficiency on Chinese-language tasks, code generation, and domain-specific benchmarks to retain its domestic mindshare. The GLM price increase, paradoxically, may help DeepSeek by making Zhipu's offerings less attractive on pure cost grounds, but it also signals that the domestic premium tier is consolidating around higher price points.
The August 3 target, while strongly indicated, remains tentative until DeepSeek makes an official announcement. Labs frequently shift launch dates by days or weeks based on final safety evaluations, benchmark runs, or infrastructure readiness. Developers should treat the date as a planning horizon, not a commitment, and continue testing against current API endpoints while monitoring for the broader rollout of the capabilities already appearing in limited traffic.
FAQ
When is DeepSeek V4 officially launching?
Multiple signals point to August 3, 2026, as the target date, most notably SiliconFlow's tenfold cache-hit price increase for DeepSeek V4 Pro effective that day. DeepSeek has not yet made an official announcement, so the date should be considered tentative.
What does the SiliconFlow price hike indicate?
SiliconFlow raised the cache-hit price for DeepSeek V4 Pro from 0.1 yuan to 1.0 yuan per million tokens starting August 3. This suggests the platform expects the official model to enter production use soon, as cache-hit pricing becomes commercially relevant only at scale.
What new capabilities have developers observed in V4 APIs?
Some DeepSeek V4-related API endpoints can now correctly answer factual questions about recent competitor model releases — specifically the launch dates of Claude Opus 4.5 and Sonnet 4.5 — without requiring internet access. The behavior is inconsistent across accounts, indicating a phased rollout.
How does DeepSeek V4's expected performance compare to rivals?
Community expectations place the official V4's performance between Zhipu's GLM 5.2 and Anthropic's Claude Opus 4.8. OpenAI's GPT-5.6 series, now priced 80% lower, adds multimodal advantages that complicate direct cost-performance comparisons.
Why have DeepSeek services been unstable recently?
Frequent service fluctuations since mid-July align with the gradual testing pattern: the team is likely routing small percentages of traffic to new model builds, monitoring metrics, and rolling back when issues arise. This is standard practice for safe frontier model deployment.