DeepSeek makes its V4-Pro 75% price cut permanent
V4-Pro is now permanently priced at $0.87 per million output tokens. DeepSeek turned a temporary discount into a strategic commitment that puts direct pressure on Anthropic and OpenAI to justify their price gap.
Summary
On May 25, DeepSeek confirmed that the 75% price cut on V4-Pro is now permanent. The model is priced at $0.435 per million cached input tokens and $0.87 per million output tokens. For comparison: Claude Opus 4.7 sits at $5/$25 per million — roughly 11 to 28 times more expensive in the same categories. It's the first time a frontier-class 1.6T MoE model with 1M context is permanently priced below the Chinese domestic floor.
In practice
For builders working with LLMs, the routing logic changed this week. DeepSeek V4-Pro becomes the natural option for agentic or batch workloads not bound by data-residency rules. Anthropic and OpenAI are in an uncomfortable position: either cut Sonnet and GPT-5.5 Mini pricing, or make a more concrete capability argument to justify the gap.
Context
The reasoning behind this decision is hardware. DeepSeek is pre-committing to a cost structure that assumes Huawei Ascend 950 supernodes arrive in volume in H2 2026 — a bet on Chinese hardware sovereignty. If the hardware arrives on time, inference costs drop further. If it doesn't, margins tighten. What is already clear: Chinese labs have compressed their release cadence from monthly to weekly, forcing teams to redesign routing logic without constant re-evals.
Why it matters
- V4-Pro is now permanently priced at ~11–28× below Claude Opus 4.7 per million tokens
- First frontier MoE model with 1M context to be permanently priced below the Chinese domestic floor
- Direct pressure on Anthropic and OpenAI to cut prices or justify the gap with capabilities
- An explicit bet on Huawei Ascend 950 hardware arriving at volume in H2 2026