DeepSeek's V4 API prices rose as much as 1,100% on August 17, the sharpest move in a wave of hikes by independent Chinese AI vendors that analysts say corrects years of subsidized inference pricing.
DeepSeek's V4 API price increases of up to 1,100% took effect August 17, the sharpest in a wave of hikes by independent Chinese AI vendors that analysts say corrects years of subsidized inference pricing. The company adopted peak and off-peak rates for its V4-Pro and V4-Flash models, with off-peak prices set at half of peak-hour levels.
"This round of repricing is not a uniform industry increase but independent model vendors leading, while large tech firms with proprietary compute hold steady or adjust indirectly," Liang Wei, partner at Grant Thornton, said. Zhipu, Moonshot's Kimi, and MiniMax have all raised API prices in recent months, while Alibaba, ByteDance, and Tencent have held rates.
DeepSeek's V4-Pro peak-hour cache-hit input rose to 0.3 yuan per million tokens from 0.025 yuan, a 1,100% jump, with output at 27 yuan from 6 yuan. Off-peak V4-Pro rates are 0.15 yuan for cache-hit input, 4.5 yuan for cache-miss input, and 13.5 yuan for output. Peak hours run 9:00-12:00 and 14:00-18:00 Beijing time, seven hours daily, with the remaining 17 hours off-peak. V4-Flash peak prices are 0.1 yuan, 3 yuan, and 9 yuan per million tokens for cache-hit input, cache-miss input, and output respectively.
The repricing marks a shift from the low-cost token dumping that defined China's model market, as frontier models converge on capability and competition moves to application-level economics. DeepSeek's V4 family had been the benchmark for low-cost Chinese model pricing, and its decision to raise API prices signals the end of sustained discounting.
The 1,100% Jump and What It Buys
DeepSeek announced the new prices August 13, the same day it released the official V4-Pro build 0813, which supports thinking and non-thinking modes, a 1 million token context window, and up to 384,000 token output. The model's DeepSWE score, an AI programming agent benchmark, jumped from 7.3 for the preview to 62.7, surpassing Claude Opus 4.8. On the AI security agent benchmark Cybergym and the agent benchmark AutomationBench, it beat Claude Fable 5, widely considered the strongest current model.
The price increase coincides with the open-sourcing of DeepSeek Harness and the announcement that Kingsoft's WPS Lingxi Pro would be among the first products to integrate V4-Pro. DeepSeek is repositioning from a pure model seller toward a company monetizing complete task delivery, where token pricing is one lever among several.
The Cost Gap Narrows Against Qwen
The pricing change reshapes the competitive comparison with Alibaba's Qwen 3.8 Max. Before August 16, DeepSeek V4 Pro was roughly 7.4 times cheaper than Qwen on a 12-task tool-use suite measured by Composio. After the cutover, that gap fell to 4.8 times off-peak and 2.4 times at peak, according to pricing applied to Composio's measured token counts. Qwen 3.8 Max scores 58 on the Artificial Analysis Intelligence Index versus DeepSeek's 53 and Kimi K3's 60.
The geography of DeepSeek's peak windows matters. Peak hours align with the Asian business day, so a developer in Singapore pays peak rates on nearly every call, while a developer in San Francisco pays off-peak on all of them. Alibaba's Qwen 3.8 Max open weights, released August 12, shipped text-only without vision or the 1M-token context window, under a custom license with revenue-share terms, a departure from the Apache 2.0 license of earlier generations. DeepSeek's V4 Pro ships its full model under MIT at 893 GB, ungated.
Independent testing also raises questions about real-world agent performance. Composio ran DeepSeek V4 Flash through four agent harnesses on 30 multi-step tasks spanning Gmail, GitHub, Slack, and Google Sheets; only 129 of 240 runs passed, and just six of 30 workflows completed across all harnesses. Success rates ranged from 14 of 30 on OpenCode to 17 of 30 on Oh My Pi, with cost per successful task between $0.073 and $0.195.
"Companies can consider high-performance inference engines as solutions for specific workloads, rather than completely replacing existing products," Carmi Levy, a technology analyst, said. Sanchit Vir Gogia of Greyhound Research added that what matters is whether a model's performance meets actual enterprise workflows, not whether it tops every benchmark.
For investors, the repricing improves monetization prospects for independent Chinese AI vendors whose inference costs had been running below sustainable levels. DeepSeek's move from flat pricing to time-of-use rates resembles peak-valley electricity tariffs, a model that raises utilization of compute clusters while easing peak-hour pressure. The question now is whether the higher prices hold as Alibaba, ByteDance, and Tencent keep their own API rates flat, and whether DeepSeek's benchmark gains translate into enterprise adoption at the new price points.
This article is for informational purposes only and does not constitute investment advice.