Key Takeaways:
- DeepSeek's flagship V4 Pro left preview Aug. 12 with agent benchmarks that narrow the gap with Anthropic's Claude Fable 5 to single digits — at roughly one-forty-sixth the price.
Key Takeaways:

DeepSeek's flagship V4 Pro left preview Aug. 12 with agent benchmarks that narrow the gap with Anthropic's Claude Fable 5 to single digits — at roughly one-forty-sixth the price.
DeepSeek's flagship V4 Pro exited preview Aug. 12 with agent-benchmark scores that trail Anthropic's Claude Fable 5 by an average 5.3 percent while costing about 46 times less on blended input-output pricing.
"This closes the inference cost gap," Clement Delangue, chief executive of Hugging Face, said, citing per-task costs that could widen to about $31 for Claude Fable 5 versus $0.04 for DeepSeek.
The 0813 build, served through the deepseek-v4-pro endpoint without a code change, keeps preview pricing at $0.435 per million input tokens and $0.87 per million output, with a one-million-token context window and 384,000-token output ceiling. On DeepSeek's own agent suites, the official build jumped to 62.7 on DeepSWE from 12.8 in preview, and to 67.2 on DSBench-Hard from 31.1.
The release pressures premium-priced rivals — Claude Fable 5 lists at $10 per million input and $50 per million output — and shifts the buying calculus for enterprises weighing top performance against cost per task.
DeepSeek's published results show the official build posting 87.9 on Terminal Bench 2.1, ahead of Claude Opus 4.8's 85.0 and a fraction behind Moonshot's Kimi K3 at 88.3. On CyberGym, a cybersecurity offensive-defensive suite, V4 Pro scored 83.3, edging Fable 5's 83.1. With tools enabled on Humanity's Last Exam, it reached 60.0 against Fable 5's 63.0.
The gains concentrate in long-horizon agent work. DeepSWE, DeepSeek's internal software-engineering benchmark, rose nearly fivefold from the preview's 12.8 to 62.7. DSBench-FullStack climbed to 71.1 from 41.8, and AutomationBench to 31.8 from 12.8.
The numbers are vendor-reported. DeepSeek has not disclosed the test infrastructure behind the runs, and two of the ten benchmarks — DSBench-FullStack and DSBench-Hard — are internal sets with no external leaderboard. The Hugging Face model card still lists the V4 series as preview, and the 0813 weights have not been published. Independent verification is pending. Third-party scoring is more muted: OpenRouter, citing Artificial Analysis's Composite Intelligence Index, rated V4 Pro at 45.3, placing it "better than 70 percent of models."
The price differential is the story's center of gravity. Claude Fable 5 costs $10 per million input tokens and $50 per million output, a blended rate near $30. DeepSeek's blended rate is about $0.65 — roughly 46 times lower. Measured per benchmark task, Artificial Analysis put V4-Flash at about 3 cents versus $3.15 for Claude Fable 5.
OpenRouter data shows a cache hit rate of 86 percent, pushing the weighted input price users actually pay to $0.064 per million tokens. The 1.6-trillion-parameter mixture-of-experts model activates 49 billion parameters per token, with DeepSeek's Compressed Sparse Attention and Heavily Compressed Attention cutting single-token inference compute to 27 percent and KV cache to 10 percent of the V3.2 generation at the million-token setting.
The low price may not hold. DeepSeek posted a notice on its pricing page that it plans "a significant increase" in overall API pricing, with peak-hour rates for V4 Pro doubling to 6 yuan per million uncached input and 12 yuan per million output during 9:00-12:00 and 14:00-18:00 Beijing time. The company also promised price cuts after the bulk rollout of Ascend 950 super nodes in the second half of the year.
Tencent has already moved: CodeBuddy and WorkBuddy — IDE, plugins, and CLI — upgraded to the official V4 Pro build, a sign of enterprise adoption in China. Anthropic faces the sharper pressure, with its internal Claude Opus 5 reportedly outperforming Fable 5 on multiple benchmarks at about half the price, complicating its own premium positioning.
For developers, the calculus shifts from raw benchmark supremacy to cost per completed task. DeepSeek's MIT-licensed weights, downloadable since April with more than 1.4 million downloads on Hugging Face in the past month, give teams a path to self-hosted inference that closed-source rivals cannot match. The open question is whether the 0813 build's performance holds up under independent testing — and whether the price increase lands before that verification arrives.
This article is for informational purposes only and does not constitute investment advice.