DeepSeek's V4-Flash costs about 3 cents per benchmark test, roughly 105 times less than Anthropic's Claude Fable 5, resetting what enterprises should pay for everyday AI.
DeepSeek's V4-Flash costs about 3 cents per benchmark test, roughly 105 times less than Anthropic's Claude Fable 5, resetting what enterprises should pay for everyday AI.

DeepSeek's V4-Flash is by far the least expensive well-known AI model to run, at an estimated 3 cents per benchmark test versus $3.15 for Anthropic's Claude Fable 5, a roughly 105-fold gap that pressures Western AI pricing.
The estimate comes from San Francisco-based research firm Artificial Analysis, which measures cost per completed task rather than posted token prices. "A model with a low headline price can still prove expensive if it requires significantly more steps to produce an answer," the firm said in its benchmark report.
DeepSeek charges $0.14 per million input tokens and $0.28 per million output tokens for V4-Flash, released into public API beta on July 31. The model scored 50 out of 100 on Artificial Analysis' Intelligence Index, matching Google's Gemini 3.6 Flash and one point behind Meta's Muse Spark 1.1 and Z.AI's GLM-5.2. Moonshot AI's Kimi K3 scored 57, while Anthropic's Claude Opus 5, Fable 5, and OpenAI's GPT-5.6 Sol each scored nine or more points higher.
The cost gap reopens the return-on-investment debate that DeepSeek's R1 model triggered in early 2025, when it caused a global technology stock selloff and raised questions about U.S. AI spending. DeepSeek, which sources say is preparing for a potential IPO, is also readying a more powerful V4-Pro model expected in August.
V4-Flash-0731 is not a bigger model than the April preview — it is the same 284-billion-parameter mixture-of-experts build, with about 13 billion parameters firing per token and a one-million-token context window. Only post-training changed. The result: DeepSWE, an agentic coding benchmark, jumped to 54.4 from 7.3, a sevenfold improvement with zero change in inference cost.
That decoupling matters for anyone budgeting an agentic stack. If a single post-training cycle can lift a budget model past its own flagship — V4-Flash-0731 outscores V4-Pro-Preview on all nine agentic benchmarks DeepSeek published — then capability and compute spend have partly separated, and a frontier lead built on model size alone is not a durable moat.
Weighted four-to-one toward input, the typical shape of an agent workload, V4-Flash blends to roughly $0.17 per million tokens. Claude Opus 4.8, by comparison, blends to about $9.00 per million tokens at list pricing. On DeepSeek's own Terminal Bench 2.1, Opus 4.8 scores 85.0 versus V4-Flash's 82.7 — 2.3 points for roughly 52 times the cost per point.
Whether that trade is worth it depends on the workload. For code review on a non-critical internal tool, the premium is hard to justify. For an autonomous agent touching production infrastructure, 2.3 points may be the whole argument.
There is a catch for open-weight users: the Hugging Face repository still hosts April's preview checkpoint under an MIT license. Build 0731 lives only behind DeepSeek's endpoint and third-party hosts such as OpenRouter. Teams that download V4-Flash today get the build that scores 7.3 on DeepSWE, not 54.4.
The pricing pressure lands on Western providers at a moment when enterprise procurement teams are comparing vendors on cost per completion. DeepSeek's 3-cent-per-test result gives them a reference point that undercuts Claude Fable 5 by about 105 times and GPT-5.6 Sol by more than 60 times. Alibaba, which unveiled its largest model, Qwen3.8-Max, on Monday, is pushing the same "good enough at lower cost" playbook. For investors, the question is whether premium AI pricing survives a market where a 50-point model costs 3 cents to run — and whether U.S. AI infrastructure spending, the pillar of the current capex cycle, still earns its keep.
This article is for informational purposes only and does not constitute investment advice.