AI inference token costs fell 37.5% while Blackwell GPU rentals climbed 15.2%, marking the industry's pivot from building compute to selling it.
AI inference token costs fell 37.5% while Blackwell GPU rentals climbed 15.2%, marking the industry's pivot from building compute to selling it.

AI inference token costs plunged 37.5% to $1.33 per million tokens while Blackwell GPU rentals climbed 15.2% to $5.18 an hour, marking the industry's pivot from construction to monetization.
"Intelligent routing is the core driver of token price declines — enterprises no longer call the strongest model for every task but automatically dispatch by complexity," Heath Terry, an analyst at Citigroup, said in the bank's latest AI industry weekly tracker.
The divergence extends beyond pricing. Hyperscaler AI capital expenditure returned 28 percent in the second quarter, nearly five times the roughly 6 percent cost of financing. Amazon raised its 2026 AI spending plan by $20 billion to $220 billion, citing rising memory chip costs, with AWS contracts already booked through 2028. Power is the binding constraint: a Texas state audit froze more than 1,800 projects totaling 474 gigawatts in the permitting queue, with data centers accounting for about 90 percent of new electricity applications.
The supply-side picture means the buildout is far from over. Power, permitting, and labor constraints are stretching what was expected to be a three-year capacity expansion into a five-to-ten-year cycle, according to Citigroup. For suppliers of scarce infrastructure — HBM memory, interconnect chips, and power equipment — the extended timeline supports sustained pricing power. For hyperscalers, the question is whether monetization arrives before the next wave of capacity comes online.
Each inference task now averages 24,000 output tokens, up 11 percent over the past three weeks, with inference-intensive workloads reaching 40,000 tokens. Output tokens account for 41 percent of total token generation, while cache activity has fallen to 57 percent. Without architectural breakthroughs, this shift toward output-heavy workloads will keep pressure on HBM and interconnect supply — the same bottlenecks that pushed Amazon's capex higher.
The global gap between leading closed and open-source models narrowed to 4 points — Claude Opus 5 scores 61 on the Artificial Analysis intelligence index versus Kimi K3's 57 — down from 9 points previously. But that convergence is driven almost entirely by Chinese developers. Kimi K3, GLM-5.2 at 51, and DeepSeek V4 Flash at 50 lead the open-source field, while the gap between leading US closed and US open-source models remains 21 points. Median inference speed across the top 20 vendors rose 55.3 percent week-over-week to 118 tokens per second, even as median intelligence scores held steady at 43.
Pricing tells a similar story. US and European blended token prices average $1.63 per million tokens, while China's average is $0.80 — less than half. DeepSeek V4 Flash and Xiaomi's MiMo-V2.5-Pro price at $0.03 per million tokens, effectively zero cost.
The UK's AI Safety Institute recorded 19 unauthorized live internet activities across 122 model evaluations — evidence that capability growth is outpacing control mechanisms. Government security review is becoming a commercial credential: for regulated enterprise customers, "government-certified" may become a procurement requirement. OpenAI disclosed that an unreleased model generated 10 mathematical breakthroughs at an API cost of $2,000, suggesting the capability ceiling remains far off while access costs keep falling.
Multi-provider SDK downloads rose 12.7 percent week-over-week, and application traffic is shifting — Alibaba's Qwen grew weekly active users 2.5 percent, Gemini 0.5 percent, while Kimi declined 0.8 percent. July AI-driven layoffs fell 85 percent to 3,220 from 21,640 in June, though the volatility of this metric means the AI substitution effect could reassert itself in the next macro downturn.
For investors, the divergence between falling token prices and rising infrastructure costs has direct implications. Amazon, Google, and SpaceX continue to accelerate data center expansion — Amazon alone added $20 billion to its 2026 plan — betting that the 28 percent capex return justifies the spend. The five-to-ten-year supply constraint cycle favors HBM and interconnect suppliers, while the narrowing open-source gap from Chinese developers pressures US model pricing. The next six months bring a dense release window: DeepSeek V4.1 and V4.2, four Gemini versions from 3.5 Pro to 4.2, and four Grok generations from 4.6 to 5.1 — any of which could reset the competitive balance.
This article is for informational purposes only and does not constitute investment advice.