Key Takeaways: OpenAI's first custom chip, built with Broadcom, beats Nvidia's Blackwell on efficiency in independent benchmarks, landing a day before Nvidia's earnings.
Key Takeaways: OpenAI's first custom chip, built with Broadcom, beats Nvidia's Blackwell on efficiency in independent benchmarks, landing a day before Nvidia's earnings.

OpenAI's first in-house inference chip, co-developed with Broadcom, delivers 1.5x to 1.9x more work per watt than Nvidia's Blackwell accelerators, a first-generation result that threatens the AI chip leader's data-center dominance.
"Jalapeño achieves high throughput and low latency simultaneously, a first in the industry," Richard Ho, vice president of hardware at OpenAI, said.
Across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5, Jalapeño posted 1.7x to 3.6x lower end-to-end latency and 2.1x to 4.1x faster interactive inference than Blackwell systems, according to SemiAnalysis's InferenceX benchmark. On DeepSeek R1, the chip exceeded 700 tokens per second per user at a concurrency of one.
The results land one day before Nvidia reports fiscal second-quarter results, with consensus revenue near $91 billion to $92 billion. Nvidia's stock has fallen after four of its last five reports despite beating estimates, and custom silicon from Google, Amazon and Meta is eroding its share.
Each rack packs 128 accelerators, 1.7 exaFLOPS of 4-bit compute, 27.5 terabytes of HBM4 memory and just under 2 petabytes per second of bandwidth. The chip, built on TSMC's N3P node, runs at a 700-watt thermal design power, compared with 900 to 1,150 watts for Nvidia's Rubin compute chip. OpenAI taped out Jalapeño in November 2025, roughly 16 months after design work began, and the B0 stepping is now in production.
The architecture strips out the complex memory systems that force GPUs to hide latency across large tensor shapes. Cores and HBM are split into slices, each with a low-latency local view, synchronized through a dedicated collective network. An out-of-order core with an L1 cache avoids the barrier latency that software-managed scratchpads impose on rivals such as Google's TPU and Amazon's Trainium.
SemiAnalysis cautioned that the comparison with Blackwell is "somewhat incomplete and unfair," because Jalapeño's true rival is Rubin, which also uses HBM4 and has begun shipping to customers. Jalapeño remains at the engineering-sample stage, with deployment planned for late 2026 and volume production in 2027. The tests excluded speculative decoding, and OpenAI did not disclose system-level power.
Jalapeño is inference-only, leaving Nvidia's training dominance untouched. But the broader shift is visible: Google's TPU 8t and 8i launched in April, Amazon's Trainium3 passed a $25 billion annual run rate, and Meta's MTIA 300 arrived in March. Broadcom, the supplier to that shift, booked over $30 billion in AI semiconductor orders in a single quarter and guided third-quarter AI revenue to $16 billion, up more than 200 percent year over year.
SemiAnalysis's Dylan Patel called Jalapeño "huge news," noting it is unusual for first-generation custom silicon to beat Blackwell at all. OpenAI wrote the chip's kernels with its Codex model, roughly 3,000 lines per kernel, and programs it through Gluon, a language built on Triton. The software stack, built from scratch in months, is where OpenAI's edge shows: it deployed DeepSeek R1, Kimi K2.5 and GPT-OSS on the chip within weeks.
Nvidia shares, trading at $213.05 and up 14.37 percent year to date, face a structural question that tonight's earnings will not answer: its biggest customers are increasingly its competitors. If Jalapeño's efficiency holds at scale, the cost advantage could shift billions in annual inference spending away from Nvidia, though the chip's 2027 volume ramp leaves the incumbent years to respond.
This article is for informational purposes only and does not constitute investment advice.