OpenAI's custom Jalapeño accelerator delivers up to 1.9x more inference work per watt than Nvidia's GB300, threatening the chipmaker's hold on AI data centers.
OpenAI's custom Jalapeño accelerator, built with Broadcom, delivers 1.5x to 1.9x more inference work per watt than Nvidia's GB300 while cutting end-to-end latency by up to 3.6x, threatening the chipmaker's grip on AI data centers.
"Jalapeño offers the best of both worlds," Richard Ho, vice president of hardware at OpenAI, said, noting AI systems typically must trade off between latency and throughput. The chip is an application-specific integrated circuit, or ASIC, designed solely for inference — the process of running a trained model to complete a task — rather than the mix of training and inference that general-purpose GPUs handle.
Each rack packs 128 accelerators with 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4 and nearly 2 petabytes per second of memory bandwidth. Each chip delivers 13.4 petaFLOPS at MXFP4, fed by 216 GB of HBM4. OpenAI plans to deploy Jalapeño in small volumes by year-end and ramp into 2027, a timeline that could pressure Nvidia as its largest customer builds in-house silicon.
The benchmark results, run on SemiAnalysis' InferenceX suite across GPT-OSS-120B, DeepSeek R1 and Kimi K2.5, show Jalapeño delivering 1.7x to 3.6x lower end-to-end latency than the best recorded results on Nvidia's GB200 NVL72 and GB300 NVL72 racks. For ultra-low-latency inference — the segment AI infrastructure providers now chase — OpenAI says the chip is 2.1x to 4.1x faster. The comparison excluded speculative decoding, a technique that uses a small draft model to predict outputs, which OpenAI argued makes for a cleaner apples-to-apples test.
Memory bandwidth is the deciding factor in inference, and here Jalapeño's rack design leans on a large on-chip SRAM cache to keep model state and key-value caches local, minimizing data movement between compute and memory. "We designed Jalapeño to minimize data movement and communication delays," the company said in a blog post. Unlike Nvidia's Groq LPX racks, which are tuned for a single phase, Jalapeño is optimized for both the compute-heavy prefill stage and the memory-bandwidth-intensive decode stage.
OpenAI also used its own models to design, architect and optimize the chip, cutting the time from inception to tape-out to nine months, including writing custom kernels as new models shipped. AMD and Nvidia have pursued similar internal automation — AMD opened its ROCm.AI tooling to the public last month — but OpenAI's nine-month cycle highlights how quickly the AI lab can iterate on silicon.
The competitive math is stark. AMD's MI455X and Nvidia's Rubin GPUs, both expected to ramp in early 2027, are general-purpose parts optimized for training and inference; Jalapeño only needs to excel at inference. OpenAI estimates each Jalapeño rack will draw 40 percent to 60 percent of the power of competing GPU systems, though it has not disclosed system-level consumption. AMD's Helios racks deliver 1.46x to 2x more compute and up to 12 percent more memory than OpenAI's rack, but only 85 percent of its memory bandwidth.
OpenAI stressed it will not abandon Nvidia or AMD, which remain among its most important investors and suppliers for training compute. Still, the shift toward custom silicon — following Google's TPU and Amazon's Trainium — gives OpenAI a lever to cut inference costs as it scales. Nvidia shares, which have priced in years of data center growth, face a longer-term question of whether its largest customer becomes a competitor. OpenAI said it will continue developing second- and third-generation Jalapeño chips, with volume production slated for 2027.
This article is for informational purposes only and does not constitute investment advice.