Key Takeaways: Nvidia has quietly revived a chip program the market had written off, and the redesign targets the most expensive phase of AI inference.
Key Takeaways: Nvidia has quietly revived a chip program the market had written off, and the redesign targets the most expensive phase of AI inference.

Nvidia has revived its Rubin CPX prefill accelerator, targeting the cost bottleneck of long-context AI inference, with production slated for the first quarter of 2027.
"Just as the market had come to believe that Rubin CPX had been dropped from Nvidia's product roadmap, my latest industry checks indicate that Nvidia has revived the program," Ming-Chi Kuo, a supply-chain analyst, said Monday in a post on X.
The redesigned chip carries 168GB of HBM4 memory per GPU, up from 128GB of GDDR7 in the earlier design but below the 288GB HBM4 in standard Rubin GPUs. Each 8-card compute tray holds roughly 1.34TB of HBM4, enough to cover most long-context prefill workloads and their KV Cache requirements. The chip draws up to 2,300 watts, matching standard Rubin power draw, and its compute performance approaches that of the standard Rubin GPU.
The restart matters because prefill — the phase where a model reads and processes input before generating output — now accounts for more than 50 percent of AI inference workloads as context windows expand. Nvidia recommends deploying CPX at a 1:1 ratio with Vera Rubin NVL72 systems, with CPX handling prefill and generating the KV Cache before transferring it over Ethernet RDMA to Rubin for the decode stage.
The new CPX moves into a standalone MGX ETL rack rather than sharing space with Rubin GPUs, allowing customers to scale prefill capacity independently. Configurations range from 64 to 256 CPX GPUs, with each 64-GPU module comprising eight compute trays of eight GPUs plus a switch tray.
Interconnect follows a layered design that trades peak bandwidth for cost. Within each tray, eight CPX GPUs connect via NVLink at 1 to 1.5TB/s per GPU — well below the 3.6TB/s on standard Rubin. Between trays and rack modules, Nvidia uses Spectrum-6 Ethernet with all-copper L1 links within a rack and OSFP fiber across modules.
The architecture reflects a fundamental split in how AI inference works. Prefill is compute-bound: the GPU processes the entire input in parallel, generating the KV Cache that stores attention states. Decode is memory-bandwidth-bound: the model generates output one token at a time, repeatedly loading weights and cached context. Running both on the same GPU forces a compromise — paying for bandwidth you don't need during prefill and compute you don't need during decode.
Nvidia's recommendation of a 1:1 CPX-to-Rubin ratio means customers deploying Vera Rubin NVL72 systems will need roughly double the GPU count for a given inference workload. That math supports continued data center capex growth as long-context models expand.
The company's second-quarter results, reported in August, showed $96.22 billion in revenue, up 106 percent year over year and above the $92.18 billion consensus. Vera Rubin is ramping into full production with CoreWeave, Nebius, Microsoft Azure, Google Cloud and Oracle Cloud among its partners.
Nvidia shares closed at $220.50 on Monday, up 1.36 percent. The stock ranks in the 98th percentile for growth on Benzinga Edge Rankings with positive price-trend ratings across short-, medium- and long-term time frames.
The CPX restart also carries implications for TSMC, which manufactures Nvidia's accelerators. The shift from GDDR7 to HBM4 increases demand for high-bandwidth memory and advanced packaging, potentially tightening supply of both. And the decision to highlight Groq 3 LPUs and LPX racks at GTC 2026 — where CPX was notably absent from the roadmap — suggests Nvidia is hedging across multiple inference architectures even as it doubles down on prefill specialization.
For investors, the key question is whether the 1:1 deployment ratio translates into higher average selling prices per rack or simply more units sold. Nvidia has not disclosed CPX pricing, but the chip's lower memory capacity and Ethernet-based scale-out suggest a lower cost per GPU than standard Rubin — potentially expanding the addressable market for inference-heavy workloads.
This article is for informational purposes only and does not constitute investment advice.