Nvidia's Vera Rubin deal with Hark moves the AI hardware contest from the chip to the rack.
Nvidia's Vera Rubin deal with Hark moves the AI hardware contest from the chip to the rack.

Nvidia's Vera Rubin deal with Hark moves the AI hardware contest from the chip to the rack.
Nvidia's multi-year partnership with AI startup Hark, announced Aug. 27, deploys gigawatt-scale computing on the Vera Rubin platform — the clearest sign yet that the AI hardware fight has moved from the GPU die to the full rack.
"Vera is important because there's only so much memory that you can put in a single server or any sort of compute platform," Jason Hardy, Nvidia's VP of storage technology, told TechCrunch.
Hardy said offloading orchestration work to the Vera CPU produced "upwards of 3x improvement in these operations" — an Nvidia figure that has not been independently benchmarked. Vera Rubin, now rolling out, pairs the Rubin GPU with the Vera CPU, the Groq 3 LPX inference accelerator, plus storage and networking racks sold as one integrated system.
The Hark deal lands as hyperscalers Amazon, Google and OpenAI push custom silicon — Trainium, TPU and the Broadcom-built Jalapeno — that threatens Nvidia's merchant-GPU margins. Nvidia's answer is to redefine the product boundary upward, making the die a component of a system it controls.
For two years the structural threat to Nvidia has been custom-silicon bifurcation: hyperscalers designing accelerators that win on performance per watt for their own workloads. A well-funded single customer can now tape out a chip that beats Nvidia on its own inference workload. That is the contestable die.
Hardy's counter reframes the fight. At gigawatt scale, he argues, the binding constraint is not raw compute but moving data — memory limits per server and flash bottlenecks become the real ceiling. The moment that framing is accepted, a rival ASIC stops being a replacement for Nvidia and becomes a component that plugs into Nvidia's interconnect and data-movement layer.
The corroboration that this is more than spin comes from OpenAI. Its Jalapeno inference chip, co-developed with Broadcom, has an explicit design goal of minimizing data movement and communication delays, the company said in a blog post. When both the incumbent and its largest customer point at data movement as the binding constraint, the framing reflects a real engineering reality — even if Nvidia's preferred resolution happens to favor Nvidia.
The logical endpoint of the full-stack pitch is NVLink Fusion, Nvidia's separately announced program that positions third-party silicon as XPUs plugging into Nvidia racks. That would turn hyperscaler ASICs into components of Nvidia's architecture rather than replacements for it. Whether hyperscalers accept that frame, or build their own rack-scale systems around their own ASICs, is the unsettled question.
The Hark deal adds to a pipeline that includes AMI, an AI data center firm promoted by Greenko founders, which ordered 9,000 Vera Rubin chips for a 5GW global data center network. Nvidia shares, which grew their market cap tenfold between the start of 2023 and mid-2025, have since been on a more modest trajectory as GPU competition concerns weigh. The bet embedded in Vera Rubin is that the moat now sits in rack-scale interconnect and system integration — where matching Nvidia means replicating an entire vertically integrated stack, not taping out one competitive die.
This article is for informational purposes only and does not constitute investment advice.