Nvidia's Vera CPU, a custom Arm-based server processor built for agentic AI, achieved up to 2.2x faster performance and supported 1.6x more concurrent AI agents than competing processors in independent benchmarks, the company said Monday, as it pushes deeper into a CPU market long dominated by AMD and Intel.
"Vera is exactly the kind of hardware solution the next iteration of AI workloads demands," Nikola Borisov, co-founder and chief executive officer of DeepInfra, the cloud platform that ran the tests, said in a statement. DeepInfra processes nearly five trillion tokens per week, with about 30 percent driven by agentic systems, giving it production-scale data on CPU bottlenecks.
The Vera CPU uses Nvidia's custom Olympus microarchitecture on a monolithic die — avoiding the latency penalties of chiplet-based designs — paired with second-generation Scalable Coherent Fabric that delivers roughly three times more core-to-core bandwidth than competing server platforms. Instead of conventional DDR5 memory, Nvidia opted for data center-class LPDDR5X, providing up to 1.2 terabytes per second of memory bandwidth, three times more bandwidth per core, and about 40 percent lower memory latency under load, according to the company. A wide 10-way decode front end sustains single-threaded performance during heavy workloads.
The benchmarks, run by DeepInfra using its own production AI agent with real captured traffic, tested Vera against leading CPUs from AMD and Intel under identical conditions. Results showed Vera supporting up to 1.6 times more concurrent AI agents at the same quality of service, with collaborative speed reaching 2.2 times that of competitors. Nvidia separately claimed the Olympus core delivers roughly two times higher performance than today's x86 processors, with up to six times faster streaming data processing in workloads developed with Redpanda and HPE, and as much as seven times faster scientific computing at Los Alamos National Laboratory versus an Intel Sapphire Rapids-based supercomputer.
Why CPUs matter again for AI
Agentic AI differs from standard inference workloads. Rather than processing a single prompt and returning an answer, AI agents continuously shift between GPU inference and CPU-driven tasks such as code execution, database lookups, tool invocation and orchestration — a pattern Nvidia calls the "agent loop." CPU responsiveness directly affects how quickly an AI workflow completes, making per-core performance and memory latency more important than raw core count.
Nvidia's approach contrasts with AMD and Intel, whose server CPUs are designed for a broad mix of enterprise, cloud and high-performance computing workloads. Nvidia is optimizing an integrated AI platform where CPUs, GPUs, networking, DPUs and system software are co-engineered. The company has lined up ecosystem support from OpenAI, Thinking Machines Lab, Perplexity and Los Alamos National Laboratory, along with multiple OEMs and cloud providers planning to deploy Vera-based systems.
Investment implications
Nvidia generated $215.9 billion in revenue in its fiscal 2026, dwarfing AMD's $34.6 billion in fiscal 2025, but the Vera CPU signals an expansion of its addressable market into a segment where AMD's Epyc and Intel's Xeon have held dominant positions. Nvidia claimed Vera delivers up to 1.9 times better agentic AI performance versus AMD's Epyc Turin processor in its own benchmarks, and CoreWeave reported 10 times more tokens per watt running DeepSeek's R1 model on the Vera Rubin platform versus the prior-generation Grace Blackwell system. Independent testing across a broader range of workloads will determine whether Vera's architectural advantages translate into market share gains, but the direction is clear: Nvidia sees the CPU as a critical component in future AI systems, and it is building accordingly.
This article is for informational purposes only and does not constitute investment advice.