Key Takeaways:
- Cerebras CS-4 runs inference up to 30x faster than GPU rivals
- WSE-3T packs 4 trillion transistors, 250 PFLOPs per wafer
- Launch follows Q2 selloff; 2027 core revenue tripling guidance stands
Key Takeaways:

Cerebras Systems unveiled the CS-4, a rack-scale system it says runs AI inference up to 30 times faster than GPU rivals, escalating its challenge to Nvidia's hold on fast-token serving. The launch, at the company's SUPERNOVA event Tuesday, arrives a week after a 13-17 percent selloff tied to confusing Q2 accounting, giving investors a concrete product to weigh against guidance to triple core revenue in 2027.
"Being 30 times faster doesn't just make a response feel fast. It gives an agentic system room for more than an order of magnitude as much reasoning, verification, or tool use in the same wall-clock time," Sean Lie, CTO and co-founder of Cerebras, said.
The CS-4 is built from three new WSE-3T wafers and delivers 750 PFLOPs of AI compute, 129.6 petabytes per second of memory bandwidth, and 7.2 terabits per second of I/O. It is up to twice as fast as the prior CS-3 and delivers up to 10x more throughput per watt, according to the company. In a head-to-head test on the GPT-OSS-120B model, the CS-4 produced more than 4,400 tokens per second per user, which Cerebras said is up to 30 times faster than GPU solutions.
The WSE-3T, the largest AI processor ever built, packs four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimeters of silicon, with 44GB of SRAM integrated on the wafer. It doubles AI compute to 250 PFLOPs per wafer and memory bandwidth to 43.2 petabytes per second versus the WSE-3. A new Direct Wafer Links mode connects wafers within and across racks without a switch, cutting wafer-to-wafer latency to as low as two microseconds and supporting models with more than 50 trillion parameters.
The Specs Behind the Speed
The CS-4 is the first product built on Cerebras' new Nexus platform architecture, which separates compute, power, and I/O into modular elements. A rear-mounted "backpack" folds power conversion, direct liquid cooling, and high-speed I/O into a package built around the wafer, cutting deployment time from days to hours and reducing components by 50 percent. Power conversion now sits about 0.5 millimeters from the processor, roughly 100 times closer than on conventional GPU boards, which Cerebras said nearly eliminates board-level power loss and lets the WSE-3T run at higher frequencies.
The system also supports standards-based RoCE v2 RDMA over Ethernet, allowing it to plug into existing infrastructure and pair with heterogeneous systems. Cerebras said the programmable I/O subsystem will enable disaggregated inference setups with partners including AMD's Helios and AWS Trainium, where a separate prefill engine hands prompts to Cerebras for low-latency decode.
"CS-4 makes dramatic improvements in system deployability, reliability, and networking, which enables scaling performance to larger models for large scale token factories," Dylan Patel, founder and CEO of SemiAnalysis, said.
What It Means for the Cloud Race
The launch is the centerpiece of Cerebras' pivot from selling hardware to renting inference capacity, a business that grew 287 percent year over year to $127.7 million in core cloud revenue in Q2. The company ended June with $25.4 billion in remaining performance obligations, anchored by a December 2025 Master Relationship Agreement with OpenAI valued at more than $10 billion, and has guided to triple core revenue in 2027.
Cerebras' wafer-scale design avoids three of the most supply-constrained components in the AI chip market — high-bandwidth memory, CoWoS advanced packaging, and 3-nanometer fabrication — a structural advantage when Foxconn's rotating CEO has cited CoWoS capacity as the ceiling on AI server production. The company said manufacturing scales more than tenfold during 2026 through contract manufacturers Flex, Sanmina, and Rocket EMS, with TSMC wafer supply secured.
The competitive stakes are clear. Nvidia dominates AI training and has been pushing into fast inference, while AMD and AWS are building their own accelerators. Cerebras claims its CS-4 advantage in tokens-per-second-per-user over GPUs reaches 30x, a gap that, if confirmed by independent benchmarks, could shift procurement spend among hyperscalers and AI labs. Cerebras did not disclose the test conditions for the GPU comparison beyond the GPT-OSS-120B model.
Cerebras shares, which settled near $219-225 after the Q2 selloff, trade at roughly 50 times trailing sales — a multiple that compresses to about 16 times if the company delivers on its 2027 revenue guidance. The CS-4's pricing and shipping timeline were not disclosed; the company said the system is the first member of the Nexus platform and that shipments are expected to begin in the coming quarters.
This article is for informational purposes only and does not constitute investment advice.