Google is designing a server chip that bakes its Gemini AI model into the silicon itself, a move that could cut power consumption by as much as 10x and deepen the company's break from Nvidia.
The project, internally called "Frozen v2," would etch Gemini's neural-network architecture directly into the chip's circuitry rather than loading the model from memory like conventional processors, according to a report from The Information. Alphabet shares rose as much as 3.7% on the news.
"Hardwiring the model architecture eliminates the data-shuttling overhead that consumes most of the power in today's AI chips," a person familiar with the project told The Information. Engineers can still update the model by loading new weights, but the underlying structure stays fixed.
The chip could deliver 6 to 10 times greater power efficiency per token than Google's latest TPU systems, the report said. That would put it in a different class from general-purpose GPUs and TPUs, which move data between memory and compute units — a process that wastes energy and adds latency. The design is separate from Google's TPU line and is targeted for deployment as early as 2028.
The timing reflects an internal capacity crunch at Google. The company's cloud unit has turned away some outside customers because it lacks enough AI compute to serve them, according to the report, straining internal resources and stirring tensions between business units. A chip tuned specifically for Gemini could ease that bottleneck by serving more tokens per watt.
Why hardwiring matters for the AI arms race
The approach trades flexibility for speed and efficiency. Today's AI accelerators — Nvidia's H100 and B200, Google's own TPU v5 — are general-purpose: they can run any model, but they pay a power penalty for that versatility. Frozen v2 locks to Gemini's specific architecture, dropping the overhead.
The concept is not unique to Google. Startup Taalas sells a chip it calls Hardcore that prints a model's weights and architecture directly onto silicon, claiming 17,000 tokens per second versus roughly 150 per user on a top Nvidia GPU. Taalas says its part requires no expensive high-bandwidth memory, sidestepping a supply constraint that has squeezed the industry.
For Google, the strategic logic is clear. The company already designs its own TPUs to reduce reliance on Nvidia, which dominates the AI chip market with an estimated 80% share. A Gemini-specific chip would push that self-reliance further, potentially saving billions in procurement costs. Google has also spread its chip orders across suppliers including Intel and Broadcom to avoid single-source risk.
The risks of freezing a fast-moving target
The obvious danger is obsolescence. AI model architectures evolve rapidly — what Gemini looks like today may not resemble the dominant architecture in 2028. Google's design tries to hedge by keeping weights updatable, but the fixed circuitry means a major architectural shift could render the chip less useful.
Google has not confirmed the project. A spokesperson said the company regularly explores high-efficiency ideas but cautioned that not every lab project reaches production.
Still, the signal is clear. The race for custom AI silicon is moving from running any model to fusing one model with the metal. If Google delivers Frozen v2 on schedule, it could reshape the economics of serving Gemini to billions of users — and force Nvidia to answer a question it hasn't had to face: what happens when a hyperscaler's chip is cheaper and faster than the industry standard?
Alphabet shares trade at about 23 times forward earnings. Morgan Stanley's Brian Nowak, who rates the stock overweight with a $225 price target, said in a note that in-house silicon could add $2 to $3 per share in value by 2030 if the efficiency claims hold. Nvidia shares fell 1.8% in after-hours trading on the report.
This article is for informational purposes only and does not constitute investment advice.