Meta's infrastructure team has forced a custom AMD MI450X chip that cuts compute silicon in half and slashes memory by two-thirds, a decision SemiAnalysis called "disastrous" for generative AI workloads.
Meta's infrastructure team has forced a custom AMD MI450X chip that cuts compute silicon in half and slashes memory by two-thirds, a decision SemiAnalysis called "disastrous" for generative AI workloads.

Meta's insistence on customizing AMD's flagship MI450X GPU — halving its compute silicon and cutting HBM4 memory stacks from 12 to 8 — has produced a chip that its own AI research team has no interest in using, according to chip research firm SemiAnalysis.
"This chip configuration is intended for RecSys workloads and was something that RecSys infrastructure teams decided," SemiAnalysis wrote in a report published this week. "TBD has no interest in this system, nor is it attractive for external customers."
The standard MI450X uses TSMC's 2nm process with hybrid bonding, 12 stacks of HBM4 memory, and the largest CoWoS (chip-on-wafer-on-substrate) reticle size on the market. Meta's custom version downgrades to 8-Hi HBM stacks and removes half the compute silicon per package. The decision was made before Meta formed its TBD Lab, the company's flagship AI research unit, which will now favor Nvidia's Vera Rubin architecture instead.
The MI450X custom chip is the latest in a pattern of costly infrastructure decisions at Meta that SemiAnalysis traced to a deeper cultural problem: a six-month performance review cycle that cuts the bottom 10% to 15% of employees every round, encouraging short-term "window washing" projects over long-term strategy. The result has been billions in wasted spending on over-engineered hardware that Meta's own model teams never wanted.
The $2.5 Billion Rivos Misfire
Meta's most expensive single misstep was the more than $2.5 billion acquisition of chip startup Rivos last year. Few inside Meta's silicon division understood the strategic rationale, SemiAnalysis reported, and those who championed the deal have since gone quiet. The prevailing theory: Meta had the cash, the custom silicon space was heating up, and it was already licensing Rivos' IP.
The deal structure created immediate friction. Rivos' founders insisted on an all-or-nothing sale, forcing Meta to buy the entire company and then cut employees in divisions it didn't want. Existing Meta chip managers treated the incoming Rivos engineers as "free headcount," pulling the startup apart across different teams until little of the original organization remained intact.
The technology rationale evaporated just as quickly. Meta had acquired Rivos for its SIMT core IP — closer to Nvidia's GPU architecture in terms of programmability than Meta's own SIMD-based MTIA chips. But after the deal closed, Meta cancelled the Olympus chip that was meant to use Rivos' GPU IP, calling the design too aggressive in system and package architecture. A replacement project called Phoebe is scheduled to tape out in 2028, though SemiAnalysis said internal optimism is low.
Around 30% of the Rivos engineers who joined Meta were let go in recent layoffs. Co-founder Mark Hayter has already left, and multiple former Rivos employees joined Gerard Williams' new chip startup Nuvacore after their first tranche of RSUs vested in May. SemiAnalysis reported that CEO and co-founder Puneet Kumar is also eyeing an exit in a year or two.
Custom Servers, Higher Costs
Meta's custom server designs have followed a similar pattern. Its H100-based Grand Teton added an extra switch tray with four Broadcom PCIe switches, 16 SSDs and eight NICs to provide more direct-attached storage for checkpointing training runs. In production, the model teams barely used the extra storage, and the design was cancelled.
The GB200-generation Ariel server went further. Meta paired one B200 GPU with one Grace CPU instead of the standard two-GPU configuration, then used a 36x2 cross-rack setup to reach 72-GPU scale. SemiAnalysis calculated the total cost of ownership was 14% higher than the standard NVL72 configuration — a premium that bought more CPU and DRAM capacity that LLM teams didn't need. Meta's GB300 servers have since reverted to standard configurations, an implicit admission that the Ariel experiment failed.
SemiAnalysis publicly urged AMD to bypass Meta's infrastructure team and work directly with TBD Lab to supply the standard MI450X instead of the custom version. "AMD needs to step in, put on their big boy pants, and work directly with teams at TBD to make sure they get the normal MI450 instead of the gimped Meta custom version which is terrible for GenAI," the report said.
For investors, the pattern raises questions about Meta's ability to efficiently deploy its AI capital expenditure. Meta trades at roughly 23x forward earnings, and the infrastructure dysfunction could pressure margins if the company continues to pay premium prices for custom hardware that underperforms standard configurations. Nvidia, meanwhile, stands to benefit as TBD Lab gravitates toward Vera Rubin, while AMD risks losing a major volume customer if Meta's infrastructure team continues to prioritize recommendation system workloads over generative AI performance.
This article is for informational purposes only and does not constitute investment advice.