IBM and Together AI's $240 million NVIDIA HGX B300 cluster on IBM Cloud targets the open-source inference market as enterprises seek cheaper AI alternatives.
IBM is deploying a $240 million NVIDIA HGX B300 inference cluster on IBM Cloud for Together AI, targeting enterprises shifting to open-source models as closed-system costs and security concerns mount.
"Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale," Vipul Ved Prakash, CEO at Together AI, said.
The multi-year agreement covers a large cluster of NVIDIA HGX B300 systems, which link the chipmaker's Blackwell processors, connected with NVIDIA Spectrum-X Ethernet networking. According to NVIDIA, the setup is built to deliver 30x more AI factory output than prior generations. Together AI, which raised an $800 million Series C at an $8.3 billion valuation in July, reports serving 400 trillion tokens monthly across its inference product.
The cluster, expected available in Q1 2027, is the first dedicated large-scale inference deployment on IBM Cloud using HGX B300 systems. The deal deepens IBM's broader collaboration with NVIDIA, which spans GPU-native data analytics, unstructured data extraction and consulting services.
The partnership reflects a broader shift in enterprise AI procurement. Open-source models such as DeepSeek, MiniMax and Kimi have gained traction as businesses seek to rein in AI costs and weigh concerns about cybersecurity incidents involving models from Anthropic, OpenAI and Meta. Inference — the process of running trained models to generate responses — has become one of the largest drivers of demand for computing capacity, prompting cloud providers and chipmakers to spend billions expanding AI infrastructure.
For Together AI, the IBM Cloud deployment extends its AI Native Cloud platform, which spans inference, training, fine tuning and agentic workflows, into enterprise-grade territory. The company, founded in 2022, says it powers more than a million developers. Its choice of IBM over other cloud providers came down to product roadmaps and the ability to deliver GPU capacity at the pace required for rapid AI scaling, the company said.
For IBM, the agreement adds a marquee AI workload to its cloud business as it competes with hyperscalers including Amazon Web Services, Microsoft Azure and Google Cloud for enterprise AI infrastructure spend. IBM shares, trading at $236.56, rose 0.11 percent on the news.
The deal also extends NVIDIA's push into dedicated inference infrastructure. The Blackwell-based HGX B300 systems are optimized for AI inference, and the 30x output claim, per NVIDIA, positions the cluster as a reference deployment for open-source model serving. NVIDIA did not disclose the test conditions for the 30x comparison.
For investors, the deal points to sustained demand for NVIDIA's Blackwell line and confirms the open-source inference market's trajectory. Together AI's $8.3 billion valuation and 400 trillion monthly tokens show the scale of inference demand. IBM's cloud revenue growth stands to benefit as the cluster comes online in Q1 2027, though the deployment timeline means financial impact is more than a year away.
This article is for informational purposes only and does not constitute investment advice.