IBM and startup ‘Together AI’ have agreed on a $240 million, multi-year partnership to build a dedicated AI inference cluster on IBM Cloud. Under the deal, IBM will deploy a large NVIDIA-powered computing cluster based on HGX B300 GPU systems and NVIDIA Spectrum-X Ethernet networking and Together AI will use it to serve open-source AI models in production.
The infrastructure is expected to be ready in Q1 2027. This arrangement is being pitched as the first dedicated large-scale inference cluster on IBM Cloud using NVIDIA’s latest HGX B300 hardware.
Together AI, which focuses on open-model training and inference, sees this deal as a way to offer “production-grade inference” to its customers.
[ALSO READ: IBM Launches Apptio AI Value ROI to Measure Business Returns on AI Investments ]
The IBM–Together AI Partnership
IBM will supply IBM Cloud infrastructure powered by NVIDIA hardware, and Together AI will be the customer running its inference services on that infrastructure. According to IBM, “under a multi-year $240M agreement”, IBM will deploy a large cluster of NVIDIA HGX B300 systems on IBM Cloud. That cluster will use NVIDIA’s Spectrum-X Ethernet networking and is intended for high-performance AI inference.
This will be the first such dedicated inference cluster on IBM Cloud using HGX B300 and Spectrum-X. The goal is to deliver “performance, efficiency and scale” so enterprises can run open-source AI models with the required throughput and low latency.
“Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale,” said Vipul Ved Prakash, CEO at Together AI. “Working alongside IBM with NVIDIA gives us that foundation. This cluster lets us bring production-grade inference to more companies, faster, and it’s a big step in our push to make open-source AI the obvious option for enterprises.”
Technology Behind the Cluster
The NVIDIA HGX B300 is at the heart of the cluster. HGX B300 is NVIDIA’s latest reference platform for data centers, featuring eight Blackwell-generation GPUs on a single baseboard. Each Blackwell GPU on B300 has 288 GB of HBM3e memory (for 2.3 TB per node) and connects via NVSwitch to provide up to 64 TB/s of aggregate GPU bandwidth. The platform delivers on the order of 144 petaflops of AI compute per node, turning it an “AI powerhouse” for large language models and other deep learning workloads.
NVIDIA’s Spectrum-X Ethernet networking is the other key component. Spectrum-X is NVIDIA’s next-generation Ethernet platform designed specifically for AI data centers. It promises high throughput and low latency – NVIDIA advertises up to 1.6× network acceleration for AI workloads compared to off-the-shelf Ethernet – and is meant to scale across thousands of GPUs. In practice, Spectrum-X gear includes high-speed Ethernet switches and “SuperNIC” network interface cards tuned for AI traffic.
IBM highlights that this combination of HGX B300 and Spectrum-X is optimized for inference: IBM says NVIDIA claims it provides “30x more AI factory output” than prior-gen platforms. (“AI factory output” is a vendor metric related to throughput on typical workloads.) This confirms that the cluster is focused on delivering very high inference performance and effectiveness. The HGX B300 baseboard has eight GPUs with 800 Gb/s external networking each (via 2×400 Gb/s Ethernet per GPU), and the cluster will be built out with multiple such boards, linked by Spectrum-X switching.
“Enterprises are in a race to adopt agentic AI at scale to drive real business outcomes,” said Alan Peacock, General Manager of IBM Cloud. “IBM and NVIDIA are delivering scalable, economical, enterprise-grade AI infrastructure that can help Together AI accelerate innovation for the next generation of AI infrastructure.”
[ALSO READ: IBM to Acquire HRL Laboratories to Advance Quantum Computing Strategy ]
“AI factories are becoming essential enterprise infrastructure—like electricity and telecommunications—turning compute and data into intelligence,” said Dion Harris, Senior Director, HPC and AI Infrastructure Solutions, NVIDIA. “With NVIDIA HGX B300 systems and NVIDIA Spectrum-X Ethernet networking on IBM Cloud, IBM and Together AI will deliver an accelerated computing platform to help enterprises deploy open-source AI with the performance, efficiency and scale required for real-time AI services.”
This IBM and Together AI agreement fits into a wider shift. Earlier in the AI boom, many big announcements were about getting more GPUs for training (for example, buying clusters for building new models). Now, as models proliferate, attention has turned to how to serve those models cheaply and quickly in business applications.




















