IBM has struck a multi-year agreement worth $240m with Together AI, under which the technology group will build a large cluster of Nvidia HGX B300 systems on IBM Cloud for open-source model inference.
The cluster is expected to become available in the first quarter of 2027.
IBM said it would be the first dedicated, large-scale cluster designed for inference on IBM Cloud using HGX B300 systems, linked by Nvidia Spectrum-X Ethernet networking.
Nvidia claims that the configuration is built to deliver 30 times more AI factory output than earlier generations.
IBM Cloud general manager Alan Peacock said: “Enterprises are in a race to adopt agentic AI at scale to drive real business outcomes. IBM and Nvidia are delivering scalable, economical, enterprise-grade AI infrastructure that can help Together AI accelerate innovation for the next generation of AI infrastructure.”
Together AI, which will operate inference workloads on the capacity, chose IBM and Nvidia because of their product roadmaps and their capacity to supply GPUs at the speed and cost the company requires.
The arrangement is intended to support Together AI’s move further into the enterprise market and to widen access to open-source models among developers and businesses.
Together AI CEO Vipul Ved Prakash said: “Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale.
“Working alongside IBM with Nvidia gives us that foundation. This cluster lets us bring production-grade inference to more companies, faster, and it’s a big step in our push to make open-source AI the obvious choice for enterprises.”
Last month, Together AI completed an $800m Series C financing round at a valuation of $8.3bn, which it is using to expand its AI Native Cloud.
Its platform covers inference, training, fine tuning and agentic workflows, and the company reports that its inference product now handles 400 trillion tokens a month.