IBM and Together AI Scale Open-Source AI Inference on IBM Cloud

IBM and Together AI Scale Open-Source AI Inference on IBM Cloud

First large-scale inference cluster with Together AI on IBM Cloud using NVIDIA HGX B300 systems to help enterprises run AI workloads, designed for fast and efficient production.

IBM and Together AI are collaborating to deliver IBM and NVIDIA AI infrastructure.

Under a multi-year US$240m agreement between IBM and Together AI, IBM is positioned to deploy a large cluster of NVIDIA HGX B300 systems on IBM Cloud with expected availability in Q1 2027.

Together AI will use this cluster to provide open-source model inference. This deployment is the first dedicated, large-scale cluster built for inference on IBM Cloud using HGX B300 systems and NVIDIA Spectrum-X Ethernet networking. According to NVIDIA, it is built to deliver 30x more AI factory output compared to prior generations.

This collaboration aims to enable Together AI to deliver better performance and token economics to enterprises as they look to efficiently scale their AI deployments. Together AI is built on the principle that open-source models are essential for the future of AI and developers should be able to build using open, modular stacks.

The company recently raised an US$800m Series C financing round at an US$8.3 billion valuation to expand its AI Native Cloud and its platform spans capabilities across inference, training, fine tuning and agentic workflows. The company reports that it has seen significant momentum for its inference product, now serving 400 trillion tokens monthly.

Together AI selected IBM with NVIDIA because of their innovative product roadmaps and their ability to deliver GPU capacity at the pace required for rapid AI scaling and lowest token cost.

Building on IBM’s expertise in delivering enterprise-grade cloud capabilities, this collaboration aims to help Together AI to continue its expansion into the enterprise space while making open-source AI more accessible to developers and enterprises around the world.

“Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale,” said Vipul Ved Prakash, CEO, Together AI. “Working alongside IBM with NVIDIA gives us that foundation. This cluster lets us bring production-grade inference to more companies, faster, and it’s a big step in our push to make open-source AI the obvious choice for enterprises.”

“Enterprises are in a race to adopt agentic AI at scale to drive real business outcomes,” said Alan Peacock, General Manager, IBM Cloud. “IBM and NVIDIA are delivering scalable, economical, enterprise-grade AI infrastructure that can help Together AI accelerate innovation for the next generation of AI infrastructure.”

“AI factories are becoming essential enterprise infrastructure – like electricity and telecommunications – turning compute and data into intelligence,” said Dion Harris, Senior Director, HPC and AI Infrastructure Solutions, NVIDIA. “With NVIDIA HGX B300 systems and NVIDIA Spectrum-X Ethernet networking on IBM Cloud, IBM and Together AI will deliver an accelerated computing platform to help enterprises deploy open-source AI with the performance, efficiency and scale required for real-time AI services.”

Browse our latest issue

Intelligent CIO North America

View Magazine Archive