IBM and Together AI's $240M Bet Is Really About Inference, Not Training
A new dedicated inference cluster is a small signal of a much bigger shift in where AI infrastructure spend is actually going.
IBM and Together AI signed a multi-year, $240 million agreement to build a dedicated large-scale inference cluster on IBM Cloud, running on NVIDIA HGX B300 systems with Spectrum-X networking, expected to come online in Q1 2027. IBM is calling it the first dedicated, purpose-built inference cluster of its kind on IBM Cloud. NVIDIA says the new generation of hardware behind it delivers roughly 30x the "AI factory output" of the systems it replaces. Together AI, which already serves something like 400 trillion tokens a month, framed the deal simply: enterprises want frontier-model performance without proprietary-model pricing, and that requires serious infrastructure to deliver at scale.
That last point is the one worth sitting with. For the last couple of years, most infrastructure headlines have been about training — who's building the biggest cluster to train the next frontier model. This deal is explicitly about something else: keeping already-trained open-source models running fast, reliably, and affordably for enterprise customers who chose them specifically to avoid proprietary pricing. That's a quieter story than a training-run arms race, and arguably a more important one, because inference is the cost that never stops. A training run ends. Inference runs for as long as the product exists.
Why this is a signal, not just a deal
A dedicated infrastructure agreement of this size, built specifically for serving open-source models, is a vote of confidence that "open-source plus the right infrastructure" is now a credible alternative to proprietary-model pricing at enterprise scale — not just for teams willing to accept worse performance to save money. That matters for anyone weighing open versus closed models for production use: the constraint increasingly isn't the model. It's whether the infrastructure underneath it can actually serve that model the way your customers need it served.
We'd expect to see more of this kind of deal, not fewer: purpose-built inference infrastructure, decoupled from any single model provider, sized for real production traffic rather than a training run. That's the same bet we're making at Smartbird — that the infrastructure decisions happening after a model is trained are just as consequential as the ones that get all the attention beforehand.