IBM–Together AI $240M B300 cluster: CapEx math vs token claims
Q1 2027 HGX B300 on IBM Cloud is a real spend signal — not a disclosed tokens-per-day figure
Key Developments
On 11 August 2026, IBM announced a multi-year $240 million agreement with Together AI to deploy a large inference cluster of NVIDIA HGX B300 systems on IBM Cloud, with expected availability in Q1 2027. The companies frame it as the first dedicated, large-scale inference cluster on IBM Cloud using HGX B300 plus NVIDIA Spectrum-X Ethernet networking. NVIDIA’s accompanying claim — up to 30× more “AI factory” output versus prior generations — is a vendor relative metric, not a published absolute tokens/sec figure for this deployment.
Together AI will use the capacity for open-source model inference. The same announcement notes the company’s recent $800M Series C at an $8.3B valuation and that its platform currently serves about 400 trillion tokens per month. That volume describes Together’s existing inference franchise across its stack — not the measured capacity of the IBM B300 cluster, which is still months from COD.
“Companies want the performance of the best models without the high price tag of closed systems. That only works if the underlying infrastructure is fast and reliable at scale.” — Vipul Ved Prakash, CEO, Together AI
What the Numbers Do — and Do Not — Say
For diligence, three separations matter:
- $240M is a multi-year commercial envelope. The PR does not break out silicon CapEx versus reserved cloud consumption, networking, storage, or services.
- 400T tokens/month is current platform throughput. Dividing $240M by that volume produces a meaningless “$/M tokens” for the new cluster.
- GPU count and sustained tok/s for the IBM deployment are undisclosed. Without them, tokens/day for this cluster cannot be derived honestly.
Market ranges for 8-GPU B300-class systems (~$300k–$500k) imply that if the full $240M were pure system CapEx, the envelope could cover roughly 480–800 chassis (~3.8k–6.4k GPUs). If instead the money is closer to a three-year cloud reservation at ~$5–$12 per GPU-hour continuous, the same envelope supports on the order of ~760–1,800 GPUs fully utilized. Both are scenarios — not confirmed bill of materials.
CapEx Lens vs Tenstorrent Galaxy
AGICY’s Phase 1 campus narrative uses Tenstorrent Galaxy Blackhole servers at roughly $110k per air-cooled chassis (32 Blackhole chips), with a design target near 17T tokens/year across 1,821 servers (~$200M server CapEx before facility and power). On a chassis sticker alone, HGX B300 systems at ~$300k–$500k sit about 2.7–4.5× above Galaxy — consistent with AGICY’s published CapEx-advantage thesis versus DGX-class NVIDIA racks.
Put another way: the same $240M envelope at Galaxy list pricing buys ~2,180 servers — about 1.2× AGICY’s Phase 1 silicon count — and, at AGICY’s published design density, on the order of ~20T tokens/year. That is a CapEx-substitution thought experiment, not a claim that IBM’s B300 cluster will deliver (or miss) that throughput. Workload mix, utilization, and software stack dominate realized tokens.
Scale context still cuts the other way. Together’s disclosed ~4,800T tokens/year annualized platform volume (~13T tokens/day) is roughly two orders of magnitude above a 17T/year sovereign campus design. Hyperscale open-model APIs and EU sovereign campuses are different jobs: one optimizes global token cost and time-to-capacity on rented NVIDIA; the other optimizes CapEx, air-cooled density, and jurisdiction.
Why This Matters for Sovereign AI
The deal reinforces three market truths AGICY Research tracks: (1) open-source inference is now buying multi-hundred-million-dollar enterprise cloud commitments, not just spot GPUs; (2) Blackwell Ultra / B300 plus Spectrum-X is the near-term hyperscaler default path to Q1 2027 COD; (3) public token headlines travel faster than disclosed unit economics.
For European sovereign operators, the diligence ask is unchanged: publish or measure tok/s per dollar of CapEx and per watt on the silicon you actually own. Marketing “30×” and platform-wide token tallies are not substitutes. AGICY’s comparison set remains CapEx-readable Galaxy economics and EU-aligned hosting — not a claim to match Together’s global API volume on day one.
AGICY Research notes these are third-party developments; this briefing does not claim AGICY participation in the IBM–Together agreement.
Sources
- IBM and Together AI Sign Multi-Year Agreement to Scale Open-Source AI Inference with NVIDIA AI Infrastructure on IBM Cloud — IBM Newsroom
- IBM signs $240m deal to boost AI performance on its cloud — Yahoo Finance / dpa
- IBM Cloud and Together AI expand AI infrastructure with NVIDIA — Cloud Computing News
Sources & References
EU AI Gigafactories tender and Europe’s speed-to-power race
Next →Cerebras Q2 2026: cloud up 281%, hardware down, stock drops
Stay Ahead of the Curve
Get exclusive access to AGICY research, sovereign AI intelligence, and infrastructure insights.
Request Access