One Computer the Size of a Building: How a Frontier AI Training Cluster Actually Works
Frontier AI clusters lash hundreds of thousands of GPUs into a single synchronous machine. Inside the NVLink domains, network fabric wars, oceans of fiber, and relentless hardware failures that define training at 100,000-GPU scale.
A frontier AI training cluster is not a data center full of servers. It is closer to a single computer that happens to be the size of a building: one workload, one clock, hundreds of thousands of processors advancing in lockstep. When Meta trained Llama 3, its 16,384 GPUs had to exchange gradients after essentially every step. When xAI’s Colossus came online with 100,000 H100s, NVIDIA described it as one coherent machine on a single fabric. Understanding how these systems actually work means understanding two networks, an ocean of fiber, and an uncomfortable amount of math about failure.
One job, two networks
Every modern cluster is really two machines nested inside each other. The inner machine is the “scale-up” domain: a small group of GPUs wired together so tightly they can read each other’s memory almost as if it were local. In NVIDIA’s GB200 NVL72 rack, the current workhorse of frontier builds, 72 Blackwell GPUs share an NVLink domain with 130 TB/s of aggregate bandwidth, according to NVIDIA. That link runs over a spine of more than 5,000 copper cables, roughly two miles of wire per rack; NVIDIA says choosing copper over optics saves about 20 kW per rack that transceivers would otherwise burn.
The outer machine is the “scale-out” fabric: InfiniBand or Ethernet connecting thousands of those racks. NVIDIA’s own engineering blog puts NVLink at roughly 18 times the bandwidth of the scale-out network. That asymmetry dictates how models are split: tensor parallelism, the chattiest kind, stays inside the NVLink domain, while data and pipeline parallelism cross the slower fabric between racks.
The fabric wars
For years the scale-out layer belonged to InfiniBand, the lossless HPC interconnect NVIDIA acquired with Mellanox. Meta’s 2024 buildout captured the industry’s uncertainty perfectly: the company built two identical 24,576-GPU clusters, one on Quantum-2 InfiniBand and one on RoCE Ethernet using Arista switches, then trained Llama 3 on the Ethernet one without, according to Meta’s engineering blog, any network bottlenecks.
xAI made the more aggressive bet. Colossus, its Memphis cluster of 100,000 H100s stood up in just 122 days, runs entirely on NVIDIA’s Spectrum-X Ethernet, which the company says sustains 95 percent data throughput versus roughly 60 percent for vanilla Ethernet. The market has followed: Dell’Oro Group reports Ethernet now leads AI back-end networks even as InfiniBand sales surged again in mid-2025, and the Ultra Ethernet Consortium’s 1.0 specification, released in June 2025 with backing from Meta, Broadcom, AMD, and over 100 other members, aims to make open Ethernet behave like InfiniBand at scale.
A cathedral of glass
The physical layer is staggering. SemiAnalysis estimates a standard 100,000-GPU H100 cluster needs about 98,304 optical transceivers just to connect GPUs to their first-tier switches, with over a million fiber strands running between switching layers, and draws roughly 150 MW of critical IT power, about 1.59 TWh per year. Designs diverge on how to tame this: rail-optimized topologies maximize bandwidth but drown in optics, while middle-of-rack layouts swap a quarter to a third of those links for cheap, reliable copper.
Marching in lockstep, and falling down
What makes all this so unforgiving is that frontier training is synchronous. Every iteration, every GPU must finish its slice of work and exchange results before anyone proceeds. One slow chip drags 100,000 others; one dead chip stops the job cold. The standard defense is checkpointing, periodically writing model state to storage and rewinding after a crash. Meta’s training-at-scale blog describes pouring engineering effort into cutting checkpoint and restart times, while SemiAnalysis notes that leading labs now rebuild a failed node’s memory directly over the network via RDMA, losing roughly a single iteration instead of the hundreds a checkpoint rewind can cost.
They need it, because at this scale hardware failure is not an event but a climate.
On a 100,000-GPU cluster, SemiAnalysis calculates that even if every optical link had a five-year mean time to failure, the first job-killing fault would arrive in about 26 minutes.
The best public ground truth comes from Meta’s Llama 3 405B run. Across 54 days on 16,384 H100s, the team logged 466 job interruptions, 419 of them unplanned, about one every three hours, according to Meta’s paper as reported by Tom’s Hardware. Faulty GPUs caused 30.1 percent of unexpected failures and HBM3 memory another 17.2 percent, yet automation kept effective training time above 90 percent.
The giga-clusters
Those numbers describe yesterday’s scale. Today’s flagship sites are an order of magnitude larger.
| Site | Chips | Fabric | Scale |
|---|---|---|---|
| Meta Llama 3 clusters (2024) | 2 x 24,576 H100 | RoCE Ethernet and InfiniBand | tens of MW each |
| xAI Colossus 1, Memphis | 100,000 H100, later ~200,000 | Spectrum-X Ethernet | ~150-300 MW |
| OpenAI Stargate, Abilene | up to ~450,000 GB200-class | mixed NVIDIA fabrics | 1.2 GW campus |
OpenAI’s Stargate flagship in Abilene, Texas, built by Crusoe for Oracle, spans eight buildings and about four million square feet at full build, with Data Center Dynamics reporting plans for roughly 450,000 GB200 GPUs; Epoch AI’s tracking puts it among the most complete of the Stargate sites. xAI’s Colossus 2, meanwhile, crossed the gigawatt threshold in early 2026 on its way to a reported 2 GW and roughly 550,000 GB200 and GB300-class GPUs, per industry tracker Introl, though outside estimates of deployed chip counts at any given moment vary widely and should be read as ranges, not census data.
The through-line is the same everywhere: the computer has outgrown the server, then the rack, then the room. The building is now the unit of compute, and the next generation of models will be trained by machines that consume small cities’ worth of power while praying, statistically, that not too many of their million parts break in the same three-hour window.