Datacenter Infrastructure Layer
Colocation, GPU clouds, and hyperscale campuses
Datacenter Infrastructure Layer
The infrastructure layer transforms raw silicon into usable compute capacity. This is where GPUs become clusters, clusters become supercomputers, and supercomputers become AI factories.
The Value Chain
graph LR
A[Server OEMs/ODMs] --> B[Rack Integration]
B --> C[Row/Cluster Deployment]
C --> D[Datacenter Campus]
D --> E[GPU Cloud Services]
E --> F[AI Model Training/Inference]
1. Server OEMs & ODMs (The Box Builders)
Key Players: Foxconn, Quanta, Wistron, Inventec, Supermicro, Dell, HPE, Lenovo
| Company | Role | Key Platforms | Notable Customers |
|---|---|---|---|
| Foxconn (Hon Hai) | Largest ODM | HGX H100/B200, GB200 NVL72 | NVIDIA, AWS, Google, Microsoft |
| Quanta (QCT) | Major ODM | HGX, Grace Hopper, custom | Meta, Microsoft, Oracle |
| Wistron (Wiwynn) | ODM | HGX, AMD MI300, custom | NVIDIA, AMD, CSPs |
| Inventec | ODM | HGX, liquid-cooled racks | NVIDIA, CSPs |
| Supermicro | OEM/ODM | SuperBlade, Hyper, liquid | CoreWeave, Lambda, enterprises |
| Dell Technologies | OEM | PowerEdge XE9680, XE9780 | Enterprises, CSPs |
| HPE | OEM | Cray EX, ProLiant DL380a | Government, enterprises |
| Lenovo | OEM | ThinkSystem SR680a, SR780a | Enterprises, research |
Reference Architectures:
- NVIDIA HGX H100/B200: 8-GPU baseboard, NVLink/NVSwitch, SXM5
- NVIDIA GB200 NVL72: 72 GPUs, 36 Grace CPUs, liquid-cooled rack
- AMD MI300X Platform: 8-GPU OAM, Infinity Fabric
- Google TPU v5p Pod: 8,960 chips, custom interconnect
- AWS Trainium2 UltraCluster: 100K+ chips, EFA networking
2. Rack-Scale Integration & Liquid Cooling
Key Players: Vertiv, Schneider Electric, Delta Electronics, Aavid (Vertiv), CoolIT, Motivair, Submer, Iceotope
| Technology | Providers | Capacity | Adoption |
|---|---|---|---|
| Direct-to-Chip (Cold Plate) | Vertiv, Aavid, CoolIT, Delta | 1-2 kW/GPU | Mainstream (H100/B200) |
| Immersion (Single-Phase) | Submer, Iceotope, GRC | 2-5 kW/GPU | Growing (GB200, MI300X) |
| Immersion (Two-Phase) | 3M, Engineered Fluids | 5-10 kW/GPU | Early adoption |
| Rear Door Heat Exchanger | Vertiv, Schneider, Motivair | Supplemental | Retrofit-friendly |
Power Density Evolution:
- Traditional: 5-10 kW/rack (air cooled)
- HGX H100: 40-50 kW/rack (air + liquid assist)
- GB200 NVL72: 120-140 kW/rack (full liquid)
- Future (Rubin/MI400): 200+ kW/rack (immersion required)
3. Datacenter Campus Operators (Colocation & Hyperscale)
Key Players: Equinix, Digital Realty, CoreWeave, Lambda Labs, Crusoe, Applied Digital, QTS, CyrusOne, Iron Mountain, STACK, Vantage
Colocation Giants (Multi-tenant)
| Operator | Locations | Capacity (MW) | AI Focus |
|---|---|---|---|
| Equinix (EQIX) | 250+ across 33 countries | 2,000+ | xScale for hyperscalers, GPU-as-a-Service |
| Digital Realty (DLR) | 300+ across 25 countries | 3,000+ | High-density, liquid-ready, 300MW+ campuses |
| QTS | 30+ North America/Europe | 1,000+ | Hyperscale lease, 500MW Atlanta campus |
| CyrusOne | 50+ global | 1,000+ | High-density, zero-water cooling |
| Vantage Data Centers | 25+ campuses | 2,000+ | Hyperscale build-to-suit, 1GW+ campuses |
GPU Cloud Specialists (Pure-play AI)
| Operator | GPUs | Valuation | Key Differentiator |
|---|---|---|---|
| CoreWeave | 28,000+ (H100/H200/B200) | $19B+ (2024) | Kubernetes-native, flexible contracts |
| Lambda Labs | 10,000+ (H100/A100) | $1.5B+ | On-demand, research-friendly |
| Crusoe | 5,000+ (H100) | $2.8B+ | Flared gas power, climate-aligned |
| Applied Digital | 5,000+ (planned 100K) | $1B+ | 400MW North Dakota campus |
| Together AI | 5,000+ | $1.25B+ | Open-source focus, inference optimized |
| RunPod | 3,000+ | Private | Serverless GPU, per-second billing |
Hyperscale Self-Build (Captive Capacity)
| Company | Estimated GPUs (2024) | Annual CapEx | Strategy |
|---|---|---|---|
| Microsoft/Azure | 300K+ | $50B+ | Own silicon (Maia), OpenAI partnership |
| 200K+ (TPU + GPU) | $40B+ | TPU dominance, Gemini training | |
| Amazon/AWS | 150K+ (Trainium + NVIDIA) | $50B+ | Custom silicon, UltraClusters |
| Meta | 600K+ (H100 equiv by 2024) | $35B+ | Llama training, open models |
| Oracle | 50K+ | $10B+ | Sovereign cloud, NVIDIA partnership |
| xAI | 100K+ (Memphis “Colossus”) | $10B+ | Single massive cluster |
4. Networking & Interconnect (The Nervous System)
Key Players: NVIDIA, Broadcom, Marvell, Cisco, Arista, Juniper, Celestica
| Technology | Provider | Speed | Application |
|---|---|---|---|
| NVLink/NVSwitch | NVIDIA | 900 GB/s (NVLink 5) | GPU-GPU within node |
| InfiniBand (NDR/XDR) | NVIDIA (Mellanox) | 400/800 Gbps | Cluster backbone |
| Ethernet (RoCE v2) | Broadcom, Cisco, Arista | 400/800 Gbps | Alternative to IB |
| NVLink-C2C | NVIDIA | 900 GB/s | Grace-Grace, Grace-GPU |
| Ultra Ethernet (UEC) | Consortium | 800 Gbps+ | Future standard |
Cluster Scale Networking:
- 1,024 GPU (128 node): ~500 InfiniBand switches, 10K+ cables
- 16,384 GPU (2,048 node): ~8,000 switches, 150K+ cables
- 100K+ GPU (UltraCluster): Custom topologies, 1M+ cables
5. Storage for AI (Feeding the Beast)
Key Players: WEKA, VAST Data, DDN, Pure Storage, NetApp, Hammerspace, MinIO
| Solution | Architecture | Throughput | Customers |
|---|---|---|---|
| WEKA Data Platform | Parallel FS, NVMe-oF | 10 TB/s+ | CoreWeave, Lambda, Stability AI |
| VAST DataStore | Disaggregated, QLC | 5 TB/s+ | CoreWeave, xAI, enterprises |
| DDN EXAScaler | Lustre-based, NVMe | 10 TB/s+ | NVIDIA DGX SuperPOD, labs |
| Pure Storage FlashBlade | Scale-out, DirectFlash | 2 TB/s+ | Enterprises, CSPs |
| Hammerspace | Global namespace | N/A | Hybrid/multi-cloud |
Infrastructure Investment per 100K GPU Cluster
| Component | Cost | Lead Time | Notes |
|---|---|---|---|
| Servers (HGX B200 x 12,500) | $4-5B | 12-18 months | $350-400K per 8-GPU node |
| Networking (IB/RoCE) | $500M-1B | 6-12 months | 400/800G switches, optics, cables |
| Storage (100PB+) | $200-500M | 6-12 months | All-flash, parallel filesystem |
| Liquid Cooling Infrastructure | $300-500M | 9-15 months | CDUs, manifolds, piping, chillers |
| Datacenter Build (shell + MEP) | $1-2B | 18-24 months | 150-200MW campus |
| Power Infrastructure | $500M-1B | 18-36 months | Substations, transformers, generators |
| Total Infrastructure (ex-GPUs) | $6.5-10B | 24-36 months | GPUs add $30-40B+ |
The infrastructure layer is where capital intensity peaks. A single 100K GPU AI factory requires $40-50B total investment — comparable to the GDP of a small nation. The companies that master integration, cooling, and power delivery at this scale will define the next decade of AI.