Datacenter Infrastructure Layer

Datacenter Infrastructure Layer

Colocation, GPU clouds, and hyperscale campuses

Datacenter Infrastructure Layer

The infrastructure layer transforms raw silicon into usable compute capacity. This is where GPUs become clusters, clusters become supercomputers, and supercomputers become AI factories.

The Value Chain

graph LR
    A[Server OEMs/ODMs] --> B[Rack Integration]
    B --> C[Row/Cluster Deployment]
    C --> D[Datacenter Campus]
    D --> E[GPU Cloud Services]
    E --> F[AI Model Training/Inference]

1. Server OEMs & ODMs (The Box Builders)

Key Players: Foxconn, Quanta, Wistron, Inventec, Supermicro, Dell, HPE, Lenovo

CompanyRoleKey PlatformsNotable Customers
Foxconn (Hon Hai)Largest ODMHGX H100/B200, GB200 NVL72NVIDIA, AWS, Google, Microsoft
Quanta (QCT)Major ODMHGX, Grace Hopper, customMeta, Microsoft, Oracle
Wistron (Wiwynn)ODMHGX, AMD MI300, customNVIDIA, AMD, CSPs
InventecODMHGX, liquid-cooled racksNVIDIA, CSPs
SupermicroOEM/ODMSuperBlade, Hyper, liquidCoreWeave, Lambda, enterprises
Dell TechnologiesOEMPowerEdge XE9680, XE9780Enterprises, CSPs
HPEOEMCray EX, ProLiant DL380aGovernment, enterprises
LenovoOEMThinkSystem SR680a, SR780aEnterprises, research

Reference Architectures:

  • NVIDIA HGX H100/B200: 8-GPU baseboard, NVLink/NVSwitch, SXM5
  • NVIDIA GB200 NVL72: 72 GPUs, 36 Grace CPUs, liquid-cooled rack
  • AMD MI300X Platform: 8-GPU OAM, Infinity Fabric
  • Google TPU v5p Pod: 8,960 chips, custom interconnect
  • AWS Trainium2 UltraCluster: 100K+ chips, EFA networking

2. Rack-Scale Integration & Liquid Cooling

Key Players: Vertiv, Schneider Electric, Delta Electronics, Aavid (Vertiv), CoolIT, Motivair, Submer, Iceotope

TechnologyProvidersCapacityAdoption
Direct-to-Chip (Cold Plate)Vertiv, Aavid, CoolIT, Delta1-2 kW/GPUMainstream (H100/B200)
Immersion (Single-Phase)Submer, Iceotope, GRC2-5 kW/GPUGrowing (GB200, MI300X)
Immersion (Two-Phase)3M, Engineered Fluids5-10 kW/GPUEarly adoption
Rear Door Heat ExchangerVertiv, Schneider, MotivairSupplementalRetrofit-friendly

Power Density Evolution:

  • Traditional: 5-10 kW/rack (air cooled)
  • HGX H100: 40-50 kW/rack (air + liquid assist)
  • GB200 NVL72: 120-140 kW/rack (full liquid)
  • Future (Rubin/MI400): 200+ kW/rack (immersion required)

3. Datacenter Campus Operators (Colocation & Hyperscale)

Key Players: Equinix, Digital Realty, CoreWeave, Lambda Labs, Crusoe, Applied Digital, QTS, CyrusOne, Iron Mountain, STACK, Vantage

Colocation Giants (Multi-tenant)

OperatorLocationsCapacity (MW)AI Focus
Equinix (EQIX)250+ across 33 countries2,000+xScale for hyperscalers, GPU-as-a-Service
Digital Realty (DLR)300+ across 25 countries3,000+High-density, liquid-ready, 300MW+ campuses
QTS30+ North America/Europe1,000+Hyperscale lease, 500MW Atlanta campus
CyrusOne50+ global1,000+High-density, zero-water cooling
Vantage Data Centers25+ campuses2,000+Hyperscale build-to-suit, 1GW+ campuses

GPU Cloud Specialists (Pure-play AI)

OperatorGPUsValuationKey Differentiator
CoreWeave28,000+ (H100/H200/B200)$19B+ (2024)Kubernetes-native, flexible contracts
Lambda Labs10,000+ (H100/A100)$1.5B+On-demand, research-friendly
Crusoe5,000+ (H100)$2.8B+Flared gas power, climate-aligned
Applied Digital5,000+ (planned 100K)$1B+400MW North Dakota campus
Together AI5,000+$1.25B+Open-source focus, inference optimized
RunPod3,000+PrivateServerless GPU, per-second billing

Hyperscale Self-Build (Captive Capacity)

CompanyEstimated GPUs (2024)Annual CapExStrategy
Microsoft/Azure300K+$50B+Own silicon (Maia), OpenAI partnership
Google200K+ (TPU + GPU)$40B+TPU dominance, Gemini training
Amazon/AWS150K+ (Trainium + NVIDIA)$50B+Custom silicon, UltraClusters
Meta600K+ (H100 equiv by 2024)$35B+Llama training, open models
Oracle50K+$10B+Sovereign cloud, NVIDIA partnership
xAI100K+ (Memphis “Colossus”)$10B+Single massive cluster

4. Networking & Interconnect (The Nervous System)

Key Players: NVIDIA, Broadcom, Marvell, Cisco, Arista, Juniper, Celestica

TechnologyProviderSpeedApplication
NVLink/NVSwitchNVIDIA900 GB/s (NVLink 5)GPU-GPU within node
InfiniBand (NDR/XDR)NVIDIA (Mellanox)400/800 GbpsCluster backbone
Ethernet (RoCE v2)Broadcom, Cisco, Arista400/800 GbpsAlternative to IB
NVLink-C2CNVIDIA900 GB/sGrace-Grace, Grace-GPU
Ultra Ethernet (UEC)Consortium800 Gbps+Future standard

Cluster Scale Networking:

  • 1,024 GPU (128 node): ~500 InfiniBand switches, 10K+ cables
  • 16,384 GPU (2,048 node): ~8,000 switches, 150K+ cables
  • 100K+ GPU (UltraCluster): Custom topologies, 1M+ cables

5. Storage for AI (Feeding the Beast)

Key Players: WEKA, VAST Data, DDN, Pure Storage, NetApp, Hammerspace, MinIO

SolutionArchitectureThroughputCustomers
WEKA Data PlatformParallel FS, NVMe-oF10 TB/s+CoreWeave, Lambda, Stability AI
VAST DataStoreDisaggregated, QLC5 TB/s+CoreWeave, xAI, enterprises
DDN EXAScalerLustre-based, NVMe10 TB/s+NVIDIA DGX SuperPOD, labs
Pure Storage FlashBladeScale-out, DirectFlash2 TB/s+Enterprises, CSPs
HammerspaceGlobal namespaceN/AHybrid/multi-cloud

Infrastructure Investment per 100K GPU Cluster

ComponentCostLead TimeNotes
Servers (HGX B200 x 12,500)$4-5B12-18 months$350-400K per 8-GPU node
Networking (IB/RoCE)$500M-1B6-12 months400/800G switches, optics, cables
Storage (100PB+)$200-500M6-12 monthsAll-flash, parallel filesystem
Liquid Cooling Infrastructure$300-500M9-15 monthsCDUs, manifolds, piping, chillers
Datacenter Build (shell + MEP)$1-2B18-24 months150-200MW campus
Power Infrastructure$500M-1B18-36 monthsSubstations, transformers, generators
Total Infrastructure (ex-GPUs)$6.5-10B24-36 monthsGPUs add $30-40B+

The infrastructure layer is where capital intensity peaks. A single 100K GPU AI factory requires $40-50B total investment — comparable to the GDP of a small nation. The companies that master integration, cooling, and power delivery at this scale will define the next decade of AI.