Infrastructure Economics
AI infrastructure is the largest capital expenditure wave in technology history. Understanding the economics - where the money flows, what drives the numbers, and how the cost of AI computation has evolved - matters for anyone making strategic decisions about AI investment, whether building products or evaluating the industry.
Capital Expenditure at Scale
The major tech companies' AI infrastructure investment in 2025:
| Company | 2025 AI CapEx (est.) | Primary Use |
|---|---|---|
| Microsoft (Azure + OpenAI) | ~$80B | Azure AI GPU clusters, data centers |
| Google (GCP + DeepMind) | ~$75B | TPU clusters, data centers, Gemini training |
| Amazon (AWS) | ~$75B | AWS AI instances, custom Trainium chips |
| Meta | ~$60B | Llama training clusters, internal AI infra |
| Stargate (US govt initiative) | $500B (5-year) | National AI infrastructure initiative; OpenAI/SoftBank/Oracle JV |
Data Center Cost Breakdown
For a representative 100MW AI training data center ($2โ4B total cost):
Construction costs:
Land and site prep: 5-10% ($100-400M)
Building and structure: 10-15% ($200-600M)
Power infrastructure: 20-25% ($400-1000M)
(substation, transformers, UPS, generators)
Cooling infrastructure: 10-15% ($200-600M)
(chillers, cooling towers, liquid cooling CDUs)
Networking: 5-10% ($100-400M)
(spine/leaf switches, InfiniBand, fiber)
IT Equipment:
GPU servers (H100/B200): 40-50% ($800M-2B)
10,000 ร H100 DGX = ~$300K/server ร 1,250 servers
Storage systems: 5-10% ($100-400M)
Annual operating costs (for a 100MW facility):
Electricity (at $0.06/kWh): ~$52M/year
Staff (100 engineers): ~$20M/year
Cooling water and chemicals: ~$5M/year
Maintenance and parts: ~$15M/year
Total OpEx: ~$90-100M/yearGPU Supply Chain Economics
NVIDIA dominates the AI accelerator market with 80โ85% market share (as of 2025). This concentration creates significant pricing power:
- H100 80GB SXM5: MSRP ~$30,000; market price peaked at $40,000+ in 2023โ2024 due to supply shortages; normalizing in 2025
- H200: $35,000โ45,000; 2ร memory bandwidth vs H100
- GB200 NVL72 rack: ~$3M per rack (72 GPUs + Grace CPUs + interconnect)
- NVIDIA's gross margin on AI chips: ~75% - one of the highest hardware margins in history
- TSMC manufactures NVIDIA's chips on its 4nm and 3nm processes; TSMC is the sole manufacturer
Cost per Token - The Driving Metric
The economics of AI inference are captured by cost per million tokens (input + output). This has declined dramatically:
| Model / Period | Input $/MTok | Change |
|---|---|---|
| GPT-3.5-turbo (2023) | $2.00/MTok | Baseline |
| GPT-4-turbo (2024) | $10.00/MTok | 5ร more capable, 5ร pricier |
| GPT-4o (2024) | $2.50/MTok | GPT-4 quality, 4ร cheaper |
| GPT-4o mini (2024) | $0.15/MTok | Good capability, 13ร cheaper than GPT-4o |
| Gemini Flash 2.5 (2025) | $0.075โ1.00/MTok | Frontier quality at commodity prices |
The cost of frontier-quality intelligence has dropped ~100ร in three years (2022โ2025). This compression is driven by: improved model architectures (same capability at fewer parameters), hardware efficiency gains, competitive pressure from open-weight models (Llama 3), and operational efficiency at scale.
Return on Investment Question
Whether the massive infrastructure investment generates adequate returns is the defining business question of 2025โ2027:
- AWS, GCP, and Azure are selling cloud AI capacity - revenue scales with AI adoption; clear ROI path
- Meta is using AI for recommendation, advertising, and Llama - direct revenue impact measurable
- Microsoft has embedded Copilot across Office and GitHub - measuring ROI per enterprise customer
- OpenAI, Anthropic, and similar labs - revenue ($3โ5B annually) currently well below compute costs; betting on long-term dominance and capability-driven pricing power
The optimistic case: AI becomes as fundamental as cloud computing, and the current investment is equivalent to building out cloud infrastructure in 2008โ2015 - a period when the ROI was also unclear in real time but became transformative in retrospect. The skeptical case: the compute investment outpaces commercial demand, leading to overcapacity and margin compression, similar to the fiber optic overbuild of the late 1990s.