Software companies learned to describe infrastructure as elastic. Compute appeared when requested, disappeared when released, and arrived on a monthly bill. The abstraction was strategically useful: teams could discuss features without discussing substations.
AI is breaking that abstraction.
Frontier-scale capacity is tied to accelerators, electrical equipment, cooling systems, land, network fabrics, construction schedules, and long-term power arrangements. Those assets do not scale with a slider. They are financed, permitted, installed, depreciated, and kept busy—or left idle. As AI spending grows, infrastructure is no longer a technical substrate hidden beneath the product. It is becoming part of the product’s economic design.
Capacity is a portfolio of commitments
A company buying AI capacity is making several bets simultaneously. It is betting that demand will arrive, that its workloads will fit the hardware, that model architectures will not strand the investment, and that sufficient power will be available where the equipment lands. It is also betting that software improvements will not make yesterday’s capacity plan absurdly oversized.
The International Energy Agency’s Energy and AI report makes the physical constraint impossible to ignore: training and inference happen in data centers whose expansion depends on electricity systems with different planning horizons from software. A model can be released in months. Transmission, generation, and large electrical interconnections often move on much slower clocks.
This mismatch turns capacity planning into an options problem. Committing too little can leave a product throttled just as demand arrives. Committing too much converts optimistic forecasts into depreciation and financing costs. The winning strategy is unlikely to be “own everything” or “rent everything.” It will be a deliberately mixed portfolio of reserved capacity, flexible cloud supply, specialized inference providers, and software capable of shifting work among them.
That portfolio has to reflect workload shape. Training is large, episodic, and sometimes schedulable. Consumer inference is continuous and latency-sensitive. Enterprise batch work may tolerate queues. Agent workloads are bursty and difficult to forecast because a single user request can trigger many model calls and tools. Treating all tokens as interchangeable obscures these differences.
Utilization is now a product metric
Once infrastructure carries large fixed costs, utilization stops being solely an operations concern. Product choices determine whether expensive equipment performs valuable work.
A verbose interface can consume more inference without creating more utility. An agent that repeatedly inspects the same context ties up capacity while appearing industrious. Routing every request to the largest available model may improve a benchmark while destroying unit economics. Conversely, caching, batching, model routing, speculative techniques, and better stopping conditions can create economic capacity without adding a server.
This is why cost per token is an incomplete metric. Teams should measure cost per accepted outcome: a resolved case, a merged change, a qualified analysis, or a completed workflow. A cheap model that generates frequent rework can be expensive. A costly model that completes a high-value task in one pass can be economical. The denominator must describe customer value, not machine activity.
The annual Stanford AI Index tracks the widening commercial footprint of AI, but company-level planning still requires a more intimate dataset: request distributions, concurrency, latency tolerance, success rates, retry patterns, and the share of work that truly needs frontier capability. Firms with this telemetry can negotiate and architect intelligently. Firms without it are purchasing capacity on narrative.
The model roadmap and the capital plan must meet
Infrastructure decisions also alter organizational incentives. A company with substantial committed capacity is motivated to create demand for it. That can produce good pressure—faster product delivery and stronger distribution—but it can also encourage AI features whose primary purpose is absorbing supply.
Leaders should therefore separate three questions that are often bundled together. Is the feature useful? Does it require this class of model? Should the company own or reserve the capacity used to deliver it? A “yes” to the first does not guarantee a “yes” to the other two.
The accounting life of hardware further complicates the picture. AI accelerators may remain functional for years, yet their economically competitive life depends on improvements in newer chips, models, and inference stacks. A server does not need to stop working to become a poor place to run a workload. Good platforms preserve optionality by supporting multiple model sizes, quantization strategies, and hardware targets rather than welding the product to one generation.
Energy locality will matter too. The IEA’s analysis of AI-related electricity demand shows why global percentages can mislead: data-center load is geographically concentrated, so local grids and connection queues can become binding constraints even when worldwide supply looks ample. Capacity strategy must be regional, not merely global.
For AI businesses, the next durable advantage may not be exclusive access to a slightly better model. It may be the ability to translate uncertain demand into a resilient combination of contracts, hardware, power, and efficient software. That capability belongs jointly to product, engineering, finance, and operations.
The cloud taught software leaders to ignore the machine room. AI is teaching them to read its balance sheet.
Advertisement