The software industry trained investors and operators to admire businesses that add customers without adding much physical plant. AI is breaking that reflex. The product may arrive through an API, but behind the endpoint sits an expanding estate of accelerators, substations, cooling systems, fiber, land, and long-term power commitments.
This is not a temporary inconvenience on the way back to familiar software economics. It is becoming a defining feature of the market. Microsoft has said that scaling AI infrastructure has pressured cloud gross-margin percentage, while its fiscal 2025 fourth-quarter discussion separated spending on long-lived assets from spending on servers. Alphabet’s 2025 fourth-quarter earnings call likewise described technical infrastructure investment alongside rising depreciation and data-center operating costs.
The crucial shift is conceptual: AI providers are no longer merely pricing software features. They are allocating scarce industrial capacity among training, consumer products, enterprise inference, internal productivity, and speculative future demand. Every model response carries a hidden capital-allocation decision.
The denominator matters more than the headline spend
Capital expenditure attracts attention because the numbers are large. But expenditure alone reveals surprisingly little. A costly cluster running valuable workloads around the clock may be a better asset than a cheaper cluster stranded by power constraints, networking bottlenecks, poor scheduling, or weak customer demand.
The strategic metric is productive utilization: how much commercially useful computation an operator extracts from the entire system over its economic life. That denominator includes more than accelerator occupancy. It includes memory capacity, interconnect efficiency, energy availability, model architecture, batching, software optimization, and the ability to redirect hardware as demand changes.
This is why the strongest infrastructure businesses will not necessarily be those that purchase the most chips. They will be the ones that turn heterogeneous demand into smooth utilization. Consumer traffic can fill one part of the day, batch workloads another, and latency-sensitive enterprise inference can occupy reserved pools. Model routing can send routine requests to smaller systems while preserving expensive capacity for work that earns its cost.
Seen this way, a model-efficiency improvement is not merely a research achievement. It is a financial event. Better quantization, caching, speculative decoding, or workload scheduling can release capacity without pouring concrete. Conversely, an architecture that wins a benchmark while requiring uneconomic serving resources may create prestige without durable margin.
Depreciation is where yesterday’s optimism meets today’s income statement
Infrastructure spending arrives in cash before its full accounting cost appears in earnings. Servers and networking equipment then create depreciation over subsequent periods, while power and operations remain recurring expenses. That delay can produce a flattering interval in which demand narratives accelerate faster than recognized costs.
The risk is not simply that AI demand disappoints. Demand can grow rapidly while economics deteriorate. If customers expect continual price reductions, models become more compute-intensive, and hardware generations turn over quickly, revenue growth may coexist with weak returns on installed capital. An asset can be technically busy and financially underproductive.
Hardware obsolescence further complicates the familiar data-center model. Buildings and electrical infrastructure can serve for decades; accelerators cannot be assumed to retain frontier value for anything close to that period. Operators therefore manage a layered portfolio of asset lives. The shell may be durable, the cooling design less so, and the compute inside it subject to abrupt changes in performance per watt.
OpenAI’s expanding Stargate infrastructure program makes the physical scale of frontier ambitions explicit. Projects of that magnitude are not ordinary SaaS capacity planning. They involve energy markets, construction timelines, financing partners, local politics, and supply chains. Those dependencies create defensibility, but they also create obligations that cannot be resized with a configuration change.
Buyers should interrogate the cost curve
Enterprise customers do not need to audit a provider’s data centers, but they should understand how infrastructure economics shape product behavior. A low introductory API price may be strategic rather than sustainable. Generous context windows may carry rate limits under heavy demand. A feature that depends on the largest model today may later be routed, compressed, repriced, or withdrawn.
Procurement teams should ask questions that sound more like capacity planning than conventional software selection:
- Can the workload move among models without being redesigned?
- Which quality level is actually required for each request?
- How does the provider handle capacity shortages and priority tiers?
- What happens to unit cost as context, tool use, and agent runtime expand?
- Can cached or asynchronous work receive materially different economics?
The same discipline belongs inside AI product companies. Gross margin should be measured by workflow and customer segment, not hidden behind an average token price. A research agent that runs for forty minutes has a different capital signature from an autocomplete request. Bundling them under one subscription can obscure which behavior creates value and which merely creates impressive usage.
AI will still produce excellent software businesses. But the winning operators will think like both software architects and infrastructure financiers. They will treat utilization as a product capability, model routing as margin engineering, and depreciation as strategic feedback. The interface may remain weightless; the balance sheet will not.
Advertisement