In The Window we argued that token prices are subsidized and won’t stay that way. The pressure driving that shift is not abstract. It has two very physical components.
Two bottlenecks, not one
The power side is well documented by now. Roughly 2 GW of behind-the-meter gas generation is already operating in the US, most of it serving a single AI company near Memphis. The announced pipeline is near 100 GW. But announced is not built. Transformers take two to three years to manufacture. Advanced gas turbines are allocated through 2027 and beyond. The physical constraint on electricity supply is not a funding problem anymore. It is a time problem.
The chip side runs in parallel. The best frontier inference hardware is allocated months or years out. A new generation arrives roughly every 18 months. Demand, by the numbers being committed to it, is growing faster than that rhythm.
When supply in both inputs is capped and demand keeps compounding, token costs go up before they come down.
The only exit that doesn’t require waiting
Building more generation and more chips is the obvious answer, and it is happening at a scale the energy and semiconductor industries have not seen in a generation. A $35 billion private capital platform launched in June to fund more than 20 GW of AI compute through 2028. Three Mile Island is being restarted under a 20-year Microsoft contract. The buildout is real and it is accelerating. But building takes time. The question is what changes the cost curve before the new supply arrives.
The answer is efficiency.
Less power consumed for the same output, or more output for the same power. Both are happening at once.
Model architectures are getting sharper at doing more with less. The same inference task that required a large centralized model a year ago can now run on a smaller model on cheaper hardware. The compute required per useful output is falling, not because the ambition is smaller but because the engineering is better.
Open source accelerates this. When capable models are available without a usage fee and without a call to a cloud API, the economic incentive to run them locally becomes real. A company routing inference through a hyperscaler pays for every token. A company running a local model on its own hardware pays for the hardware once.
Both tracks grow
4.2 million Australian homes had installed 26.8 GW of distributed solar generation by mid-2025, enough to push midday grid prices negative. And the large centralized grid kept building. Distributed adoption did not replace the central system. It added a second system beside it.
AI infrastructure is likely to follow the same pattern. Centralized data centers with dedicated power generation will keep growing because the largest training runs and the most demanding inference tasks still require that scale. Local and edge compute will also keep growing because the economics of efficiency make it rational. Power generation as a sector is going to expand massively regardless of how much runs locally.
They are not competing outcomes. They are two responses to the same demand signal, operating at different layers.
The companies and infrastructure that win in the next few years are probably not only the ones that build the most centralized capacity. They are also the ones that extract the most output per watt. That is an efficiency race, and it is already underway.
The sector that looked like a pure scale problem is starting to look like an optimization problem too. Those tend to reward different skills.