Simon Eskildsen runs a database company, turbopuffer, that lives entirely on CPUs rather than GPUs. Asked recently by Gergely Orosz how easy it still is to get CPU capacity from AWS, GCP or Azure, his answer was blunt: it isn’t. Reinforcement learning training needs CPUs to run the software the model is practising on, agents outside training need CPUs for everything general-purpose, and the two demands are now colliding in the same queue. Even companies with cash to spend and long lease terms on offer are being told there’s nothing left to rent.
That is the opposite of the story most people have been telling themselves about AI for the past two years. AI isn’t free anymore, and the shortage behind it is CPUs, not the GPUs and memory everyone has spent two years watching.
The free-AI story was a snapshot, not a rule
AI is essentially free was a reasonable thing to write in 2025. Frontier labs were burning venture money to win market share, providers cut prices to build usage habits, and a curious professional really could learn agentic tooling for the cost of an afternoon. It was the same logic behind the open-weight default: if capability was going to keep getting cheaper and more available, hoarding it made no sense. That framing wasn’t wrong. It was a description of a particular moment: oversupply during a land grab. But land grabs end.
What has ended it isn’t more demand for the thing everyone was watching: GPU training runs. It is a shortage nobody was noticing: CPUs. Cloud providers used to discount idle CPU capacity by up to 90% through spot pricing, because plenty sat unused. That discount has effectively disappeared. Reserving specific CPU types now needs months of lead time, and providers are turning down reservations outright when they don’t have enough of the right hardware to spare.
Katelyn Lesse, who runs platform engineering for Anthropic’s Claude Platform, has traced the mechanism back to the factory floor: GPUs and CPUs compete for the same production lines at TSMC, high-bandwidth memory for AI accelerators competes with ordinary DRAM at Samsung, SK Hynix and Micron, and AMD’s CPUs, with no fabs of its own, sit at the back of TSMC’s queue behind the chips getting the headlines. Intel owns its fabs but has been diverting capacity from PC chips to server chips just to keep up. None of this can be solved in weeks. Analysts expect it to take multiple quarters (years!) before CPU supply gets comfortable again, and memory will take longer.
Why the shortage is agentic, not just AI
The 2023–24 GPU shortage was about training bigger models. This one is about running them, over and over, as agents. A single agent doing a general-purpose task, browsing, running code, calling tools, checking its own output, spends most of its time on a CPU, not a GPU. Multiply that by every team that has gone from an engineer prompts a model to an agent runs unattended for hours, and the CPU bill stops looking like a training cost and starts looking like a headcount cost.
That distinction matters for anyone still budgeting AI the old way. Model access is still cheap, in the sense that a subscription or an API key gets you extraordinary capability for a small fee. What’s expensive now is running that capability continuously, at the scale agentic workflows demand: the compute behind long-running agent sessions, the tokens spent when one agent’s output becomes ten other agents’ input, and the review time a human still has to spend checking what came out the other end. Call it the integration tax. It’s the same cost that shows up as prompt debt inside a single system. It has always been there, hidden behind free compute until the bill finally arrives.
What’s still free, and what never was
Learning to prompt a model well is still close to free. Reading the documentation, running your first agent, understanding what these tools can and can’t do: none of that costs meaningfully more than it did a year ago, and accepting the “I haven’t tried this yet” excuse still doesn’t hold up. That part of the free-AI argument survives intact.
What doesn’t survive is the leap from learning it is free to using it at scale is free. Those are different claims, and 2025’s oversupply blurred the line between them. A product leader who scoped a roadmap assuming inference cost would keep falling is now discovering the opposite curve: costs rising. Because the workload changed shape.
The budgeting question that matters
This is not an argument against building with AI. It’s an argument for building the right thing at the right scale, and pricing conservatively rather than extending 2025’s assumptions into a 2027 plan. Before greenlighting the next agentic workflow, ask what it runs on hour by hour, not just what the model call costs per request. The model will not be the most expensive part. The agent will, running unattended all day, on infrastructure everyone assumed would stay cheap.
AI was free while the industry was buying market share with someone else’s money. It stopped being free the moment agents became the customer for compute, not just the tool for writing it. Budget for the compute you’re actually using, and prepare for when the land grab comes to an end and models’ prices rise as well.
Leave a Reply