Key Takeaways

  • Pre-training compute collapsed from roughly 50% of frontier GPU workloads down to 10%, shifting capital toward reinforcement learning and post-training data curation.
  • Frontier labs now treat model training as an engineering optimization race rather than scientific discovery, pouring cluster budgets into hill-climbing algorithms.
  • NeoClouds and third-party inference providers face razor-thin margins as commoditized hardware access turns token generation into a low-yield utility.
  • Nathan from Air Street Capital observes that the bottleneck for artificial intelligence has shifted from algorithmic discoveries to capital reserves, energy permitting, and debt mechanics.

The Collapse of Exploratory Pre-Training

For five years, the default playbook in frontier artificial intelligence was simple: gather as much text as possible, buy larger clusters, and let pre-training run. That era is over. Nathan from Air Street Capital points out that foundational pre-training compute has dropped from roughly 50% of total GPU utilization down to just 10%.

Labs are not spending less money overall. They are reallocating billions into post-training. The low-hanging fruit of raw web scraping has been harvested. Now, teams run reinforcement learning loops where models generate responses, get scored, and adjust their internal pathways repeatedly.

As Nathan explained, “It strikes me as normal that one would end up spending less money on exploratory work because the solution is more like in the plain eye.” John Coogan framed the commercial mindset behind this shift directly: “I want the solution bad enough. I'm willing to spend a ton of money on RL to try to hill climb to get there.” Instead of guessing what a raw foundation model might pick up, teams spend their compute budgets forcing models to master exact reasoning tasks.

The NeoCloud Margin Trap

While frontier labs direct their clusters toward reinforcement learning, companies providing rented chips and raw inference face a brutal market reality. Specialized cloud providers and inference wrappers took on massive equipment debt to buy hardware, expecting high software margins. Instead, they walked into a commodity trap.

Nathan summed up the dilemma facing these infrastructure operators: “You either die getting to the frontier or you live long enough to serve inference.”

Serving inference offers almost zero pricing power. When every cloud can run the same open weights, buyers shop purely on cost per million tokens and latency. If your sole business model is reselling rented H100 capacity, you are competing against subsidized hyper-scalers who treat compute as a loss leader. The margins are thin, the hardware depreciates rapidly, and debt service continues regardless of customer churn.

The Financialization Era

The frontier is no longer governed primarily by novel whitepapers written by academic researchers. It is governed by project finance, electrical substation permits, and capital allocation.

Nathan highlighted this transition directly: “We're now in the financialization stage of AI. We understand the recipe, the demand is there, we just need to deliver it in more efficient ways and get more people to understand how to get value from it.”

Progress currently depends on who can secure gigawatt-scale power purchase agreements and manage balance sheet risk. The core recipes for building and refining weights are largely understood across top labs. The winners will not be the teams hoping for a sudden algorithmic miracle, but the organizations that treat compute as a high-volume manufacturing operation.

What to Do With This

Audit your cloud spend this week and divide it into raw token calls versus proprietary evaluation data. Stop spending money fine-tuning base models from scratch if you lack an internal reinforcement learning pipeline to verify the outputs. Shift your engineering hours toward building deterministic verification loops that grade model answers, because verified domain feedback is where actual software defensibility lives.