Key Takeaways
- Frontier labs like Anthropic and OpenAI used to rent GPU clusters in blocks of 8,000 chips; now they aggressively secure slices as small as 1,000 GPUs.
- Inference platforms including Baseten, Fireworks, Together, and Morph are generating 60% gross margins on open models with vLLM and SGLang optimizations.
- Startups like Modal now monetize server blocks as small as four nodes, eliminating spare capacity across cloud providers.
- New research labs trying to train models cannot secure spot instances because cloud hosts favor cash-flowing inference businesses.
The Squeeze on Small GPU Slices
For years, frontier AI labs stayed in their lane. OpenAI, Anthropic, and their peers barged into data centers and demanded massive clusters of 8,000 GPUs or more. That left the remaining smaller slices for early-stage teams and research experiments.
That buffer is gone. Frontier labs are now sweeping up smaller allocations to meet demand. SemiAnalysis founder Dylan Patel notes that market access has compressed dramatically: "It's the toughest it's ever been to rent compute. Obviously Anthropic and OpenAI have been buying smaller and smaller cluster sizes. They used to only go 8k GPUs from a provider. We've even seen as small as 1,000 GPUs get rented by the labs now."
When frontier labs bid on 1,000-chip allocations, they swallow the exact capacity that mid-tier companies used to count on.
Why 60% Gross Margins Broke Cloud Availability
The bigger shock comes from the inference layer. Running models for end users used to be seen as an expensive loss leader. Today, specialized hosting startups are printing cash.
Patel points out that software optimizations have turned inference into a high-margin cash machine: “Baseten, Fireworks, Together, and many other companies, even small companies like Morph, are running at 60% gross margins on open models using vLLM or SGLang, using just a little bit of optimization beyond that, and able to do an awesome job.”
Because serving queries pays out immediately, inference platforms will gladly lease any loose hardware on the market. As Patel explains: "Anyone who can get access to a GPU can immediately make money. It is challenging for anyone to get any sort of compute if Modal can make money hand over fist off of even four nodes from a company."
When a four-node cluster can instantly turn a profit serving inference calls, no cloud provider leaves machines sitting on the spot market.
Research Labs Are Priced Out by Unit Economics
This creates a harsh reality for early-stage research teams. If you run a lab trying to pre-train a new architecture, you burn cash for months with zero incoming revenue until training finishes.
Cloud operators know this. Given the choice between a research startup burning venture capital and a company like Modal that generates cash on day one, data centers choose the paying inference workload every time.
Patel frames the builder's dilemma plainly: “Ultimately it's like, well great, I'm a new lab who wants to train an interesting cool new model. Well, I'm also competing against people who make money now.”
What to Do With This
Audit your compute roadmap this week. If your startup relies on renting bare-metal GPU clusters for custom pre-training, stop assuming spot pricing or short-term leases will exist when you need them. Shift your architecture to fine-tuning on top of open weights, and route your production traffic through specialized inference providers running vLLM or SGLang to lock in positive unit economics before you scale.