Key Takeaways
- Inference infrastructure providers are seeing massive valuation jumps, with Modal tripling to a $15 billion valuation and BaseTen discussing a $26 billion valuation.
- Jack Altman describes "intelligence saturation" as the inflection point where cheaper open models become good enough to handle specific production tasks reliably.
- Jason Lemkin argues enterprise adoption of open-weight models has already peaked and will decline as frontier labs slash API prices and enterprise buyers demand strict security.
- Compute ownership remains the decisive factor: providers that control dedicated compute capacity hold structural pricing power over pure software layers.
The Disagreement
For the past 18 months, buying into inference platforms was the simplest way to bet on AI growth. As Altman observed, “And it has now turned out in the last 18 months or whatever, the correct answer was just keep buying inference. You have modal, base 10, fireworks, foul together, and it's just all worked.”
Where Altman and Lemkin split is on what gets run on those inference engines over the next two years.
Altman points to a phenomenon he calls intelligence saturation. In his view, “we have now crossed sort of the threshold in a lot of areas and more and more happening where you get sort of like intelligent saturation where it is now good enough to do the thing.” When a smaller, cheaper open-weight model can accurately extract data, write code, or route customer tickets, paying frontier lab prices makes little financial sense. Companies will run open weights on specialized inference clouds like Together, Fireworks, BaseTen, or Modal.
Lemkin sees the exact opposite trend inside enterprise sales cycles. He argues open weights have hit a hard ceiling: “I think open source has reached its maximum as a market share. I think it's going to keep going down.” When talking directly to enterprise buyers, Lemkin sees zero appetite for self-hosted or open-weight models: “no one wants to use an openweight model on the floor that I talked to nobody so I just think the market share is peaked.”
From Lemkin's perspective, three forces are squeezing open source: price cuts from proprietary labs, legal anxiety around foreign-trained weights, and the convenience of managed closed APIs.
Who's Right (and When They're Wrong)
Altman is right about task economics, but Lemkin is right about enterprise procurement psychology.
If your product runs internal workflows, high-volume batch processing, or structured data extraction, Altman's intelligence saturation thesis holds. Once a lightweight model scores 98 percent accuracy on your specific prompt, paying ten times more for a frontier model burns cash. Inference routing platforms win here because developers want speed, predictable latency, and margin control.
However, if you sell software to Fortune 500 CISOs, Lemkin's warning is spot on. Enterprise legal teams still hesitate to approve non-US weights or complex open deployments. When proprietary labs cut their API prices by 80 percent every nine months, the margin savings of self-hosting open models rarely offset compliance headaches and engineering maintenance.
The final variable is compute ownership. Altman noted that “on some level this will also just come down to like all of the compute is firing all the time and like who owns it, I think is going to turn out to just be a dominantly important part of the equation.” If frontier labs own the physical clusters, they can undercut third-party inference providers whenever competition heats up.
What to Do With This
Audit your model spend this week. Break your AI queries into two buckets: complex reasoning and routine task completion. Move every routine query with consistent inputs to a dedicated inference platform running a smaller model, and reserve closed frontier APIs exclusively for ambiguous, customer-facing reasoning.