Key Takeaways

  • While hyped, AI model routing is a genuinely difficult technical problem that most enterprises are ill-equipped to solve internally.
  • Market forces will make AI inference dramatically cheaper, with Arena CEO Anastasios predicting eventual pricing pressure on currently "disgustingly high" gross margins.
  • When Anthropic goes public, its high inference margins will be exposed, acting as a major catalyst for cost reductions across the industry.
  • AI businesses built on reselling tokens or GPUs (GMV models) face severe long-term viability issues due to low terminal value and unsustainable public company margins.
  • Founders should avoid building their own routing layers and instead focus on core product value, assuming a rapid commoditization of inference costs.

The Model Routing Gold Rush is Real, But Harder Than You Think

Forget the buzz for a second. Anastasios, CEO of Arena, cuts through the noise on AI model routing: there's real value there, but it's not a trivial problem. “I absolutely think there's value in the model routing layer,” he affirms, acknowledging the current "hype cycle" but emphasizing the underlying technical challenge. Routing isn't just sending a request to the cheapest model. It demands a deep understanding of the query's nature, its difficulty, and how to rapidly onboard new models, often beyond what individual enterprises can manage.

He puts it bluntly: “Routing is a very difficult technical problem.” Think about it. You need to intelligently direct a query to the best model, not just any model. That means assessing the query's domain, its complexity, and the specific strengths and weaknesses of a constantly evolving roster of models. Doing this at scale, with low latency, and adapting to new model releases daily is a specialist's job, not something a product team casually bakes in.

Your AI Inference Bill is About to Drop

If your startup’s financial model relies on today’s AI inference prices, prepare for a shock. Anastasios expects a sharp drop. “I definitely think in the long run, the market will be efficient. Things will get cheaper,” he says, pointing to the inevitable downward pricing pressure as the AI market matures. He singles out one major player: Anthropic. Their current “disgustingly high gross margins in their inference” are a temporary luxury. Once they go public, those numbers become transparent.

“After they go public, the whole world is going to see that,” Anastasios warns. This exposure will not only invite intense competition but will also set a new standard for pricing. When the public sees how much margin leading AI labs are making on inference, it will “exert downward pricing pressure on their inference” across the board. This isn't just a prediction; it's a market dynamic almost guaranteed to play out once the financial details of these private giants become public information.

Why Most 'AI Resellers' Will Get Crushed

Beyond just dropping inference costs, Anastasios flags a deeper structural problem for many emerging AI businesses: the "GMV model." These are companies fundamentally built on reselling access to AI models or GPUs. Think about it like a reseller who just adds a small markup on someone else's product. While they might show high Gross Merchandise Volume (GMV), their actual margin and, critically, their terminal value, are shaky.

“The bigger problem with businesses like that I see these days is that a lot of them are fundamentally GMV businesses where there's like some reselling happening,” he notes. These are "tough" businesses because their economic value isn't truly proprietary. Without unique intellectual property or a deeply defensible moat, they struggle to achieve the kind of "sustainable public company margins" that investors demand. If your core business is just reselling someone else's compute or tokens, what happens when those costs plummet, or the original provider decides to sell direct? You're left with little to no competitive edge.

What to Do With This

If you're building an AI product, stop trying to build a complex model routing layer yourself unless that is your core business. Focus your engineering talent on proprietary features and leverage specialist routing solutions. Second, immediately re-evaluate your financial models and pricing. Assume AI inference costs will fall by 50% or more in the next 18-24 months; price for value, not just today's compute costs. Finally, if your current business model hinges on reselling AI tokens or GPUs, pivot now. Find a unique data moat or specific application IP, or you'll be squeezed out as commoditization accelerates.