Key Takeaways

  • Startups running on frontier models have zero pricing power against foundation model suppliers who also build competing user-facing products.
  • Underlying compute costs fall roughly 50% every 9 to 12 months, allowing founders to price for 20% to 30% gross margins 18 months out without manual prompt compression.
  • Spending scarce engineering hours on token routing instead of data integrations or network effects burns early distribution advantages.
  • Uncapped token maxing distorts customer behavior by encouraging wasteful compute usage instead of proving concrete return on investment.

The Supplier Trap at the Frontier

Grèze warns that early AI application builders face a brutal structural squeeze. If your product relies on top-tier reasoning capabilities, you are trapped buying from suppliers like OpenAI or Anthropic where you have zero pricing power. Worse, those same suppliers are building applications to take your customers.

Cursor serves as the clearest cautionary tale. “You can have huge market share and be customers love you and everything,” Grèze explains, “but if you're paying your suppliers and competing with your suppliers at 70% margin, eventually it gets like a little bit difficult.”

While routine tasks like email tagging quickly drop to cheap open-weight models, high-value enterprise trajectories remain anchored to the frontier. When your supplier captures a 70% margin on the underlying tokens and rolls out a competing client, application layer margins collapse unless the software owns proprietary context.

Why Early COGS Optimization Kills Growth

Founders often panic when they see their monthly token bills, immediately assigning top engineers to shrink prompts and route queries. Grèze considers this a major strategic mistake.

“If I have an engineer hour, what is the engineer hour best spent on?” Grèze asks. “Is it taking my current AR making it more efficient? Or is it figuring out a way to grow the product faster by working on a better network effect feature or making the model better and integrating with a new data source that makes its trajectories much better for a set of our users.”

Because base model efficiency improves rapidly, hardware deflation does the cost cutting for you. “You know where it's going to be and so that you could price your product today at a point where you'll generate 20 30% margins in 18 months.” Engineering hours spent shaving 15% off API calls today are hours stolen from locking in distribution before incumbents arrive.

The Danger of Token Maxing

Subsidizing infinite token usage to inflate engagement metrics creates a false signal of product-market fit. When enterprise software encourages unconstrained usage, it trains buyers to ignore unit costs.

“The problem with token maxing, there's two problems in my opinion,” Grèze says. “One is like companies told people, hey, you can use as much money as you want on AI, which is bad. You want people to think is the ROI of using AI here worthwhile.”

Instead of encouraging customers to burn tokens on low-value tasks, products must tie software usage directly to tangible workflow output. If a customer cannot justify the return on investment of a run, letting them run it for free only hides churn that will appear the moment you enforce realistic pricing.

What to Do With This

Audit your product backlog this week. If any engineer is scheduled to build prompt compression tools or multi-model fallback routers, cancel those tickets. Reassign that engineer to integrate another internal enterprise data silo that improves multi-step trajectory completion rates.