Key Takeaways
- Legal tech startup Harvey watched gross margins drop from positive 50% to negative 50% in June after customer token usage jumped 20-fold.
- The unit economics collapse happened because reasoning models and multi-step agents burned massive token volumes across rented OpenAI and Anthropic endpoints.
- Instead of forcing customers into consumption pricing or downgrading answer quality, Harvey returned to positive margins in a single quarter.
- Harvey repaired margins by combining model routing, evaluation adjustments, and post-training on open-weight models like Harvey Tenet.
- As Jordi Hays observed, every application-layer AI startup faces margin shocks when agents scale, making fast routing changes a survival skill.
The 20x Token Explosion
In June, legal startup Harvey ran into a wall that every AI builder eventually meets. Token volume expanded 20-fold as customers adopted multi-step agentic workflows and advanced reasoning models. Because Harvey ran those workloads through rented OpenAI and Anthropic endpoints, its gross margins inverted, swinging from positive 50% down to negative 50%.
John Coogan outlined the problem on TBPN: “Harvey's gross margin fell from about 50% to negative 50% by June as agent token use spiked 20fold on rented open AI and anthropic models, Bloomberg reports.”
When unit economics drop that fast, most founders panic. They add hard usage caps, force customers into usage-based billing tiers before they understand the product, or route queries to cheaper, dumber models.
Harvey founder Gabe avoided that playbook. As Coogan quoted from Gabe's response: “The hardest thing about building Harvey is doing what's best for our customers despite pressure to do what's easy. The easy thing would have been to force our customers into consumption pricing before they were ready and serve them worse models to protect our margins.”
Fixing Unit Economics in Ninety Days
To pull gross margins back into positive territory within three months, Harvey rebuilt how it processed queries. Instead of sending every request to expensive closed-source frontier models, the team focused on model routing and post-training smaller, open-weight models like Harvey Tenet.
Coogan highlighted Gabe's approach: “This meant optimizing our product through routing... and post-training so we could serve frontier intelligence at an affordable price.”
By routing basic subtasks to post-trained open models and reserving high-end frontier APIs for complex legal reasoning, Harvey reduced token expenses without hurting the user experience.
The turnaround highlights a central debate for vertical software builders: should you fine-tune open models, or will falling API prices solve the margin problem on their own? Hays framed the skeptic position: “doesn't make sense to train your own models because the public basically publicly available frontier models are dropping in cost and the quality is increasing. So, you're just wasting your own and again that's a bet that he's making.”
Waiting for frontier API price drops is a gamble when active burn is bleeding cash. As Hays noted, “every application layer company has gone through a moment like this even if it was for a single day. So it matters how quickly you can respond and adjust.”
What to Do With This
Audit your application's token consumption by subtask. Identify the top 20% of repetitive prompts that burn 80% of your frontier model budget, and set up a basic router to direct those queries to a smaller open-weight model this week.