Key Takeaways
- AI token spend is no longer a negligible cost for businesses; for some Ramp customers, it now rivals 10% of their entire payroll expenditure.
- Ramp customers saw their AI token expenditures grow 21x in the last year alone, signaling the definitive end of the 'Uber era of AI' where usage was cheap and often went unnoticed.
- The new product,
token-spend.fm, highlights a crucial shift: AI costs require active management, much like cloud infrastructure, with dedicated tools for visibility, alerts, and optimization. - A key cost-saving strategy emerging from this data is the migration of repetitive, low-stakes AI tasks from expensive frontier models to cheaper, open-source alternatives.
The Method: Uncovering the Hidden AI Tax
For years, AI API calls felt like free money, a trivial rounding error in the company budget. Not anymore. Eric Lyman, co-CEO of RAMP, put it starkly: “A few years ago, AI spend was a routing error... In May, it hit almost 10% of our payroll spend. The equivalent was on tokens on payroll.” This isn't just about big companies; it hits smaller, ambitious builders too. Their new product, token-spend.fm, isn't just a dashboard; it’s a systematic approach to turning an invisible cost into an actionable line item.
The method centers on four critical steps:
1. Gain granular visibility: You can't manage what you can't see. The first step is to track every token spent across every model and every user. This means moving beyond aggregate API bills to understand who's using what and for what purpose. For example, some users might be generating expensive text summaries when a cheaper, local model could do the job.
2. Alert on anomalies: Once visibility is established, the system flags unusual spikes or high-cost patterns. “It's like, I didn't realize I spent $100 to check the weather,” Lyman explained. These alerts provide real-time feedback, helping teams catch runaway costs before they become substantial problems.
3. Identify optimization opportunities: The data reveals patterns. Are teams using expensive frontier models for repetitive tasks that don't demand cutting-edge performance? Ev Randall, also on the podcast, reinforced this point: “We're seeing an immense amount of app companies... 80% plus um open-source usage... or like 90% of them are trying to get to that ratio plus.” This insight drives the shift towards open-source models for specific, high-volume tasks.
4. Empower with cost awareness: The goal isn't just to cut costs from the top down. It's to make every engineer and product manager aware of the real-time cost of their AI usage. Just as cloud dashboards made engineers mindful of compute spend, token spend management aims to embed cost-consciousness into AI development cycles.
Where This Breaks Down
This method shines brightest for companies with distributed AI usage and a rapidly escalating spend. If your startup has only one or two engineers making a handful of API calls a day, the overhead of implementing a complex tracking and optimization system might outweigh the current savings. The gains from shifting to open-source models also depend on your internal engineering capacity. If your team can't easily integrate and manage open-source solutions, relying on frontier models might still be the most practical (if not cheapest) path for now. This approach also assumes that AI usage isn't purely exploratory; if your team is still in the experimental phase, optimizing for cost might prematurely stifle innovation.
What to Do With This
Stop treating AI API calls as an invisible cost. This week, pull your company's last three months of AI-related expenses from your cloud provider or directly from API invoices. Categorize the top three models consuming the most tokens. For each of those, identify one high-volume, repetitive workflow that relies on it. Next, spend 2-3 hours researching if a smaller, fine-tuned, or open-source model could handle that specific task for significantly less. Run a small-scale test to compare output quality and cost, making a data-driven decision about where to optimize.