Key Takeaways
- Variable token consumption breaks traditional enterprise IT budgeting because model queries do not follow predictable, fixed-seat SaaS contract pricing.
- High-profile engineering organizations have misjudged consumption velocity; Uber exhausted its entire 12-month artificial intelligence budget inside of four months.
- Mid-market companies face immediate margin erosion from unmonitored model usage, including one services firm that generated a single monthly Claude bill in the hundreds of thousands of dollars.
- Traditional annual budget cycles fail under consumption models, requiring finance teams to institute monthly or quarterly reforecasting cycles alongside hard token caps per user.
The Uncapped Meter in Modern Software Budgets
Traditional enterprise software procurement operates on predictable rails. Finance teams negotiate annual seat licenses, model out a 5% to 10% annual price uplift, and lock the line item into the budget. Variable token consumption destroys this predictability. When engineers and business units plug large language models directly into production workflows without consumption caps, the budget turns into an open-ended utility bill.
Even tech-forward enterprises fail to forecast this drain. Justin D'Onofrio points to Uber burning through its entire annual artificial intelligence allocation in four months: “The one that was going around the past couple months has been Uber spending their full year AI budget in four months. We are talking about a sophisticated engineering forward organization not being able to size the prize with AI.”
If top-tier tech firms with dedicated infrastructure teams struggle to project consumption rates, mid-market private equity portfolio companies face far steeper downside risks. Unchecked experimentation without cost monitoring creates immediate balance-sheet surprises.
Six-Figure Surprises in the Middle Market
The issue is not limited to tech giants testing frontier models at massive scale. Portfolio companies in standard services sectors are seeing cloud invoices explode overnight. In one instance, a services firm approached Accordion after discovering an unexpected six-figure monthly model invoice.
“Services company came to us pulled up a Claude bill from a recent month,” D'Onofrio explains. “It was many hundreds of thousands of dollars and they said what is going on with this bill?”
This gap occurs because engineering and product teams can spin up API access keys in minutes, tying business processes to external foundation models without establishing rate limits. D'Onofrio emphasizes that AI spend requires a dedicated operational owner who sets hard spending ceilings on individual seats and organizational units: “If you think about your budget line items for where you should invest, one is definitely that personal productivity, but again, cap it and have an owner associated with making sure that that spend does not get out of control.”
Moving FP&A to Continuous Consumption Tracking
Static annual budgets cannot protect margins when underlying infrastructure costs fluctuate on a per-query basis. If a company implements an automated customer support bot or an internal research agent, usage spikes directly increase cost of goods sold or operating expenses in real time.
Controlling this volatility requires FP&A teams to shift from static annual cycles to rapid cadences. D'Onofrio outlines the required shift: “If you are not doing a quarterly reforecast or a monthly reforecast specific to those AI spend buckets, you need to be, because the cost can balloon, changes can happen quickly in the value capture.”
Without structured consumption governance, productivity gains get erased by backend infrastructure bills.
Why It Matters
For private equity sponsors, uncapped AI spend poses a direct threat to projected EBITDA margins and exit multiples. When portfolio companies deploy autonomous agents or LLM-driven features without unit-level cost tracking, gross margins compress before operating partners even detect the shift. Sophisticated deal teams are beginning to treat token governance and compute commitments as standard technical due diligence items rather than minor operational details.