Key Takeaways
- AI models are diverging rather than commoditizing because the majority of enterprise GPU budgets now go toward custom post-training rather than foundation pre-training.
- In automated financial research, context retrieval accounts for roughly 80% of compute costs and 70% of final output quality.
- Frontier models struggle with institutional finance queries because their training defaults to wide, high-volume web search patterns rather than targeted data retrieval.
- AlphaSense runs open-source models on Cerebras hardware to cut inference latency by 10x to 20x while dropping research loop costs by up to 40x.
The Post-Training Divergence
Many technology analysts spent the last eighteen months predicting that foundational models would become cheap, interchangeable utilities. That thesis is falling apart in production environments. The split is happening because the economics of model development have shifted from raw pre-training runs to task-specific post-training.
“Now, the majority of the GPU spend for these firms is not pre-training, it's post-training,” Ackerson explains. “In post-training, you can apply a lot more differentiation in terms of what tasks you're post-training on, what data sets you're acquiring to post-train on. So these models have actually differentiated, which is kind of contrary to a lot of predictions that thought they would all sort of commodify and converge.”
Foundation labs spend billions teaching models general reasoning across public internet scrapes. But general reasoning does not tell a model where to locate historical credit agreements, how to parse SEC filings, or how to cross-reference private transcript databases. By directing compute budgets toward post-training on proprietary document sets, vertical software vendors build domain moats that raw foundation models cannot match.
The 80% Bottleneck in Financial Search
Frontier reasoning models look impressive in conversational interfaces, but their underlying retrieval mechanics break down in specialized financial workflows. When a general model attempts research, it behaves like an open-ended search engine: it queries dozens of generic sources, brings back noisy context, and floods the context window with irrelevancies.
“The frontier LLMs, the way that they're trained to search is very much sort of web search frame where you just kind of spray and pray. You just search a ton of stuff,” Ackerson notes. That wide retrieval loop creates severe cost and latency issues in deal workflows.
If the retrieval phase fails to isolate the exact table or disclosure immediately, the downstream synthesis model hallucinates or inflates compute costs. To solve this, AlphaSense trains small, specialized routing models that pinpoint the exact internal index needed for a specific task. By pairing these domain-specific models with dedicated hardware from Cerebras, they achieve 10x to 20x lower latency and slash operational research costs by up to 40x compared to frontier LLM APIs.
Why It Matters
This dynamic changes how investors should underwrite vertical software companies. The defensibility of enterprise AI applications does not lie in building bigger foundation models or creating thin wrappers on top of OpenAI. Real margin expansion and technical moats come from proprietary data routing, specialized post-training, and infrastructure optimization that dramatically lowers inference costs on complex research tasks.