Key Takeaways
- Relying exclusively on closed commercial AI APIs caps gross margins between 20% and 30%, while post-training open-weights models pushes margins above 80%.
- Training foundation models from scratch to top public leaderboards is a distraction that fails to reflect actual production video workflows.
- Higgsfield dynamically routes over 40% of user requests to specific cost-effective architectures rather than defaulting to expensive frontier models.
- Video models operate like modern rendering engines that require thousands of words of prompt detail and reference assets, not simple text-to-video queries.
Why Chasing Leaderboards Destroys Margins
Higgsfield scaled to a $1 billion annualized revenue run-rate in 18 months while spending over $4 million per month on internal compute. Yet founder Alex Mashrabov argues that building proprietary foundation models from scratch is a trap.
For two years, early-stage AI startups treated public leaderboards as validation. Founders raised large rounds to train base architectures from zero, aiming to beat benchmark evaluation scores. Mashrabov rejects that strategy: “chasing benchmarks is valuable but I don't believe this is just sort of corporate scops frankly.”
Public benchmarks measure one-shot output against generic prompts. Real video production demands something entirely different. “A lot of benchmarks today for video is really text to video which does not represent actual workflows at all,” Mashrabov says. “The way to think about video models today, it's just modern rendering engine.” Real users submit technical camera directions, character consistency constraints, and multiple reference images. A model tuned only for generic prompt leaderboards fails inside real creative pipelines.
The Post-Training Margin Advantage
The choice between foundation training and post-training comes down to basic unit economics. Reselling closed commercial APIs leaves founders with software margins that look more like physical retail. “The margin on own models and open weights models is over 80%,” Mashrabov explains. “And then it almost doesn't matter. And for closed source models, it's probably between 20 and 30%.”
At 25% gross margins, an AI company cannot sustain long-term customer acquisition costs or support engineering hubs like Higgsfield built in Kazakhstan. At 80% margins, the business generates enough free cash flow to absorb infrastructure shifts.
Instead of starting from zero, practical AI companies treat open weights as a base layer. “Whenever someone says we build our own models very likely what they mean is something what see what's happened with Cursor,” Mashrabov points out. “A lot of companies they actually take open weights model and just post train on own data.”
To preserve those margins at high volume, Higgsfield implemented dynamic model routing. The application layer inspects each user request before firing compute. Mashrabov notes: “we choose which model we can use. So like we choose what model to use in over 40% cases.” Routing structured requests to fine-tuned open models lets Higgsfield keep closed commercial models reserved strictly for edge cases.
What to Do With This
Audit your inference spending across the last 30 days. Pull the top 20% of repeated API prompts currently running on closed commercial endpoints, fine-tune an open-weights model on that exact prompt-response dataset, and deploy it to dedicated GPUs. Once quality matches your baseline, route those production requests internally to immediately raise your product gross margin toward 80%.