Key Takeaways

  • OpenAI's Luna Terasol model now presents a highly competitive offering, often outperforming open-source alternatives when considering real-world cost-per-task, not just raw token prices.
  • Founders should discard cost-per-token as their primary metric. The true measure of AI efficiency lies in cost-per-task, which accounts for the entire workflow and output quality.
  • A model's 'harness'—the specific API integration settings and configurations—can dramatically alter its performance, as demonstrated by OpenAI tripling ARC AGI V3 scores with 6x fewer tokens by adjusting just two settings.
  • Public benchmark scores for AI models, like ARC AGI V3, are highly dependent on the integration and testing environment, meaning founders cannot blindly trust them for their specific applications.

The Hidden Lever of AI Performance

Forget the headline-grabbing model releases for a moment. The real battlefield for AI efficiency isn't just about the raw power of the underlying model; it's increasingly about how you integrate it. John Coogan observed OpenAI's recent moves as “pushing the model frontier access across efficiency.” What does that mean? They dropped the cost of their Luna Terasol model. Jordi Hays backed this up, saying, “This is the cheapest model. Yes. Massively reduced cost. You can see on the kind of Yeah. Pluto curve. This is like actually much cheaper than a lot of like open source models.”

This price reduction makes Luna Terasol highly competitive, but the deeper insight is about how founders should truly measure value. Coogan quickly clarified, “there you could measure it on on cost per token, but if if a certain model takes 10 times the amount of tokens and it's only half the cost, you wind up spending more.” This simple truth—that cost-per-task beats cost-per-token—is a fundamental shift in how to evaluate AI tools.

But the bombshell came from OpenAI's own internal discovery. They investigated a surprisingly low score of 5.6 for one of their models on the ARC AGI V3 benchmark. The problem wasn't the model's core intelligence. “Apparently OpenAI was able to investigate uh the low score of 5.6 Soul on ArcGIV3,” Coogan explained, “And the harness was not letting it remember what it had learned.” The 'harness'—the API settings governing how the model interacts with its environment—was throttling its ability to perform.

Then came the revelation: “We found that enabling two API settings tripled our scores with 6x fewer output tokens.” Let that sink in. Not a new model, not a massive training run, but two specific API settings made a threefold difference in benchmark scores, while also drastically cutting token usage. This points to a reality that ambitious founders must internalize: your integration layer is a massive, untapped performance lever.

Where This Breaks Down

While the 'harness' insight offers huge potential, it's not a universal panacea. This approach works best when your AI provider exposes granular API settings and when your use case demands sophisticated contextual memory or multi-turn interactions. If your application relies on extremely simple, single-prompt queries, the impact of such fine-tuning might be minimal, or the overhead of managing complex integrations could outweigh the gains. It also breaks down if you fall into the trap of over-optimizing for synthetic benchmarks without verifying the real-world value to your users. A higher score on ARC AGI V3 means nothing if it doesn't translate to a better, more efficient user experience in your product.

What to Do With This

Stop taking published benchmark scores or simple cost-per-token comparisons at face value for your specific use case. This week, pick one key AI-powered workflow within your product or internal operations. Identify the core model you're using. Then, go deep into its API documentation. Look for parameters beyond the obvious—things that control memory, context window management, or specific interaction protocols. Experiment with enabling or adjusting two or three such settings. Don't just measure token count; measure the actual cost and quality for completing a specific task. You might just uncover a hidden performance boost that feels like unlocking a brand new, vastly superior model, without spending a dime on upgrades.