Key Takeaways

  • Grok 4.7 launched with benchmark results targeted specifically at legal tasks evaluated by cost efficiency.
  • Anthropic released Opus 5.5, highlighted by vision-to-simulation demos that turned a napkin sketch of a trebuchet into an active virtual simulation.
  • OpenAI started rolling out GPT-6 Soul and Luna as foundation providers switch focus to pacing the frontier with cheaper inferences.
  • Foundation model price cuts provide immediate relief to vertical AI companies like Harvey, whose agentic architectures struggle with negative gross margins.

The Margin Rescue for Vertical Agents

For vertical AI startups, the biggest threat to survival has not been accuracy. It has been token consumption. Companies like Harvey face negative gross margins because chaining autonomous agents eats through massive compute budgets on every legal query.

John Coogan points out that the newest frontier wave attacks this exact problem. “Tokens are getting cheaper by the day,” Coogan said. “There's a bunch of model launches. Opus 5.5 launched today. Grok 4.7 was yesterday.”

The real battle is no longer who can spend the most capital on raw compute. It is who can drive task-specific inference costs into the floor. Grok 4.7 made this explicit by anchoring its release around legal workload efficiency. “The interesting thing about that particular benchmark was that it was a legal task benchmark based on cost,” Coogan noted. “And so Grok 4.7 is particularly cheap for the type of legal work that at the very least Harvey wants to do.”

When foundation providers compete on cost per vertical task, application layer margins flip from red to black without founders rewriting their underlying agent logic.

Pacing the Frontier Through Efficiency and Multimodal Simulation

The second shift across Opus 5.5 and the early rollouts of GPT-6 Soul and Luna is functional utility over raw parameter counts. Providers are shifting how they define frontier progress.

As Coogan explained regarding recent releases, “This is the first model that we've released since we said we were pacing the frontier. This is an expression of the pacing... this is, you know, Fable 5.1 class model, but it's much cheaper.”

Alongside cost reduction, capabilities are expanding into spatial simulation. Coogan highlighted an Anthropic demo where Opus 5.5 bridged the gap between static image input and live physics: “I saw a cool demo where someone sketched out on a piece of paper sort of a trebuchet and Opus 5.5 took that image into simulation and created like a virtual version of it.”

Instead of waiting months for massive parameter jumps, builders are getting specialized efficiency and direct physics simulation at a fraction of last year's pricing. With GPT-6 Soul and Luna entering rollout, the pressure on per-token pricing will only accelerate.

What to Do With This

Audit your agentic workflows this week to separate high-reasoning steps from repetitive domain tasks. Route high-volume legal, parsing, and extraction calls away from legacy frontier endpoints to Grok 4.7 or Opus 5.5 to immediately expand your gross margins.