Key Takeaways

  • OpenAI ran an experiment deploying 10,000 agents over 88 hours, consuming 130 billion output tokens with an estimated public API cost of $40 million.
  • MIT researcher Alex Zhang estimates that 95% of compute in unguided agent swarms is pure waste, exploring dead ends that never contribute to the final answer.
  • Swarm coordination is not an emergent trait of smarter base models; it requires specialized post-training and scaffolding designed specifically for convergence.
  • Pointing massive compute at hard problems only works when you build strict verification loops that prune useless exploratory branches before token costs explode.

Burning $40 Million to Solve One Proof

When researchers talk about scaling test-time compute, they usually mean asking a model to think for five minutes. OpenAI took that concept to an extreme. During a recent compute run, they orchestrated 10,000 parallel agents across 88 hours. The system churned through 130 billion tokens, equivalent to roughly $40 million at public API rates, to tackle complex proofs.

As Zhang noted, the sheer ability to throw that level of compute at a task marks a shift: “It is very exciting that we even have the option to point $40 million at a problem and solve it.”

Yet throwing raw compute at a problem creates an immediate efficiency trap. When you spin up thousands of subagents without tight constraints, they do not automatically coordinate like a disciplined research team. Instead, they produce vast amounts of digital exhaust.

Why Smarter Models Do Not Solve Coordination

The popular assumption among founders is that raw model capability solves agent orchestration. If GPT-5 or a future model is smart enough, the logic goes, throwing a swarm of them at a codebase will just work.

Zhang rejects this view. “The agent swarm design is not something you can just take for granted. Like it's not like GPT6 Astra is just super smart and then it just got agent swarms running well.” He pointed to historical training runs where OpenAI had to deliberately post-train models to act like a coordinated swarm rather than isolated instances.

Without explicit training for collective problem-solving, swarms degenerate into token furnaces. Zhang estimates that “95% of the swarm is entirely useless or like what it's exploring is entirely you're just burning tokens.” When subagents lack reliable ways to evaluate peer output, they branch into irrelevant tangents. They hallucinate plausible-looking paths, duplicate work, and drown the correct signal in gigabytes of generated slop.

Engineering Swarms That Actually Converge

The primary challenge in multi-agent systems centers on forcing convergence rather than adding more workers.

“I think we take for granted what it means for a swarm to converge to an answer,” Zhang explained. In a formal proof or a code repository, broad exploration is useless if the system cannot systematically discard incorrect trajectories and synthesize the remaining truths.

For founders building agent systems today, the takeaway is clear. If you build multi-agent architectures that rely purely on prompting off-the-shelf models to collaborate, you are paying for 95% waste. Making swarms economically viable requires programmatic scaffolds, deterministic verifiers, and task-specific post-training that cut dead branches immediately.

What to Do With This

Audit your agent architecture this week. Look at the total token consumption across multi-step or multi-agent workflows in your staging logs. Calculate the percentage of subagent outputs that actually make it into the final user-facing response. If that number is under 20%, stop adding subagents and build a deterministic validator to kill failing paths at step two instead of step ten.