Key Takeaways

  • Asking large generative models to predict complex multi-agent execution plans up front creates fragile pipelines that break when agents touch shared resources.
  • John Lindquist uses Typesafe AI's Jev model to run micro decision checks at every operational step, preventing agents from colliding on files or spatial coordinates.
  • Claire Vo terms this approach "efficient inefficiency": brute-force evaluating and ranking every possible immediate move at runtime instead of relying on slow, monolithic planning.
  • Running ultra-fast classification checks at each tick eliminates the high latency and token costs of querying frontier reasoning models for routing decisions.

Stop Asking Frontier LLMs to Plan Multi-Agent Paths

When developers build multi-agent setups, the default impulse is to feed the full problem into a large frontier model. You ask it to map out an entire sequence of actions, anticipate dependencies, and assign subtasks to individual workers.

That strategy falls apart the moment two agents need the same file or work on dependent branches. As Lindquist points out, parallel execution requires constant collision avoidance so agents “never end up on the same space” or disrupt shared state.

Frontier reasoning models are too slow and expensive to query on every micro-step. If an agent must wait two seconds for a large model to evaluate whether a file lock is clear, the whole parallel system grinds to a halt. Trying to solve this by having the large model forecast every step in advance fails because runtime environments are dynamic. The plan goes stale after the first step completes.

Vo highlights the shift in approach: “In the past, last week we were tossing these problems to these like big brainy LLMs and being like, 'Think really hard. Just think really hard and come back with a plan.' And instead here, you can really very quickly evaluate all those options, stack rank them, and go through the sort of like inefficient universe of options and come to the right kind of conclusions.”

The Mechanics of Efficient Inefficiency

Instead of a master plan generated upfront, Lindquist structures parallel workflows around high-frequency state evaluation. At every operational tick, a specialized classification model like Jev ingests the state of all active agents and tests immediate candidate actions.

Lindquist explains the setup: “If you think of this in agentic programming or any sort of parallel work you'd be doing to ensure that you have things running and you want to analyze to help steer them in different directions so they never collide or disrupt each other or touch the same files, you can ask Jev to take in every single step and check.”

Vo calls this "efficient inefficiency." In classical computing, generating and ranking every possible permutation of an immediate move seems wasteful. But when done with an ultra-light, low-cost decision model, brute-force ranking of immediate actions beats high-level generative reasoning. You do not care if evaluating ten branch options takes twenty micro-classifications, provided those classifications return in milliseconds for pennies.

This pattern cleanly separates generation from coordination. Let your heavy models write the actual code or synthesize the data. Let lightweight, deterministic decision models handle traffic control, file locking, and routing.

What to Do With This

Audit your multi-agent architecture this week and strip out any prompt that asks an LLM to "plan the full sequence of tasks" before execution. Replace that planning block with a step-level loop that passes candidate actions into a lightweight classifier to check for state conflicts before firing each worker.