Key Takeaways

  • Frontier labs trained the entire industry to assume a language model must always be a token-by-token autoregressive transformer decoder.
  • Systems like GEV prove that modifying the model output space produces much faster and cheaper inference for targeted tasks like verification, gaming, and classification.
  • Neo-labs and startups waste capital when they try to clone OpenAI or Anthropic on standard pretraining compute instead of finding architectural edges.
  • In Recursive Language Models (RLMs) and agent swarms, the primary bottleneck is latency caused by chaining standard text-to-text inference calls.

The Trap of Text-to-Text Dogma

Almost every developer builds software under the assumption that a language model is a text box that takes tokens in and emits tokens out. Frontier labs spent billions optimizing this exact setup. As a result, the market treats the autoregressive decoder as the only way to build intelligent software.

Alex Zhang from MIT argues that this assumption limits what builders can create. “For the longest time because the labs are the only places that control, you are never going to use something other than GPT or comparable models because they are the best models,” Zhang says. “Because of that, people have gotten accustomed to this idea that a language model is just an autoregressive decoder.”

A language model is simply a statistical model of language. It does not have to be a transformer decoder that emits one word at a time. When builders limit themselves to text-to-text APIs, they inherit every latency penalty and compute inefficiency that comes with generating full string sequences.

Rethinking Output Spaces to Fix the Latency Bottleneck

When builders run complex multi-step systems, the standard autoregressive approach grinds to a halt. Chaining sequential LLM calls multiplies latency until the system becomes unusable in production.

“The biggest bottleneck in RLMs or swarms or systems like these is they are slow,” Zhang points out. “When you do multiple language model calls all the time, you are not distributing your compute correctly.”

This is where systems like GEV enter the equation. Instead of forcing a model to generate verbose text strings for verification or discrete actions, researchers can tune the output space directly. Mapping the backbone representation directly to task-specific outputs removes the need to generate dozens of intermediate tokens.

Zhang notes: “What is really interesting about GEV is that it opens up this question: are language models correct in the form that they are in? Can we consider a different design space other than text to text? We now have a new thing to tune, which is what is the output space and how does this affect inference latency.”

The Compute Reality for Neo-Labs

For smaller research teams and startups, competing on raw pretraining scale against frontier labs is financial suicide. “If their strategy is just to replicate OpenAI or Anthropic, that is a horrible strategy,” Zhang argues. “You kind of just have to think of it in terms of what advantage you have.”

The structural advantage for smaller teams lies in non-autoregressive output designs, specialized execution setups, and custom verification layers. If your product requires thousands of verification or routing calls per minute, ditching standard token generation will give you an order-of-magnitude speed advantage over competitors who simply query frontier model APIs.

What to Do With This

Audit your multi-step agent workflow tomorrow morning. Identify every step where an LLM call generates a text response merely to make a binary decision, classify an intent, or output a structured score. Replace those full text generation calls with a non-autoregressive classifier or a small, task-specific model head to cut your system latency in half.