Key Takeaways
- Standard next-token prompting forces models into brittle token limits that fail on long tasks.
- Recursive Language Models (RLMs) treat code execution as their single tool, offloading raw text to an external environment.
- Models trained in this setup generalize to tasks 8 to 30 times longer than their training samples without extra fine-tuning.
- Teaching an agent to iterate over data programmatically allows one problem-solving strategy to transfer across unrelated domains.
- Future frontier model providers will likely package multi-agent swarms directly inside standard API endpoints.
Stop Feeding Raw Context to Next-Token Prompts
Most developers building AI agents try to shove full documents and conversational history directly into the context window. They hope attention mechanisms will sort out the details. When the document grows from 2,000 words to 50,000 words, accuracy falls off a cliff. The model hallucinates, loses instructions, and burns tokens.
Alex Zhang at MIT points out that this failure comes from forcing models to operate purely through autoregressive next-token prediction. Instead of treating text generation as the only tool, an RLM strips the model's direct view of long context and places the data into an external code environment.
As Zhang puts it: “An RLM is basically just a harness design where the only tool in the harness is code.”
The model never sees the entire corpus at once. Instead, it inspects variables, runs sub-queries, and processes data programmatically. By turning massive inputs into small, isolated chunks, the model solves each micro-step inside its normal comfort zone.
Generalizing Across Tasks by Changing a Variable
When models process data through executable scripts, their problem-solving logic stabilizes. A prompt that reads a single paragraph looks totally different to a standard model than a prompt reading a hundred pages. But in Python, reading one item or a thousand items uses the exact same loop.
Zhang discovered that this programmatic interface gives models an unexpected superpower: inductive generalization. As Zhang explained, “if you sufficiently offload context and write ask the model to write code over that context, you get this really really weird but useful property which is that when the model recognizes during training how to solve a task, it turns out that the solution to many tasks is very similar.”
Because the programmatic structure stays identical, “when you train the RLM on one of them, it generalizes this behavior to the second one.”
Even more striking is length scaling. Models trained on small inputs scale naturally to inputs 8 to 30 times longer. Zhang noted that “they naturally learn to generalize, for example, to longer tasks because the strategy is basically the same. You're just modifying like a length variable.”
Swarms Behind the API
This dynamic points to a major shift in how developers will interact with frontier models. Right now, founders build complex orchestration setups on top of raw completion endpoints. They handle retries, prompt chaining, and memory buffers manually.
That wrapper layer will soon move directly inside the model API. Zhang predicts that “what we think of as a language model like the thing that we query might actually be like a swarm or like a scaffold or like some weird harness design that scales very well, but the user just doesn't see it.”
If the model provider executes subagents and code loops under the hood, proprietary wrappers built only on token orchestration will lose their value. The defensible layer is not how you string prompts together; it is the external environment and proprietary execution tools you give the model.
What to Do With This
Audit your agent pipeline this week and identify where you pass raw text directly into prompt templates. Replace those text blobs with a Python execution environment where the model inspects data through variable queries and sub-functions. Measure task accuracy when inputs scale 10x past your current prompt limits.