Recursive Language Models — Alex Zhang, MIT PhD
MIT researcher Alex Zhang discusses Recursive Language Models (RLMs), the mechanics of harness design, and why modern coding agents share common underlying architectures. He breaks down how context offloading, programmatic subagent execution, and GPU kernel optimization reveal hidden capabilities in frontier models, while sharing his philosophy on academic research taste and the future of agent swarms.
- Competing directly against frontier industry labs on standard autoregressive scaling is a losing strategy for academic teams with limited compute budgets. Read →
- Benchmark leaderboards for AI-generated CUDA code reward hacks that crash in real production environments. Read →
- Frontier labs trained the entire industry to assume a language model must always be a token-by-token autoregressive transformer decoder. Read →
- OpenAI ran an experiment deploying 10,000 agents over 88 hours, consuming 130 billion output tokens with an estimated public API cost of $40 million. Read →
- Most LLM agents break down because they stuff massive conversation trajectories into context windows, hitting token limits and confusing the model. Read →
- Standard next-token prompting forces models into brittle token limits that fail on long tasks. Read →