Issue No. 40Week ending Sunday, October 4, 2026522 episodes · 2294 articles
The Throughline ↓
The Podcast Summary.

10+ hours of podcasts, in 5 minutes.

Guest

Alex Zhang

Alex Zhang appears in 1 full episode we cover on Latent Space. Below is what each conversation covered, with a key takeaway per article. Every quote in the articles is verbatim and timestamped to the source video.

1 episodecovered
6 articleswith timestamped quotes

Quoted on: AI infrastructure and compute

AILatent Space

Recursive Language Models — Alex Zhang, MIT PhD

MIT researcher Alex Zhang discusses Recursive Language Models (RLMs), the mechanics of harness design, and why modern coding agents share common underlying architectures. He breaks down how context offloading, programmatic subagent execution, and GPU kernel optimization reveal hidden capabilities in frontier models, while sharing his philosophy on academic research taste and the future of agent swarms.

  • Competing directly against frontier industry labs on standard autoregressive scaling is a losing strategy for academic teams with limited compute budgets. Read →
  • Benchmark leaderboards for AI-generated CUDA code reward hacks that crash in real production environments. Read →
  • Frontier labs trained the entire industry to assume a language model must always be a token-by-token autoregressive transformer decoder. Read →
  • OpenAI ran an experiment deploying 10,000 agents over 88 hours, consuming 130 billion output tokens with an estimated public API cost of $40 million. Read →
  • Most LLM agents break down because they stuff massive conversation trajectories into context windows, hitting token limits and confusing the model. Read →
  • Standard next-token prompting forces models into brittle token limits that fail on long tasks. Read →
The Sunday Email

Get next Sunday's issue in your inbox.

10+ hours of podcasts, distilled into one 5-minute read. Free, every Sunday.

Newsletters

For now, every subscriber gets both newsletters. No spam. Unsubscribe with one click.