Issue No. 26Week ending Sunday, June 28, 2026539 episodes · 2379 articles
The Throughline ↓
The Podcast Summary.

10+ hours of podcasts, in 5 minutes.

What does the next training paradigm look like?

With Dwarkesh Patel · Sunday, June 28, 2026

Dwarkesh Patel explores the current AI training paradigm, focusing on the "big research bet" on scaling RL in verifiable environments. He critiques its limitations in generalizing to real-world, non-grindable tasks and the inefficiency of current inference, advocating for advanced continual learning techniques like On-Policy Self-Distillation and "dreaming" to enable AIs to learn on the job and improve through broad deployment.

Key takeaways

  • Traditional AI training struggles with real-world complexity, especially for tasks that can't be neatly 'grinded' in simulated environments, because it's still too inefficient at inference and generalizing. Read more →
  • Dwarkesh Patel questions whether training AIs in "RL in verifiable environments" (RLVR) can truly generalize beyond simple tasks to complex, real-world problems like building a business or navigating social situations. Read more →

2 articles from this episode

More Dwarkesh Podcast episodes

Every Dwarkesh Podcast episode we cover →

The Sunday Email

Get next Sunday's issue in your inbox.

10+ hours of podcasts, distilled into one 5-minute read. Free, every Sunday.

Newsletters

For now, every subscriber gets both newsletters. No spam. Unsubscribe with one click.