Issue No. 26Week ending Sunday, June 28, 2026539 episodes · 2375 articles
The Throughline ↓
The Podcast Summary.

10+ hours of podcasts, in 5 minutes.

Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown

With Sarah Guo, Noam Brown · Sunday, June 28, 2026

This episode features OpenAI research scientist Noam Brown discussing the shortcomings of current AI model evaluation benchmarks, particularly their failure to account for large-scale test-time compute. He explains how this oversight impacts the assessment of model capabilities and has significant implications for AI safety and responsible scaling policies. Brown also shares insights into the true nature of recursive self-improvement and the potential of latent capabilities in current models.

Key takeaways

  • Current AI model benchmarks, often presented as a single-point "grid," fail to account for the amount of compute spent during evaluation, masking true model capabilities. Read more →
  • Early LLMs (pre-GPT-5.2) were useless for complex reasoning tasks like building a poker bot; they simply couldn't get started. GPT-5.2 could help, but required constant oversight. Read more →
  • Your current AI models, even public ones like GPT-5.5, hold significant untapped "latent capabilities." OpenAI recently disproved the complex Erdos unit distance conjecture internally, a feat Noam Brown believes GPT-5.5 could mimic with roughly $100,000 in dedicated compute. Read more →
  • Forget the science fiction fantasy: Brown says we're not heading for an "overnight intelligence explosion" where AI instantly becomes superhuman across the board. Read more →

4 articles from this episode

More No Priors episodes

Every No Priors episode we cover →

The Sunday Email

Get next Sunday's issue in your inbox.

10+ hours of podcasts, distilled into one 5-minute read. Free, every Sunday.

Newsletters

For now, every subscriber gets both newsletters. No spam. Unsubscribe with one click.