6 quotes from 1 episode on No Priors, each with a timestamped link to the source.
6 quotes1 episode
The short version
Noam Brown argues that evaluating AI models requires measuring the test-time compute budget spent on a task. Model capabilities are fundamentally bottlenecked by the weeks or months of processing time required to complete complex work.
Most interesting insights
Current models act as complements to humans because the systems lack strategic judgment and research taste.
“…they don't have very good research taste right now.”
Plotting model performance against cost, tokens, or time reveals an accurate picture of capabilities. Current scaling frameworks fail to measure this extended processing factor.
“My claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of test-time compute that's going into the model…”
Achieving powerful results forces models to run for extended periods. This heavy reliance on test-time compute makes system evolution a drawn-out process.
“If it requires so much test time on compute to unlock the full capabilities of the model…”
“…then that means you're bottlenecked by time; things can only go so fast because the models need to run for long enough to actually do something really, really powerful.”
Current AI model benchmarks, often presented as a single-point "grid," fail to account for the amount of compute spent during evaluation, masking true model capabilities.
Modern models like OpenAI's GPT-5.5 can achieve significantly higher performance by "thinking" for extended periods—weeks or even months—a crucial factor ignored by standard evaluations.
Forget the science fiction fantasy: Brown says we're not heading for an "overnight intelligence explosion" where AI instantly becomes superhuman across the board.
Current AI models are fundamentally bottlenecked by large-scale "test-time compute," meaning they need weeks or months of processing to unlock full capabilities.
How we attribute quotes. Every quote was matched against the episode transcript, so the words and the timestamp are real (we trim filler words like "um", nothing else). The name comes from our written summary of the episode. YouTube gives us no voice-by-voice transcript, so open the timestamp to hear who is talking. See a wrong name? Tell us and we fix or remove it.