Key Takeaways
- OpenAI demonstrated automated problem-solving across 722 complex mathematical manuscripts, tackling problems at a speed no human mathematician has matched.
- Reinforcement learning with verifiable rewards (RLVR) lets models push math and code arbitrarily far without waiting on new human data.
- Domains without objective, verifiable ground truth risk hitting a plateau because they remain bottlenecked by human-generated training datasets.
- Coogan argues the industry should retire the search for human-shaped AGI and recognize machine intelligence as an uneven, jagged spike.
The Jagged Frontier of Machine Intelligence
Most people look at artificial general intelligence through a human lens. They expect a system that writes essays at a college level, reasons about philosophy, and solves calculus, moving upward smoothly across all skills at the same time.
Recent results from OpenAI challenge that mental model entirely. On TBPN, John Coogan and Jordi Hays examined OpenAI's release showing AI solving over 700 complex mathematical manuscripts. Hays pointed out the sheer scale: “I think the full number is 722 like manuscripts and then so it's kind of hard to like how do you delineate like an actual problem.”
To Coogan, this output looks completely unlike human cognition. No person in history has ever analyzed and resolved hundreds of high-level mathematical manuscripts in a single burst. “This is a less of a field the AGI moment and more of a field the super intelligence moment for me personally,” Coogan explained, “because no human intelligence has ever been able to solve this many math problems this quickly at this level of mathematical ability.”
The capability does not represent a smooth upward lift across every intellectual task. It represents a sharp, narrow spike.
Why RLVR Changes the Rules
Citing analysis by researcher François Chollet, Coogan explained why math and software development are breaking away from everything else. The secret lies in reinforcement learning with verifiable rewards (RLVR).
In mathematics or programming, the computer knows instantly if an answer is correct. A compiler rejects broken code. A mathematical proof either verifies logically or it fails. Because the reward function is objective, a model can play millions of games against itself, generating synthetic trials and learning from errors without needing a person to grade its homework.
Compare that to writing marketing copy, negotiating a commercial contract, or diagnosing subjective symptoms. In those areas, feedback requires messy human intervention. If you want the model to write better novels, you need humans to read them and label what feels compelling.
Coogan framed the split plainly: “What if the jagged frontier is mainly math plus code, which you can push arbitrarily far with reinforcement learning, learning with verifiable rewards, RLVR, and everything else starts to plateau because it's still bottlenecked by human generated data.”
If that hypothesis holds, our language for this technology is backwards. “I still like the machine intelligence name because it has a different shape,” Coogan noted. “We're trying to mold the machine intelligence into something that slots in neatly into like a human a general framework, but it keeps breaking out and spiking in these odd areas because of the ability to train in these environments.”
What to Do With This
Audit your product roadmap this week. If you are building features around open-ended human judgment, customer taste, or subjective text, assume foundation model progress will move at an incremental pace tied to human data collection. If you are building workflows that can be mechanically checked (compilers, formal proofs, unit tests, structural validation, schema enforcement), design your software around the assumption that machine reasoning will soon be orders of magnitude faster and cheaper than any human team.