Key Takeaways

  • An Anthropic pre-training researcher with three years across OpenAI and Anthropic resigned publicly, claiming labs are racing toward self-improving superintelligence.
  • Anthropic alignment scientist Evan Hubinger backed the resignation, stating he believes the risk of human extinction from AI exceeds 10% within the next ten years.
  • Jordi Hays argued that assigning a 10% extinction probability requires sharing a testable methodology and explaining what evidence would change that belief.
  • The debate reveals a sharp split between internal safety researchers who choose public resignations and builders who view walking away as ineffective theater.

The Disagreement

When an engineer leaves frontier model development, the industry usually looks at vesting schedules or compensation packages. This departure was different. A pre-training researcher walked away from Anthropic and the entire field, posting: “I spent the last three years doing pre-training research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”

Shortly after, Anthropic alignment scientist Evan Hubinger validated the stance. “Jacob is correct here,” Hubinger said. “We really do earnestly believe AI could kill all humans. I personally think it's above 10% within the next decade.”

That 10% figure sparked immediate pushback from investors and operators. Hays targeted both the lack of mathematical rigor behind the probability and the decision to quit rather than stay.

“If you publicly assign a greater than 10% probability to human extinction within a decade, you owe people a clear explanation of how you reach that number plus what would change your mind,” Hays noted. He questioned the strategic value of walking away from the lab entirely: “Believing in your whole heart that things are that serious and then just rage quitting and not leaning in and trying, do you really think the rage quit method is going to work?”

On one side stand researchers who treat runaway recursive self-improvement as an imminent engineering reality. On the other stand founders and investors who view loose probability estimates as emotional rhetoric that distracts from technical progress.

Who's Right (and When They're Wrong)

Hays is right on epistemology: probability numbers without explicit models are assertions, not arguments. In engineering, assigning a 10% failure rate to a bridge requires stress tests, material science data, and failure mode simulations. In frontier AI safety, numbers like 10% P(doom) often operate as social signals. They convey deep concern rather than quantifiable statistical risk. When researchers throw out double-digit existential odds without showing their work or defining falsifiable criteria, they invite valid skepticism from technical peers.

Hays is also right about operational impact. Leaving the frontier lab removes your ability to touch the weights, audit training runs, or install governance guardrails. If an airplane engine has a flaw, the propulsion engineer does not fix it by stepping out of the hangar.

Yet the resigning researchers have a point that critics dismiss too quickly. Inside frontier pre-training, researchers witness the velocity of capability jumps firsthand. When competitive pressure between Anthropic, OpenAI, and Meta forces faster release cycles, internal dissent often gets sidelined. If an engineer concludes that commercial pressure has completely overridden safety commitments, staying inside the building can turn into quiet complicity. Quitting publicly is their final attempt to create an external speed bump.

The breakdown happens when resignations lack concrete technical documentation. A public post that screams fire without publishing the thermal data changes zero executive decisions. It only creates a temporary news cycle.

What to Do With This

Audit how your technical team communicates critical system risks. If an engineer flags a high-severity bug or data exposure with vague doom language, mandate a written post-mortem format: state the exact failure mechanism, calculate empirical odds from historical logs, and list three falsifiable tests that would prove the risk is resolved.