Key Takeaways
- Anthropic pre-training researcher Jacob Coxon resigned from the artificial intelligence industry after three years across OpenAI and Anthropic, warning both labs are racing toward self-improving superintelligence.
- Anthropic alignment scientist Evan Hubinger puts the probability of AI-driven human extinction at greater than 10% within the next ten years.
- Silicon Valley investors like Andreessen Horowitz general partner Martin Casado push back, arguing that doom probabilities lack technical falsifiability and create strange cognitive dissonance compared to real defense projects.
- AI labs approaching public offerings face a collision between existential doom rhetoric and mandatory SEC S-1 risk disclosures, where vague extinction claims invite heavy legal liability.
The Disagreement
When an engineer builds pre-training pipelines for three years at both OpenAI and Anthropic, people listen when he walks away. Jacob Coxon left his post at Anthropic and abandoned the artificial intelligence industry entirely. His statement pulled no punches: “I spent the last three years doing pre-training research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving super intelligence and gambling with our lives.”
Shortly after, Anthropic alignment scientist Evan Hubinger defended the broader concern, saying, “We really do earnestly believe AI could kill all humans. I personally think it's above 10% within the next decade.”
The pushback from operators and investors was immediate. If you assign a specific statistical probability to human extinction within ten years, you must present a testable, falsifiable mechanism. Otherwise, it functions as speculative theater. Martin Casado, general partner at Andreessen Horowitz, pointed out the strange culture surrounding these claims: “I worked on an actual thermonuclear weapons project, Martin Casado from A16Z says, and the dissonance was less bizarre.”
Casado's critique highlights the divide. On one side sit safety researchers who treat recursive self-improvement as an urgent hazard. On the other side sit builders and engineers who view ungrounded probability estimates as unscientific claims that justify corporate regulatory moats.
Who's Right (and When They're Wrong)
The safety researchers are right about one core mechanical risk: frontier labs face intense competitive pressure to shorten safety evaluations and deploy models before they are understood. When capital spending runs in the tens of billions, commercial momentum sidelines internal caution.
Where the doomer position collapses is its refusal to provide falsifiable milestones. Assigning a 10% extinction probability without defining the exact steps, empirical thresholds, or observable failure modes turns engineering safety into unfalsifiable belief. If a researcher cannot state what experimental result would lower their probability from 10% to 1%, the number is a personal impression rather than an objective measurement.
This debate will soon face a harsh legal reality in the public markets. When AI labs file an S-1 registration statement with the SEC for an initial public offering, corporate lawyers must write concrete risk factors. Stating in an SEC filing that your flagship product carries a 10% chance of ending human civilization within ten years invites direct regulatory scrutiny and shareholder lawsuits. Frontier labs will either have to put precise mathematical models behind their safety claims or drop the doom rhetoric to avoid severe legal exposure.
What to Do With This
Audit your product roadmap and risk statements for unfalsifiable claims. Strip out vague worst-case hand-waving and replace it with three concrete metrics you can track every quarter, such as automated test failure rates, prompt injection breach counts, or latency spikes during failure states.