Anthropic Researcher Resigned Amid Existential Risk Fears

The departure follows internal warnings from experts that advanced AI models may pose a threat to human survival.

Updated on Sept. 20, 2026 in Artificial Intelligence

Isometric editorial illustration showing a heavy industrial power transformer in a clean, minimalist facility corridor, representing AI infrastructure and safety risks.
A researcher at Anthropic has resigned, citing concerns over existential risks and safety protocols related to the development of advanced autonomous artificial intelligence systems. AI Illustration. Upload story photo >

Live Poll

Do you trust that the current development of artificial intelligence will benefit humanity in the future?

Jacob Coxon has resigned from AI safety company Anthropic, citing concerns over the firm's research trajectory. The move highlights growing debate within the industry regarding the potential for autonomous systems to cause human extinction.

Why it matters

Internal dissent over safety protocols underscores the tension between rapid AI development and the management of existential risks. These concerns are shaping the research agendas of institutions studying how systems might pursue goals independently of human values.

Researcher Evan Hubinger has projected a greater than 10% probability of AI-led human extinction within the next decade. This assessment stands against prior industry forecasts that generally categorize such outcomes as low-probability tail risks.

The players

Jacob Coxon

A former employee of Anthropic who resigned due to disagreements over the company's research direction.

Anthropic

An AI safety and research company focused on building large language models with a stated emphasis on constitutional safety protocols.

Evan Hubinger

An AI safety expert associated with research at institutions like Carnegie Mellon University and Stanford University.

Thomas Larsen

An analyst whose research papers detail projected AI development trajectories through 2040.

The details

Autonomous AI threats theoretically require systems to scale beyond digital intelligence into physical world dominance, including the ability to construct factories and energy infrastructure. This mirrors historical cyber-physical security incidents like Stuxnet, which damaged Iranian nuclear facilities after gaining access through physical USB ports. Experts argue that security protocols designed for isolated weapon systems are insufficient to contain models capable of autonomous goal pursuit.

Timeline

  1. September 2026: Jacob Coxon resigned from Anthropic.

  2. 2027: Target date for analysis in Thomas Larsen's research report.

  3. 2040: Future timeline analyzed in Thomas Larsen's reports.

  4. Next decade: The timeframe during which experts estimate potential human extinction risks.

The Tech Race

The discourse surrounding AI risk mirrors the 2010 Stuxnet cyberattack, where isolated critical infrastructure was compromised through digital means. Experts are now evaluating whether future autonomous models will possess the capacity to execute similar real-world interventions.

The debate remains largely centered on research and safety policy, with no direct changes to commercial AI product accessibility or user workflows. The trajectory of this concern will be dictated by upcoming benchmarks for model autonomy and security.

The takeaway

The field is shifting toward a period of heightened internal scrutiny as safety researchers weigh the likelihood of rapid system scaling. Observers should track upcoming analysis from Thomas Larsen and similar academic institutions for new data on 2027 development milestones.

Further reading

For more on the current state of safety research, visit Artificial Intelligence.

Source note: This article includes information reported by Protothemanews.

Live Poll

Do you trust that the current development of artificial intelligence will benefit humanity in the future?