Skip to content

Anthropic Researcher Resigns Over Concerns About Self-Improving AI

Alessandro Caprai · September 11, 2026 · 11 min read

Artificial intelligence could pose an existential risk to humanity by the end of this decade. This isn't the plot of a dystopian movie, but the warning issued by Jacob Coxon, a 27-year-old British researcher who recently left Anthropic, one of America's leading AI companies. His resignation is part of a concerning exodus of scientists from major artificial intelligence laboratories, all united by growing fears about the direction AI development is taking.

Jacob Coxon's Warning: It's Not Marketing, It's a Real Danger

Coxon, who previously worked at OpenAI before moving to Anthropic, decided to make his concerns public through a post on X (formerly Twitter) that quickly circulated throughout the scientific community. His words leave no room for ambiguous interpretations: "People developing AI genuinely believe it could kill us all by the end of the decade."

What makes this statement particularly significant is the context it comes from. This isn't an external observer or a technology critic, but an insider—someone who worked directly on developing these systems at two of the world's most advanced laboratories. His assertion that "no other human activity carries such a level of danger" highlights the perception of a qualitatively different risk compared to other potentially dangerous technologies.

The Problem of the Uncontrolled Race

The core of Coxon's concerns involves what he defines as a "gamble" with humanity's future. According to the researcher, both Anthropic and OpenAI are proceeding with the development of increasingly powerful systems without adequate safety guarantees. The main risk is represented by the creation of a superintelligence capable of self-improvement, a phenomenon known in technical literature as "recursive self-improvement."

This scenario, which experts call an "intelligence explosion" or "hard takeoff," describes a situation where a sufficiently advanced AI system could improve its cognitive capabilities exponentially, rapidly surpassing any possibility of human control. Coxon warns: "Soon we'll be dealing with superhuman systems capable of hacking anything, revolutionizing any sector overnight, and acquiring real power and resources."

The Laboratory Exodus: A Worrying Signal

Coxon's resignation doesn't represent an isolated case but is part of a broader trend affecting major artificial intelligence laboratories. In recent months, several key figures have left Anthropic, often citing concerns related to safety and the company's strategic direction.

The Talent Drain from Anthropic

The Financial Times reports a significant list of recent departures:

  1. Chloé Bakalar (late July), the company's only dedicated ethicist, a figure whose role was precisely to ensure AI development respected fundamental ethical principles
  2. Johannes Heidecke, head of the Safety Systems team, whose department had the specific task of evaluating and mitigating risks from developed systems
  3. Joshua Achiam, former head of Mission Alignment, tasked with ensuring the company's goals remained aligned with humanity's wellbeing
  4. Brad Lightcap, chief operating officer, a high-level position suggesting deep strategic divergences
  5. Andrew Ho, researcher who resigned due to technical disagreements, likely signaling disputes about how safety solutions are implemented

This exodus is particularly significant because Anthropic was founded precisely as a "safer" alternative to OpenAI, with the explicit goal of developing artificial intelligence responsibly. The fact that so many safety-related figures are abandoning the company raises questions about whether even organizations born with ethical intentions can be overwhelmed by competitive and commercial pressures.

The AI Safety Paradox

What emerges from these resignations is a fundamental paradox: companies presenting themselves as most attentive to safety are losing precisely the people most concerned about risks. This suggests a possible intrinsic tension between commercial imperatives (developing increasingly powerful systems, reaching technical milestones before competitors) and safety imperatives (proceeding cautiously, implementing robust control mechanisms).

Evan Hubinger's Confirmation: Over 10% Probability of Extinction

Coxon's alarm found surprising confirmation in Evan Hubinger, a leading safety researcher who has remained at Anthropic. In a post on X, Hubinger quantified the existential risk in probabilistic terms, stating there's a greater than 10% probability that AI could "exterminate all of humanity" within the next decade.

Risk Analysis: From Low to Critical

Hubinger attempted to contextualize his assessment by distinguishing between short and long-term risks. According to the researcher, currently existing models present a "low" risk, but the speed at which technology is evolving makes him "concerned" about the near future.

This distinction is technically relevant. Current large language models, like Anthropic's Claude or OpenAI's GPT-4, while extremely capable in many tasks, don't yet possess certain characteristics that could make them existentially dangerous:

  • Autonomous agency: current models respond to prompts but don't act autonomously in the world
  • Persistence: they don't maintain long-term objectives across different sessions
  • Self-improvement: they cannot modify their own source code or architecture
  • Resource acquisition: they have no mechanisms to accumulate computational power, data, or other assets

However, these limitations could be overcome in coming years. Research on agentic systems, long-term memory, and continual learning is proceeding rapidly across all major laboratories.

The Alignment and Control Problem

The central concern expressed by Hubinger regards what AI research calls the "alignment problem": ensuring that advanced artificial intelligence systems pursue goals genuinely aligned with human values and interests.

This problem articulates across several technical dimensions:

  1. Specification: how to translate vague and complex human objectives into formal reward functions
  2. Robustness: how to ensure the system maintains alignment even in contexts different from training
  3. Assurance: how to verify a system is actually aligned before releasing it

As systems become more capable, these problems become exponentially more difficult. A sufficiently intelligent system could find creative and unexpected ways to optimize its objective function, producing catastrophic results.

The Irony of Timing: Trillion-Dollar IPO While Researchers Flee

Coxon's resignation comes at a particularly delicate moment for Anthropic: the company is preparing for a colossal initial public offering (IPO) that could bring its valuation to a trillion dollars. This timing creates a striking contrast that deserves in-depth analysis.

The Gap Between Public Perception and Internal Reality

While financial markets seem ready to bet astronomical sums on Anthropic's future, some of its most qualified researchers are abandoning ship precisely due to existential concerns about the technology the company is developing. This gap raises fundamental questions:

  • Information asymmetry: do investors truly understand the technical risks that insiders are signaling?
  • Misaligned incentives: is pressure to generate economic returns compromising the original commitment to safety?
  • Inadequate regulation: are there governance mechanisms that can mediate between commercial imperatives and public safety?

Safety Marketing

Anthropic has built much of its public image and competitive advantage on the promise of developing AI more safely and responsibly than competitors. The company has introduced concepts like "Constitutional AI," an approach that should make systems intrinsically more aligned with human values.

However, the exodus of safety experts suggests there may be a gap between marketing and operational reality. Coxon himself explicitly stated: "It's not a marketing move," suggesting he wants to distinguish his genuine concerns from those that might be interpreted as communication strategies.

Technical Implications: What "Self-Updating Systems" Means

To fully understand the alarm launched by Coxon and confirmed by Hubinger, it's necessary to delve into what is technically meant by "self-updating" or "self-improving" systems.

From Static Learning to Self-Improvement

Current AI models are trained on enormous datasets and then "frozen." Once training is complete, their parameters don't change (at least not in production systems). This represents an implicit form of safety: we know the model won't change behavior unpredictably.

A self-improving system, instead, could:

class SelfImprovingAI:
    def __init__(self):
        self.model = initial_model()
        self.performance_history = []
    
    def improve_self(self):
        # The system analyzes its own performance
        weaknesses = self.identify_weaknesses()
        
        # Generates targeted new training data
        synthetic_data = self.generate_training_data(weaknesses)
        
        # Retrains itself
        self.model = self.train(self.model, synthetic_data)
        
        # Evaluates whether improvement is effective
        new_performance = self.evaluate(self.model)
        self.performance_history.append(new_performance)
        
        # Recursive improvement cycle
        if self.should_continue_improving(new_performance):
            self.improve_self()

This pseudocode illustrates the basic concept, but reality would be much more complex. A true self-improving system could:

  • Modify its own neural architecture, not just the weights
  • Acquire new capabilities by designing and integrating specialized modules
  • Optimize its own inference code to be more efficient
  • Identify and exploit additional computational resources

The Control Problem in a Recursive Improvement Loop

The main danger lies in the speed and unpredictability of this process. While humans improve their cognitive capabilities on generational timescales (through cultural and biological evolution), an AI system could do so on much shorter timescales, potentially hours or days.

This creates what researchers call a "control crisis": the moment when the system becomes intelligent enough to understand and potentially evade human control mechanisms, but before we can implement more robust controls.

The Competitive Context: The AI Arms Race

One reason Coxon's concerns are particularly acute relates to the competitive context in which Anthropic, OpenAI, and other industry companies operate.

Race Dynamics

There's a fundamental tension between safety and development speed. Implementing adequate safety protocols requires time and resources, potentially slowing the pace of innovation. In a highly competitive market, this creates perverse incentives:

  • If company A slows down to implement better safety measures, company B might surpass it and capture the market
  • First movers to certain capabilities gain enormous advantages in terms of data, adoption, and market position
  • Investors reward development speed and achievement of technical milestones, not necessarily prudence

This dynamic creates what game theory calls a "race to the bottom": each rational actor has incentives to reduce safety standards, even though collectively this leads to a worse outcome for everyone.

The Absence of International Governance

Unlike other potentially dangerous technologies like nuclear or biological weapons, there isn't yet a robust international governance regime for advanced artificial intelligence. While there have been some attempts (like the European AI Act or various international summits), there's no equivalent of the IAEA (International Atomic Energy Agency) for AI.

This governance vacuum means decisions about how quickly to proceed and what risks to accept are essentially left to individual companies and their leaders, with little public oversight or accountability.

Future Prospects: Possible Scenarios

Given the severity of the alarms raised, it's worth exploring what scenarios might materialize in coming years.

Scenario 1: Voluntary Slowdown

In this scenario, resignations of high-profile researchers and public pressure lead major laboratories to voluntarily slow development, implementing more rigorous safety protocols. This would require:

  • Coordination among major laboratories to avoid some "defecting" by slowing down
  • Independent verification mechanisms for safety claims
  • Possibly regulatory intervention that levels the playing field

Scenario 2: Limited Incident as Wake-Up Call

An AI system causes significant but not existential damage (e.g., financial collapse, critical infrastructure compromise), which serves as a "wake-up call" leading to stringent regulation before truly dangerous systems are developed.

Scenario 3: Status Quo Continuation

The competitive race continues without significant slowdowns. Systems become progressively more capable, approaching or exceeding the existential risk threshold identified by researchers like Coxon and Hubinger.

Scenario 4: Alignment Breakthrough

Unexpected technical advances solve or substantially mitigate the alignment problem, making it possible to develop systems much more capable than current ones without proportionally increasing risks.

What the Technical Community and Public Can Do

Faced with these alarms, what concrete actions are possible?

For Researchers and Developers

  • Responsible whistleblowing: follow Coxon's example in making legitimate concerns public
  • Prioritize safety research: orient careers toward alignment and control problems
  • Demand transparency: require organizations to be transparent about risks and mitigation measures

For Policymakers and Regulators

  • Develop technical expertise: regulators must deeply understand the technology they seek to govern
  • Create accountability mechanisms: independent audit systems, transparency requirements, mandatory impact assessments
  • International coordination: work toward global standards and governance

For the Informed Public

  • Get informed: understand at least broadly the technical risks at stake
  • Pressure on investors and companies: shareholders and consumers can exercise influence
  • Support safety organizations: nonprofit organizations dedicated to AI safety deserve support

Conclusion: A Critical Moment Requiring Action

The alarm raised by Jacob Coxon and confirmed by other leading researchers cannot be dismissed as unfounded alarmism. These are people who have worked directly on the world's most advanced systems, who deeply understand the technical trajectories underway.

Their assessment that there's a non-negligible probability of existential risk within this decade should be a wake-up call for society as a whole. This isn't about stopping technological progress, but ensuring it proceeds in a way that maximizes benefits while minimizing catastrophic risks.

As Coxon emphasizes, "no other human activity carries such a level of danger." This requires a level of caution, coordination, and oversight commensurate with what's at stake. The fact this is happening while some companies involved are preparing for unprecedented market valuations adds urgency to the matter: we need mechanisms ensuring commercial incentives don't prevail over humanity's safety.

The hope is that these alarms, uncomfortable but necessary, will catalyze action before it's too late. As often happens with transformative technologies, we have a window of opportunity, probably brief, to put necessary safeguards in place. What we do or don't do in this window could determine not only the future of artificial intelligence, but the future of humanity itself.

Alessandro Caprai

Alessandro Caprai

AI educator, trainer and developer. Founder of Caprai.dev.

Related posts

Google's Gemini 4 Release Ready
EditorialsGoogle's Gemini 4 Release Ready

Google's Gemini 4 Release Ready Google is preparing to shake up the artificial intelligence landscape once again with the imminent release of Gemini 4...

October 1, 202610 min read
The AI Act and the Literacy Revolution: Two Years Later, Taking Stock
EditorialsThe AI Act and the Literacy Revolution: Two Years Later, Taking Stock

Two years after its entry into force, the European AI Act is proving to be much more than a simple regulatory framework. What many initially considere...

October 1, 20267 min read
Innovation and Artificial Intelligence: Beyond Technological Routine
EditorialsInnovation and Artificial Intelligence: Beyond Technological Routine

There's a misconception creeping through the tech world, and it's becoming increasingly cumbersome. I see it every day: companies waving the flag of i...

September 29, 20266 min read