by datastudy.nl

Sunday, September 13, 2026

AI

Anthropic lead warns AI existential risk tops 10 percent

AI existential risk is the chance AI causes human extinction. Anthropic's alignment lead puts that probability above 10 percent within a decade, with no plan to solve it.

Pie chart showing Evan Hubinger's AI existential risk estimate: greater than 10 percent chance AI could kill all humans within the next decade, with 90 percent representing survival probability. The AI existential risk figure exceeds 10 percent per Anthropic's Alignment Science Lead.
Anthropic Alignment Science Lead Evan Hubinger estimates a greater than 10 percent chance AI could cause human extinction within a decade, with 90 percent representing survival probability. Source: Hubinger's post on X, September 9, 2026. Data Today benchmark.

The people building the most advanced AI systems in the world think those systems might kill you, and they are building them anyway. AI existential risk has moved from academic speculation to an on-the-record admission from inside a frontier lab.

On September 9, 2026, Jacob Coxon, a researcher who spent three years doing pre-training work at Anthropic and OpenAI, announced his resignation from Anthropic on X. He accused both companies of "racing straight to self-improving superintelligence and gambling with our lives." The people building AI, he wrote, "earnestly believe that it could kill us all by the end of the decade."

Hours later, Evan Hubinger, Anthropic's Alignment Science Lead, confirmed that Coxon was correct. "We really do earnestly believe AI could kill all humans," Hubinger wrote on X, in a post viewed more than 10 million times. He put his personal estimate at greater than 10 percent within the next decade. Then he added the qualifier that should make anyone building on these platforms stop: Anthropic does "not yet have a plan" to solve alignment for superintelligence and "are not clearly on track to" develop one.

What exactly did the Anthropic researchers say?

Coxon's resignation post was short and unsparing. He said he had spent the past three years doing pre-training research at Anthropic and OpenAI and that "neither company is acting responsibly." The two labs are "locked in a race" to build self-improving superintelligence, he wrote, and are pushing ahead "despite the risk."

Hubinger's response was remarkable for its candor. He agreed that Coxon's characterization was "correct," then went further. He said he worries about self-improving AI and that "it is happening faster than we thought." He acknowledged that the risk from current models is "low" but said he is worried the technology could soon improve itself to the point where it poses an existential threat.

The mechanism Hubinger described is recursive self-improvement: AI systems that build progressively smarter versions of themselves in a loop that could spiral beyond human control. ABC News reported that Anthropic published research on this topic in June 2026, noting that "full recursive self-improvement also might increase the risks of humans losing control over AI systems." The company's own framing: if systems can build their own successors, "the ways we secure them, monitor them, and shape their behavior all grow much more important."

What makes this exchange different from previous AI safety warnings is the source and the specificity. Anthropic was founded in 2021 by former OpenAI researchers who left specifically over safety concerns. The company's Responsible Scaling Policy is one of the most detailed in the industry. When the alignment lead at the lab most associated with safety says the company has no plan and is not on track to make one, that is a structural admission, not a personal opinion.

How fast is the safety warning timeline accelerating?

The Coxon resignation and Hubinger confirmation did not come out of nowhere. They cap a summer of escalating warnings from inside the AI industry, and the pace is itself the story.

In July 2026, more than 1,300 employees of frontier AI companies signed a public letter called "Pacing the Frontier," warning of "a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems." The same ABC News report noted that the letter called on the US government to support an international effort to develop tools that can "deliberately pace the frontier of automated AI development."

The chart below shows the cumulative count of major public AI safety warnings from inside the industry in 2026, rising from zero in May to five by September.

Step chart showing cumulative major public AI safety warnings from inside frontier labs in 2026: 0 in April and May, 1 in June after Anthropic's recursive self-improvement research, 2 in July after the Pacing the Frontier letter with 1,300 signatories, holding at 2 in August, then jumping to 5 in September after Pachocki's caution warning, Coxon's resignation, and Hubinger's confirmation.
Cumulative count of major public AI safety warnings from frontier lab employees in 2026, rising from 0 in May to 5 by September. Events: Anthropic's June recursive self-improvement research, the July Pacing the Frontier letter with 1,300+ signatories, and three September events: Pachocki's caution warning, Coxon's resignation, and Hubinger's confirmation. Source: ABC News, BBC News, The Verge. Data Today benchmark.

After the June research publication and the July letter, September brought three events in rapid succession. OpenAI chief scientist Jakub Pachocki called for "extreme caution" over AI's progress, warning that "more intervention may be needed to ensure humans remain in control of the future." Then Coxon resigned. Then Hubinger confirmed the >10 percent estimate on the same day.

Both Anthropic and OpenAI publicly responded to the Pacing the Frontier letter, according to the same ABC News report. Anthropic said its research on recursive self-improvement "points to the need for tools to deliberately pace the frontier." OpenAI posted that "at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement." Both responses share a tell: they point to a future need for pacing while continuing to accelerate today.

Why should builders care about internal lab warnings?

If you are shipping products on top of Claude, GPT, or any frontier model, this is supply chain risk dressed up as philosophy.

First, regulatory risk is now being driven from inside the labs. When 1,300 industry employees ask the government to step in, and the alignment lead at a major lab says on the record that the company has no plan, lawmakers have cover to act. The US government has so far taken a light touch on AI regulation. Internal admissions like Hubinger's change the political calculus. If you are building a business that depends on unrestricted access to frontier models, budget for the possibility that access gets throttled.

Second, the "no plan" admission has engineering consequences. Anthropic's Responsible Scaling Policy includes safety thresholds that, if crossed, trigger deployment freezes. If the company's own alignment lead says they are not on track to solve the problem, the probability of a self-imposed freeze or a forced slowdown goes up. Your roadmap should account for capability plateaus, not just improvements.

Third, the recursive self-improvement timeline matters. Hubinger said it "is happening faster than we thought." Bill Gates has separately said that AI danger thresholds have already been crossed. When the people inside the labs and the people watching from outside agree that the timeline is compressing, the window for building guardrails into your own systems is shrinking too. If your agents operate with real permissions, sandbox escapes and monitoring failures are the core design problem.

What this means for you:

  • Dependency risk: If Anthropic pauses deployment for safety reasons, your Claude-based features go dark. Diversify across providers or build fallback architectures.
  • Regulatory risk: The Pacing the Frontier letter and Hubinger's admission give regulators ammunition. Expect scrutiny of agentic systems, not just base models. Audit your agent permissions and data access now.
  • Talent risk: Coxon's departure signals that safety researchers will vote with their feet. If your team includes alignment or safety engineers, their concerns about your deployment choices carry more weight this week than last week.
  • Roadmap risk: If recursive self-improvement arrives as fast as Hubinger suggests, model capabilities could jump discontinuously. Build for flexibility in how you call models, not for a specific model's current behavior.

What should you watch in the coming weeks?

The immediate question is whether other Anthropic researchers follow Coxon. One resignation is a data point. Three is a pattern. Watch the X accounts of Anthropic's safety team for the next few weeks.

The second question is whether the Pacing the Frontier movement produces concrete policy. The letter asked for international coordination on AI development pacing. If the US government moves from rhetoric to regulation, the builders who prepared for it will have a head start.

The third question is whether Anthropic's Responsible Scaling Policy actually binds. The policy specifies safety thresholds that trigger pauses. If the company's own alignment lead says they are not on track to solve the problem, does the policy force a slowdown, or does commercial pressure override it? Both Anthropic and OpenAI are reportedly preparing for IPOs. The tension between safety commitments and shareholder expectations is about to be tested in real time.

Finally, watch the recursive self-improvement research itself. Anthropic's June paper and Hubinger's comments suggest the company is tracking this capability closely. If a future model demonstrates meaningful self-improvement, the alignment problem shifts from theoretical to operational overnight. Your agent infrastructure should be designed to handle a model that is significantly smarter on Friday than it was on Monday.

The admission that changes the conversation

Hubinger did not say something the industry has never heard. AI researchers have debated existential risk for years. What is new is the source, the setting, and the specificity. The Alignment Science Lead at the lab most associated with safety, speaking publicly, on the record, with a number attached: greater than 10 percent, within a decade, no plan in hand.

If you are building on these systems, whether you personally agree with the 10 percent matters less than whether your architecture, your business model, and your risk assessment account for a world where the people who make your most critical dependency are telling you they might not be able to control it.

Sources