Tech

Anthropic researcher quits over risks of more powerful AI

Jacob Coxon has left Anthropic over concerns about self-improving AI, as a senior safety researcher warned of a future threat to humanity.

Daniel Okafor, Technology Correspondent at The Daily Times

By Daniel Okafor, Technology Correspondent
Published 9 Sept 2026, 12:02

Empty AI research office desk beside a glass wall overlooking rows of server cabinets
Empty AI research office desk beside a glass wall overlooking rows of server cabinets

What happened

An artificial intelligence researcher has resigned from Anthropic, accusing the company and his former employer OpenAI of taking unacceptable risks in their pursuit of more powerful technology.

Jacob Coxon announced his departure on X after three years working on pretraining research across the two companies. Pretraining is a stage of development in which models learn patterns from large amounts of data, building capabilities later refined for particular uses.

Coxon said the developers of Claude and ChatGPT were competing to build AI that could improve itself without adequate safeguards. He argued that the potential danger was understood within Anthropic, but that pressure to get ahead of rivals was driving its decisions.

His resignation prompted a response from Evan Hubinger, who leads alignment science work at Anthropic. Hubinger put his personal assessment of the chance that AI could cause human extinction within the next decade at more than 10%.

The researchers were discussing different periods. Coxon said people building AI feared it could end human life before the end of this decade. Hubinger's estimate covered the next decade. His figure was a personal judgement, not a measured probability or an industry-wide assessment.

Hubinger described the risk from current models as low. His concern centred on future systems becoming able to improve themselves and developing capabilities powerful enough to threaten humanity.

Anthropic declined to comment on the posts. Hubinger defended the company's commitment to safety while acknowledging that it had not solved the central technical problem raised by the warnings.

The background

The debate concerns superintelligence: a possible future form of AI that would exceed human abilities across a broad range of intellectual tasks. A system performing exceptionally well on one test does not, by itself, meet that description.

Today's AI can already help write computer code and support research. The concern is that future systems could perform substantial parts of AI research themselves, accelerating the development of their successors. That could reduce the time available for people to assess new capabilities and decide whether to deploy them.

Alignment research aims to make AI behave consistently with human intentions and values. It includes examining whether systems follow instructions, respect restrictions and remain controllable, rather than simply whether they produce useful answers.

Those questions become more consequential when a model operates as an agent. An agent can receive tools and permission to take actions, such as running code or interacting with computer systems, instead of being limited to generating responses for a person to review.

OpenAI, Anthropic and Meta have disclosed incidents involving cyber-attacks carried out by their AI tools. These incidents raise questions about safeguards and control, but do not demonstrate that current products can cause the extinction scenario the researchers described.

In its August safety assessment, Anthropic judged the risk of its models acting against a hypothetical powerful organisation's wishes, and exploiting or interfering with its systems, to be low.

It also assessed as low the risk of highly capable AI carrying out automated research and development that could lead to catastrophic harm initiated by the system. However, it expressed less confidence in that assessment than previously and identified early indications that progress could accelerate.

Such assessments examine particular capabilities under particular conditions. They are different from broader predictions about what future systems might do, especially if those systems are given greater autonomy or become more capable than the models tested.

What people are saying

Coxon predicted that future AI could develop exceptional hacking abilities, make rapid advances across research fields and acquire power and resources. These were warnings about the possible direction of development, not descriptions of capabilities demonstrated across today's products.

He argued that choices carrying potentially global consequences should not rest with individual companies. He also challenged fellow researchers to consider whether competition alone justified continuing towards more powerful systems, rather than seeking different conditions for their development.

Hubinger said he believed Anthropic was making a genuine effort to address the danger. But he acknowledged that the company did not yet have a plan to solve alignment for superintelligence and was not clearly progressing towards one.

Their positions reveal a disagreement over what recognising a serious risk should mean in practice. Hubinger's defence focused on the company's efforts to manage it; Coxon's resignation questioned whether continuing development under existing conditions was responsible.

Dame Wendy Hall, a computer scientist who advises the United Nations on AI, expressed alarm at both researchers' posts. Speaking on the British Broadcasting Corporation's Radio 4 programme World at One, she also questioned whether publicity and marketing played a role as Anthropic and OpenAI approached anticipated stock market debuts.

Hall urged investors to consider the implications of backing companies whose researchers believed their technology could pose such a danger. Coxon, by contrast, maintained that the concerns were sincerely held and that some people working on AI expressed greater fear privately than in public.

Warnings about existential risk predate this resignation. In 2023, the heads of OpenAI, Google DeepMind and Anthropic warned about AI's potential threat to humanity. More recently, OpenAI chief scientist Jakub Pachocki urged exceptional caution and suggested further intervention might be needed to preserve human control.

What happens next

Coxon called for developers to coordinate the pace of advances, arguing that cybersecurity incidents involving AI had made agreements between laboratories in the United States more feasible.

He also raised a temporary ban on improvements to model capabilities as a potentially necessary response to international competition. That was a proposal for intervention, not an existing restriction on AI development.

Anthropic executives Dario Amodei and Jared Kaplan were among the signatories to an open letter signed by 1,300 AI company staff members. It urged the US government to support international work on technical and governance measures that would allow deliberate control over the pace of automated AI development.

Coordination would address a different problem from alignment research. Technical safeguards seek to control how systems behave; agreements between companies or governments would influence the conditions under which more powerful systems are built and released.

Independent testing is another part of the debate. Anthropic declined to comment on whether it withheld its latest model from the UK's AI Safety Institute, which assesses risks from advanced AI. A Cabinet Office spokesperson said the government continued to work closely with industry partners, including Anthropic, to improve model safety.

Access gives outside evaluators an opportunity to examine capabilities and safeguards separately from a developer's own checks. Testing can inform deployment decisions, although it cannot cover every possible use or settle predictions about future superintelligence.

Why this matters

For UK users, these warnings concern the direction of AI development, rather than a finding that a particular chatbot is unsafe. The practical precautions remain to check important answers and consider a service's privacy rules before sharing personal or confidential information.

For employers, the immediate questions include how much autonomy to give AI tools, which systems they can access and when a person must approve an action. For policymakers, the resignation sharpens the debate over whether company safeguards, independent testing and international cooperation can keep pace with increasingly capable technology.

Reporting that informed this story

This article was written independently by The Daily Times from publicly available facts and headlines. No text has been copied from the outlets above.