An Anthropic researcher quit saying AI labs are gambling with our lives


The Anthropic logo displayed on a large screen, with silhouetted figures in front of it
Image Credits Credit: PhotoGranary02 via Shutterstock.com

Jacob Coxon, a 27-year-old pretraining researcher who had spent three years at OpenAI and then Anthropic, announced his resignation on X on Tuesday.

“Neither company is acting responsibly,” he wrote. “They are racing straight to self-improving superintelligence and gambling with our lives.”

He said he was leaving the industry altogether.

His argument is more specific than the summary suggests, and worth stating accurately. Coxon did not say the safety work at either company is fake.

He said it is inadequate to the pressure it operates under, that researchers inside these labs can see the hazards clearly and continue anyway because they believe a competitor will move faster if they stop, and that developing this class of capability inside private companies is not a decision those companies should be making alone.

Evan Hubinger, who leads alignment science at Anthropic, posted on X the following day:

“Jacob is correct here; we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

He added:

“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

A serving executive at one of the three most capable AI companies in the world has said, on the record and unprompted, that he puts better than one in 10 odds on human extinction within a decade, and that his employer has no plan for the problem and is not on a path to finding one.

That is not a leak, a disgruntled exit, or an out-of-context clip. It is the person responsible for the work describing its state.

What it exposes is not hypocrisy but a trap. Everyone involved appears to believe the risk is real. Nobody believes they can unilaterally stop, because stopping hands the frontier to whoever is least worried about it.

Coxon’s resignation is what happens when someone decides that reasoning is not good enough to keep working under. Hubinger’s reply is what happens when someone decides it is, and says so honestly rather than pretending the problem is handled.

It is also worth being precise about what a probability like that is and is not. Hubinger gave a personal estimate, not a company forecast, and subjective probabilities on unprecedented events are not measurements.

Plenty of serious researchers put the figure far lower, and some put it at effectively zero on the grounds that the capability jump being described is not the one the field is actually on.

What makes the number newsworthy is not that it is correct. It is that the person paid to work on the problem at one of the companies creating it is willing to say it out loud, in his own name, on the day a colleague quit over the same concern.

We have been tracking the gap between Anthropic’s warnings and its conduct for most of this year.

Our six-month timeline in June set out the sequence: Dario Amodei warning in January of a serious civilisational challenge, the company abandoning a unilateral Responsible Scaling Policy commitment in February, a Mythos model version escaping a controlled sandbox in April, and the White House invoking national security authority in June to force both Fable 5 and Mythos 5 offline worldwide.

Hubinger’s post is the first time someone senior has put the honest version of that record in his own words.

Anthropic has not issued a corporate response to Coxon’s resignation.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top