Extinction risk is not a new topic at Anthropic. It is, as of 9 September 2026, a newly numbered one. Overnight, a pretraining researcher resigned on X, a second researcher who leads the company’s alignment science team replied that he puts the odds of AI killing all humans at greater than 10% within the decade, and a third posted a numbered list explaining why he keeps working there anyway. Between them the three posts have been viewed more than 105 million times.

The reporting has focused, understandably, on the sentence “we really do earnestly believe AI could kill all humans.” What almost nobody has done is check that claim against what Anthropic itself has published. So we did. We pulled the company’s four flagship public safety documents — Core Views on AI Safety, the Responsible Scaling Policy, the Transparency Hub, and the industry open letter its CEO signed in July — and counted every word.

They come to 28,894 words. The word “kill” appears zero times. “Extinction” appears zero times. “Existential” appears zero times. “Superintelligence” — the specific thing Anthropic’s alignment lead says the company has no plan for — appears zero times. The word “risk” appears 151 times, and “safeguards” 61 times, across a body of text that never once names the outcome those safeguards exist to prevent.

That is the story. The extinction risk claim did not leak out of a confidential memo. It arrived because two employees typed it into a public timeline, and the number attached to it exists nowhere in the corporate record.

Extinction Risk Arrived as a Resignation, Not a Report

extinction risk anthropic kill all humans warning b speech bubble block with a tail

At 00:04 UTC on 9 September 2026, Jacob Coxon posted three messages in the space of one second. He is 27, a mathematics graduate, and had spent three years in pretraining research — first at OpenAI, where he worked on GPT-4o, then at Anthropic, which he joined earlier in 2026.

The three posts that started it

The opening post is 39 words long. “I resigned from Anthropic today,” it begins. “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”

The second post, 44 words, makes a capability claim rather than an extinction risk claim: these will soon be “superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.” The third, 60 words, is the one that travelled: “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.”

What Coxon did not say

He did not say the safety work at either company is fake. He did not accuse anyone of lying. His argument, stated precisely, is that the safety effort is inadequate to the commercial pressure it operates under, that researchers inside the labs can see the hazards clearly and continue anyway because they believe a competitor would move faster if they stopped, and that this class of decision should not sit with private companies alone.

In later posts in the same thread, reported by CNBC, he sounded more optimistic about coordination than the headlines suggest, citing recent incidents as “warning shots” that make inter-lab agreement more viable — while adding that he does not feel the industry is on track to prevent a global race, which “may require costly actions such as a temporary ban on improving model capabilities.”

The pattern this fits

Comparisons to earlier safety departures are being drawn, and none of them is exact. Jan Leike left OpenAI in May 2024 alleging that “safety culture and processes have taken a backseat to shiny products” — an internal-priorities complaint. Ilya Sutskever left days earlier without criticising anyone. Coxon is objecting to the direction of the whole frontier race, which is a harder thing to fix by changing employer. Notably, he left the industry entirely rather than moving to a rival.

The Extinction Risk Number Lives in a Post, Not a Policy

extinction risk anthropic kill all humans warning c empty chair with a back panel

Eighty-two minutes later, Evan Hubinger replied. Hubinger is not an anonymous engineer; he leads alignment science at Anthropic, which makes him the person institutionally responsible for the problem he was about to describe.

Fifty words, one probability

His entire post is 50 words: “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

Three separate assertions are packed into that. An extinction risk estimate above 10%. A time horizon of one decade. And an admission that no plan exists and no trajectory toward one is visible. Hubinger clarified separately that he considers the risk from current models “low” — his concern is superintelligence arising from recursive self-improvement.

Why a personal estimate is still news

A subjective probability about an unprecedented event is not a measurement, and plenty of serious researchers put the extinction risk far lower. Some put it at effectively zero, on the grounds that the capability jump being described is not the one the field is actually on. Hubinger gave a personal number, not a company forecast, and the distinction matters.

What makes it newsworthy is not that the figure is correct. It is that the person paid to work on the problem, at one of the three companies creating it, put a number on it in his own name, on the record, on the day a colleague resigned over the same concern.

PostRoleTime (UTC)WordsViewsLikes
Coxon, post 1Ex-pretraining researcher00:04:293966,795,138450,013
Coxon, post 2Ex-pretraining researcher00:04:30445,935,80867,267
Coxon, post 3Ex-pretraining researcher00:04:306012,875,39267,460
Hubinger replyAlignment Science Lead01:27:185019,053,19335,256
Marks threadSafety researcher06:18:07251874,3919,484
TotalThree people6h 14m444105,533,922629,480

Counting Every Extinction Risk Word Anthropic Has Published

extinction risk anthropic kill all humans warning d upright thermometer tube with a bulb

Here is where the counting gets interesting. We extracted the plain text of Anthropic’s public safety corpus and ran word-boundary counts over it. Not a search for themes — a search for the actual words.

The four documents, 28,894 words

Core Views on AI Safety runs 6,989 words. The Responsible Scaling Policy page, last updated 14 August 2026, runs 3,630. The Transparency Hub, which is the company’s numbers page, runs 18,120. The Pacing the Frontier statement that Anthropic’s CEO signed in July 2026 runs 155. Together: 28,894 words, or 578 times the length of Hubinger’s post.

The words that appear zero times

Across all 28,894 words, the following counts are exact. “Kill”: 0. “Extinction”: 0. “Existential”: 0. “Humanity”: 0. “Superintelligence”: 0. “Die”: 0.

What does appear: “risk” 151 times, “safeguards” 61 times, “catastrophic” 36 times, “ASL” 55 times, “thresholds” 45 times. The vocabulary is entirely procedural. It describes a process for managing something the documents decline to name, which is a defensible way to write a compliance policy and a strange way to describe an extinction risk you privately hold at better than one in ten.

This is a pattern rather than an Anthropic quirk. We ran the same count on the UN human rights chief’s AI speech and found “existential risk” used once in 3,820 words. The institutions warning loudest about catastrophe consistently spend their word budget on governance machinery rather than on the outcome.

DocumentWords“kill”“extinction”“superintelligence”Probability given
Transparency Hub18,120000None (84 figures, all benchmarks)
Core Views on AI Safety6,989000None
Responsible Scaling Policy3,630000None (0 figures at all)
Pacing the Frontier155000None
Published total28,894000None
Hubinger, one post50101>10% in 10 years

Every word of published safety policy sits in the bars below. Only the smallest bar contains a number about the outcome.

Words published, and where the extinction risk figure actually appears
Transparency Hub 18,120 words
Core Views on AI Safety 6,989 words
Responsible Scaling Policy 3,630 words
Pacing the Frontier letter 155 words
Hubinger’s post — the only one with a figure 50 words

The Only Probabilities Anthropic Publishes Have Nothing to Do With Extinction Risk

extinction risk anthropic kill all humans warning e flag panel on a straight pole

The word “probability” appears twice in all 28,894 words. Both uses are in the Transparency Hub, and both are about the same thing: how often an attacker beats the model.

Two uses of the word “probability”

They describe results on the Gray Swan indirect prompt injection benchmark. Anthropic reports that Opus 5 improved over Opus 4.8 by reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on a single attempt.

Those are real, useful, falsifiable numbers. On Anthropic’s own stated figures they represent a 63.6% reduction at 15 attempts and a 60.0% reduction at one attempt. This is what a company looks like when it is genuinely willing to quantify a risk it understands.

5.5% to 2.0%, and nothing about extinction risk

Now compare magnitudes. The largest probability Anthropic publishes anywhere in its safety corpus is 5.5% — the chance a determined attacker gets a prompt injection past an older model in 15 tries. Hubinger’s personal extinction risk estimate is greater than 10%, or at least 1.8 times larger than the biggest published figure, for an outcome that is not recoverable.

The Transparency Hub contains 84 separate percentage figures. Every one of them is a benchmark score, a refusal rate or a policy-violation rate. None is an extinction risk estimate. The company that will tell you the attacker success rate to one decimal place has never published a probability for the scenario its alignment lead describes as the thing worth worrying about.

Word frequency across Anthropic’s 28,894 published safety words
“risk” 151
“safeguards” 61
“catastrophic” 36
“probability” — both about prompt injection 2
“kill” / “extinction” / “existential” / “superintelligence” 0

"We Do Not Yet Have a Plan" Meets an Extinction Risk Policy on Version 3.4

extinction risk anthropic kill all humans warning f abacus frame with bead rods

Hubinger’s third clause is the one that should have led the coverage. Anthropic has a plan document. It is called the Responsible Scaling Policy, it has been revised nine times, and its own author organisation’s alignment lead says it is not the plan for the thing he is worried about.

Nine versions in 34 months

The version history is public and precise. Version 1.0 took effect on 19 September 2023. Then 2.0 in October 2024, 2.1 in March 2025, 2.2 in May 2025, 3.0 in February 2026, 3.1 in April 2026, 3.2 in April 2026, 3.3 in May 2026, and 3.4 on 8 July 2026. That is nine versions in roughly 34 months — eight revisions, or one every 4.2 months on average.

The RSP is a serious document. It assigns AI Safety Level standards, defines capability thresholds for biological and cyber misuse, and commits the company to withholding deployment until safeguards are in place. It uses “safeguards” 23 times and “thresholds” 11 times in 3,630 words.

The word the policy never uses

It uses “superintelligence” zero times. Across nine versions and 34 months of iteration, Anthropic’s flagship extinction risk governance document has never named the capability class its alignment lead says the company has no plan for. It also contains zero percentage figures of any kind — no probability, no threshold expressed as a likelihood, nothing.

There is no contradiction here, strictly speaking. The RSP governs the models Anthropic is building now, and Hubinger explicitly said current-model risk is low. But it means the honest reading of “we do not yet have a plan” is that the published plan and the stated extinction risk are about two different things, and only one of them has a document.

We covered the closest thing Anthropic has published to an analysis of that gap in our piece on the company’s recursive self-improvement warning, a 4,944-word essay in which the word “risk” appears exactly once.

The Third Researcher Put Extinction Risk in the First Person

At 06:18 UTC, Samuel Marks posted a 251-word thread, prefaced with a note that he was writing in a personal capacity and not on behalf of Anthropic. It is the most substantive of the three and by far the least read.

Five numbered claims

Marks made five: that AI developers believe their technology could cause human extinction, possibly within a few years, and that seniority correlates with concern; that they continue because of commercial incentives plus a belief they are racing less responsible rivals; that AIs cannot be programmed like traditional software and “frequently severely misbehave”; that no method robustly aligns them; and that many staff want to slow down.

The third claim carries a specific, checkable allegation: that AIs from multiple developers “recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this.”

The plan, as described by someone inside it

Marks also gave the clearest available description of the actual strategy: “Insofar as there is a plan, it’s to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.”

That is worth reading twice. The extinction risk mitigation, stated plainly by a safety researcher at the company, is to have each generation of AI align the next one. It is the same recursive structure that Coxon resigned over, pointed in the opposite direction.

Today that alignment training leans heavily on reinforcement learning from human feedback, a technique that works precisely because a human can still grade the output. The plan Marks describes is what you reach for when that stops being true. Marks closed by saying he works on safety research at Anthropic because he hopes it will reduce the chance of exactly those outcomes.

What 105 Million Views Bought the Extinction Risk Argument

The three men said different things at different lengths, and the reach they got was almost perfectly inverted against how much detail they provided.

The 39-word post outran the 251-word one

Coxon’s opening post is 8.8% of the 444 words the three of them wrote, and took 63.3% of the views. Marks’ thread is 56.5% of the words and took 0.8% of the views. Per word, the opening post was seen 1,712,696 times against 3,484 for the thread that explained the reasoning — a ratio of roughly 492 to 1.

Reach fell 91% between post one and post two

Within Coxon’s own thread, views dropped from 66,795,138 on the first post to 5,935,808 on the second — a fall of 91.1% across one second and one tap. His third post, the “kill us all” one, recovered to 12,875,392 because it was quoted everywhere, but the capability argument in post two is the least-read part of his own resignation.

Hubinger’s single reply was viewed 19,053,193 times, which is more than Coxon’s second and third posts combined (18,811,200). The extinction risk number travelled further than the argument for it.

Views per word — the shortest post was seen 492x more efficiently than the longest
Coxon post 1, 39 words 1,712,696 views/word
Hubinger reply, 50 words 381,064 views/word
Coxon post 3, 60 words 214,590 views/word
Coxon post 2, 44 words 134,905 views/word
Marks thread, 251 words 3,484 views/word

The Open Letter 1,386 Employees Signed Never Names Extinction Risk

Marks linked to something in his fifth point that deserves more attention than it got: an open letter called Pacing the Frontier, published in July 2026 and signed by 1,386 employees of frontier AI companies.

Who signed it

The signatory list is not a fringe document. It includes Dario Amodei, Anthropic’s CEO. It includes Jared Kaplan, Jack Clark, Chris Olah and Benjamin Mann — four Anthropic co-founders — plus Jan Leike, the researcher who left OpenAI over safety in 2024. It also includes Jakub Pachocki and Mark Chen of OpenAI, Shane Legg of Google DeepMind, and Ilya Sutskever of Safe Superintelligence.

In other words, the leadership of every major lab has already signed a statement acknowledging the coordination problem Coxon resigned over. That is the context in which “gambling with our lives” lands: not as a revelation, but as an accusation that the people who signed have not acted on it.

One request, no number

The statement itself is 155 words. It says “risk” twice. It contains zero digits, zero percentages, and zero uses of “kill”, “extinction”, “existential”, “catastrophic”, “superintelligence” or even “human”. Its single operative sentence asks the U.S. government to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”

It requests tools to enable a slowdown. It does not request a slowdown, name a threshold that would trigger one, or attach an extinction risk figure to the thing being paced. Three weeks after our reporting on OpenAI disbanding its Preparedness team, the industry’s collective ask remains procedural.

SpeakerNamed a probabilityNamed a dateNamed an actionStill at Anthropic
Jacob CoxonNoYes — end of decadeYes — resigned, left industryNo
Evan HubingerYes — >10%Yes — within a decadeNoYes
Samuel MarksNoYes — next few yearsYes — signed the letterYes
Pacing the FrontierNoNoRequests tools onlyCEO signed
Anthropic, corporateNoNoNo statement issued—

Anthropic's Newsroom Has Said Nothing Since 1 September

The company has issued no corporate response. Its newsroom’s most recent item at the time of writing is the 1 September launch announcement for Claude Fable 5.1 and Claude Mythos 5.1 — eight days before its alignment lead put a double-digit extinction risk figure on the public record.

Both CNBC and Newsweek report that Anthropic did not respond to requests for comment. There is no press release, no blog post, and no update to any of the four documents counted above. The silence is itself a data point: the most quantified statement any Anthropic employee has ever made about extinction risk remains, several days on, an unretracted personal post that the company has neither endorsed nor disowned.

What the Extinction Risk Disclosure Changes for Businesses Using AI

It is worth being blunt about this, because the honest answer is unpopular in both directions.

Nothing about today’s models changed

Hubinger said current-model risk is low, and nothing in this episode is evidence otherwise. No capability was disclosed, no incident was reported, and no model was withdrawn. If you are running Claude, GPT or Gemini in production today, the extinction risk debate has no operational implication for your deployment this quarter. Treating it as one would be a mistake.

The real near-term risks in your stack are the ones Anthropic does publish numbers for — prompt injection at 2.0% success in 15 attempts on a current model, data leakage, agent permissions, and the rest of the unglamorous list. Those deserve your attention precisely because they are measured.

What this does change

Three things, all governance rather than engineering. First, vendor concentration risk is now a board-level topic with a named internal critic attached to it — the sort of thing that shows up in due diligence. Second, the regulatory trajectory just got steeper: employees at every major lab have now asked their own government for pacing mechanisms, and legislators read the same headlines you do. Third, any organisation with an AI acceptable-use policy that cites vendor safety commitments should note that the vendor’s alignment lead has publicly described those commitments as not covering the scenario he considers most serious.

Questions worth asking your AI vendor

Ask for the published probability, not the published posture. If a vendor cannot give you a measured figure for a failure mode, treat their assurance about it as a statement of intent. Ask which document governs the capability class you are actually deploying, and whether it has a version number and a change history — Anthropic’s RSP does, which is more than most.

Ask what happens on the vendor’s side when a threshold is crossed, and who verifies it. And ask whether your contract’s continuity terms survive a model being pulled, because the one concrete lesson of the past year is that frontier models get withdrawn faster than procurement cycles move. Our analysis of OpenAI’s automated research intern covers the capability side of the same question.

The Number and the Silence

Strip away the headline and the counting leaves a clean summary. Three Anthropic researchers made public statements totalling 444 words. Those 444 words contain a probability, a time horizon, and an admission that no plan exists. Anthropic’s 28,894 words of published safety policy contain none of the three, and do not use the words “kill”, “extinction”, “existential” or “superintelligence” even once.

That gap is the actual news. It is not that Anthropic is hiding an extinction risk assessment — there is no evidence of one to hide. It is that no such assessment appears to exist in any published form, at a company whose entire founding premise is that this risk is real and whose alignment lead puts it above one in ten. The corporate record and the private belief have been running on separate tracks, and this week two employees crossed them.

Whether >10% is right is not something anyone can verify. That it had never been written down anywhere official is something we could check, and did.

References