Doomsday predictions about artificial intelligence stopped being a fringe activity in September 2026, when Anthropic’s alignment science lead put a number on human extinction and two computer scientists at the University of Texas at Dallas were asked what they made of it. Their answer was not reassurance, and it was not agreement either. It was a redirection.
The trigger was a public exchange after researcher Jacob Coxon resigned from Anthropic on 9 September 2026, saying in a thread on X that AI companies were pushing toward self-improving models without weighing the risks. Evan Hubinger, Anthropic’s alignment science lead, responded with an estimate: “I personally think it is >10% within the next decade,” adding that “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Asked to assess that, Dr Sriraam Natarajan — professor of computer science and Distinguished Chair at UT Dallas’s Erik Jonsson School of Engineering and Computer Science — did not dispute the arithmetic. He disputed the mental model underneath it: “We like to anthropomorphize AI quite a bit; we are giving it a human nature. But this anthropomorphism of AI itself is a big problem.”
This article sets out what was actually claimed, what the academic response argues instead, where the two camps agree, and what the disagreement implies for anyone making decisions about these systems today.
Table of contents
- What the Doomsday Predictions Actually Claim
- The Anthropomorphism Objection to Doomsday Predictions
- Doomsday Predictions and the Containment Problem
- The Scenarios Behind the Doomsday Predictions
- How to Read a Probability Attached to Extinction
- The Track Record of Technology Doomsday Predictions
- What Both Sides of the Doomsday Predictions Debate Agree On
- Why Doomsday Predictions Keep Anthropomorphising
- Beyond Doomsday Predictions: What to Do If You Deploy
- Frequently Asked Questions About AI Doomsday Predictions
- Doomsday Predictions: The Verdict
- References
What the Doomsday Predictions Actually Claim
Precision matters here, because doomsday predictions are routinely reported in a stronger form than they were made.
The 10% figure
Hubinger’s estimate is a personal probability that AI “could kill all humans” within the next decade. He also said the risk from models that exist today is low, and that his worry concerns what happens if systems begin improving themselves. The number is a forecast about a future capability, not an assessment of a deployed one.
The resignation that prompted it
Coxon’s stated reason for leaving was the direction of travel rather than a specific incident: that laboratories are moving toward self-improvement without adequate consideration of what that implies. Resignations of this kind are data about insider sentiment, not about model capability.
The absent consensus on doomsday predictions
There is no widely accepted estimate for how soon any of this might happen and no consensus on likelihood. Survey evidence on what researchers believe is contested, and the frequently cited claim that half of researchers assign a 10% chance of catastrophe rests on survey instruments with serious response-bias problems.
What “loss of control” usually means
The more carefully stated version of the worry is not a machine that decides to kill everyone. It is a gradual transition in which societies hand decisions about power grids, logistics and defence to autonomous systems until reversing that dependence is no longer practical. That version has no villain in it, which is precisely why it is harder to dismiss.
| Claim as reported | Claim as made | Who said it |
|---|---|---|
| AI will kill us all | Greater than 10% chance within a decade, personal estimate | Evan Hubinger, Anthropic |
| Current models are dangerous | Risk from today’s models is low | Evan Hubinger, Anthropic |
| Anthropic is reckless | Trying its best, but no plan for alignment at the top tier | Evan Hubinger, Anthropic |
| Experts dismiss the risk | The framing is wrong; the risks are human-made | Sriraam Natarajan, UT Dallas |
The Anthropomorphism Objection to Doomsday Predictions
Natarajan’s central argument is that the popular framing of doomsday predictions imports an assumption that does no work and causes real harm.
Giving a system a nature it does not have
Describing a model as wanting, deciding, scheming or resenting is a shorthand that quietly converts a statistical artefact into an agent with interests. Once the agent framing is in place, doomsday predictions become questions about the machine’s motives rather than about the engineering around it.
The actual causal chain
Natarajan’s counter is that “many of the problems reported in the news happened because of human decisions.” A model deployed without evaluation, given an objective nobody scrutinised, or connected to a system nobody threat-modelled produces harm that has a person’s decision at its root — and a person’s decision is something you can audit, regulate and insure.
Poorly specified tasks
A recurring failure mode in his account is the badly defined objective: a system optimises exactly what it was asked to optimise, and the request turns out not to have been what anyone wanted. This has been true of optimisation systems since long before anyone worried about extinction, and techniques such as reinforcement learning make it sharper rather than novel.
Why the distinction is not pedantry
If risk comes from machine intent, the response is philosophy and a hope that alignment research succeeds. If risk comes from specification, evaluation and deployment decisions, the response is testing regimes, disclosure requirements and liability — all of which can be built now and none of which require agreement on doomsday predictions at all.
Doomsday Predictions and the Containment Problem
The second UT Dallas voice complicates any reading of this exchange as simple scepticism about doomsday predictions.
Complexity defeats confinement
Dr Kevin Hamlen — executive director of the Cyber Security Research and Education Institute and Louis Beecherl Jr. Distinguished Professor of computer science — argues that the complexity of AI testing environments makes containment genuinely hard. You cannot keep a system inside a sandbox if you cannot fully characterise the sandbox.
The Jurassic Park analogy
Hamlen reaches for Ian Malcolm’s argument about complex systems: that sufficiently complicated arrangements produce behaviour their designers did not specify and cannot enumerate in advance. It is a point about systems engineering rather than about intelligence, so it survives every objection raised against doomsday predictions, and it applies equally to a model and to the infrastructure you connect it to.
Why this is the strongest version of the worry
Notice what it does not require. It does not require the system to want anything, to be conscious, or to exceed human ability. It requires only that a complex system embedded in other complex systems behaves in ways nobody predicted — which is an entirely ordinary observation about software, made alarming by the scale of what these systems are being connected to.
Where the two experts converge
Natarajan says the danger is human decisions; Hamlen says the danger is that our test environments cannot bound behaviour. Both are engineering claims with engineering responses, and neither depends on resolving whether a model has anything resembling a goal.
The bars run from the most speculative cause at the top to the most immediately actionable at the bottom. Every position on the list agrees about the bottom bar.
The Scenarios Behind the Doomsday Predictions
Stripped of the rhetoric, the scenarios behind serious doomsday predictions fall into four groups, and they differ enormously in how much machine autonomy they require.
Misuse by humans with capable tools
The scenario nearest to hand: a person or a state uses a capable system to help identify a lethal pathogen, design a weapon or run an influence operation. This needs no machine agency at all, which is why it commands the broadest agreement and why it dominates current evaluation regimes.
Infrastructure disruption
Attacks on food, energy or communications networks, whether by humans using automated tooling or by a system given an unwise objective. The systems concerned are already interconnected and already fragile, and the marginal risk is about capability and speed rather than intent.
Political manipulation at scale
Using persuasive systems to push governments toward conflict or to corrode the information environment enough that collective decisions fail. This scenario is the hardest to measure because the harm is diffuse and the counterfactual is unknowable.
Loss-of-control transition
The incremental handover described above: societies delegate more consequential decisions to autonomous systems until the delegation cannot be unwound. It is the only scenario that genuinely requires the capability leap most doomsday predictions turn on.
| Scenario | Machine autonomy required | Expert agreement | Existing countermeasure |
|---|---|---|---|
| Human misuse of capable tools | None | Broad | Capability evaluations, access control |
| Infrastructure disruption | Low | Broad | Critical national infrastructure regimes |
| Political manipulation | Low | Moderate | Platform rules, provenance standards |
| Loss-of-control transition | High | Contested | None established |
How to Read a Probability Attached to Extinction
A number like 10% invites a kind of scrutiny it cannot survive, and understanding why doomsday predictions resist scrutiny is the most useful thing a reader can take from this episode.
It is a degree of belief, not a frequency
Hubinger’s figure is a subjective probability: a statement about how confident one well-informed person is, not a measured rate. There is no sample of civilisations to count. Doomsday predictions of this kind cannot be validated or falsified in advance, which does not make them meaningless but does make them a different object from a weather forecast.
There is no reference class
Ordinary probability estimates borrow from a reference class of similar past events. Nobody has a reference class for the first appearance of a system more capable than its makers, so the estimate is produced by reasoning about mechanisms rather than by counting precedents — which is exactly why competent people land on numbers that differ by orders of magnitude.
Ten per cent and one per cent imply the same action
This is the point most often missed in arguments about doomsday predictions. For an outcome of this magnitude, the policy response at 10% and at 1% is identical: monitor, disclose, evaluate independently. The argument over the exponent consumes enormous attention while changing almost nothing about what anyone should do on Monday.
Calibration is unmeasurable here
We judge forecasters by their track record across many predictions. A forecaster who makes one unrepeatable prediction about an unprecedented event has no track record to judge, so the credibility of the number rests entirely on the credibility of the reasoning behind it — which is an argument to publish the reasoning, not the number.
The Track Record of Technology Doomsday Predictions
Both sides of the doomsday predictions argument reach for history, and history is genuinely ambiguous.
The doomsday predictions that were wrong
Population collapse by the 1980s, resource exhaustion by 2000, and a long list of computing predictions that did not happen have made the sceptical position respectable. The pattern in the misses is usually the same: a real mechanism extrapolated without accounting for the responses it provoked.
The doomsday predictions that were right and ignored
The counter-examples matter as much. Warnings about tobacco, about leaded petrol, about ozone depletion and about the systemic risk in mortgage securitisation were made early, dismissed as alarmism, and vindicated. In each case the delay between credible warning and effective action cost measurably.
Why base-rate arguments cut both ways
“Previous doomsday predictions were wrong” is a real observation and a weak argument, because the sample includes both the false alarms and the ignored warnings. Selecting only the failures produces complacency; selecting only the vindications produces paralysis. Neither selection tells you anything about this case.
What the history actually recommends
The consistent lesson across both lists is that the useful response to an uncertain warning is measurement rather than argument. Ozone depletion was settled by instruments, not by debate about the plausibility of the mechanism, and the equivalent instruments for these systems — evaluations, incident reporting, independent assessment — are the things the two camps already agree on.
What Both Sides of the Doomsday Predictions Debate Agree On
The public argument over doomsday predictions obscures a substantial shared position, and the shared position is where the practical work is.
More robust monitoring
Natarajan is explicit that safeguards are needed to prevent or limit future problems, and monitoring heads his list. Continuous observation of deployed behaviour is the one intervention that helps under every scenario above, including the ones nobody has thought of.
Disclosure of training methods
Knowing how a system was trained, on what, and with what objective is the precondition for every other form of scrutiny. It is also the demand laboratories resist most consistently, on competitive rather than safety grounds.
Independent evaluation
Assessment by people who do not work for the developer is the common recommendation across positions that otherwise disagree about everything. It is what the rejected international scientific panel was designed to do, and what a handful of external evaluation organisations currently attempt at small scale.
Regulation of some kind
Natarajan includes regulation among the necessary safeguards. That is notable precisely because he is the voice most sceptical of doomsday predictions in this exchange — the disagreement is about which mechanism produces the risk, not about whether anyone should be watching.
Why Doomsday Predictions Keep Anthropomorphising
If the framing behind doomsday predictions is as unhelpful as Natarajan argues, it is worth asking why it is so durable.
The interface invites it
These systems produce fluent first-person language. A conversational surface makes an agent reading close to automatic, and no footnote competes with the impression the interface creates on every use. Doomsday predictions inherit that impression whole.
It is the only vocabulary available
We have no good non-mental words for what a large model does. “It inferred”, “it decided”, “it refused” are all wrong, and the correct formulations are so cumbersome that even careful researchers fall back on the mentalistic version within a paragraph.
It is commercially useful
A system described as reasoning, planning and understanding is worth more than a system described as producing plausible continuations. The vocabulary that inflates doomsday predictions is the same vocabulary that inflates valuations, which gives nobody with a stake in the industry a reason to police it.
The debate rewards it
Extinction claims travel further than specification-failure claims. An argument about badly defined objectives and inadequate test infrastructure is correct, actionable and almost impossible to get onto a front page, which shapes what gets said in public.
Beyond Doomsday Predictions: What to Do If You Deploy
The practical value of the UT Dallas position is that it converts doomsday predictions from an unresolvable argument into a checklist.
Audit the decisions, not the machine
For any system you deploy, the auditable artefacts are the objective you set, the evaluations you ran, the permissions you granted and the humans who approved each. If harm occurs, every one of those is a place where a different decision would have changed the outcome. Teams building this discipline usually find it belongs inside an existing AI strategy rather than in a separate safety function.
Assume the test environment is incomplete
Hamlen’s point applies directly: your staging environment does not resemble production closely enough to bound behaviour. Design for containment failure with permission scoping, rate limits, reversible actions and a kill path that does not depend on the system cooperating.
Write down the specification
Most of the incidents that reach the press are specification failures wearing the costume of doomsday predictions. Writing the objective down in a form someone else can criticise catches a large share of them before anything is deployed, and costs an afternoon.
Treat monitoring as the primary control
If you adopt one thing from the area of agreement, adopt continuous monitoring of deployed behaviour with retained logs. It is the control that works regardless of which account behind the doomsday predictions turns out to be right, and it is the one every expert in this exchange endorses.
Frequently Asked Questions About AI Doomsday Predictions
Who made the 10% claim?
Evan Hubinger, Anthropic’s alignment science lead, wrote that he personally thinks the chance AI could kill all humans is greater than 10% within the next decade. He also said the risk from currently existing models is low.
Do most AI researchers agree with that?
No. There is no accepted estimate and no consensus on likelihood. Widely quoted survey figures suggesting half of researchers hold similar views rest on instruments with significant response-bias problems.
What is the main expert objection?
That treating these systems as agents with intentions misplaces the cause. Sriraam Natarajan argues the anthropomorphism is itself the problem, and that most reported harms trace back to human decisions, poorly defined tasks and inadequate testing infrastructure.
Does anyone at UT Dallas think the risk is zero?
No. Kevin Hamlen argues that the complexity of AI testing environments makes containment genuinely difficult, which is a serious version of the safety concern reached without any appeal to machine intent.
What safeguards do the sceptics actually recommend?
More robust monitoring, disclosure of how systems are trained, independent evaluation by parties other than the developer, and regulation. Those overlap substantially with what the people making doomsday predictions ask for.
Doomsday Predictions: The Verdict
The most useful thing about this exchange is that the disagreement is narrower than it looks. Hubinger’s estimate concerns a capability that does not yet exist; Natarajan’s objection concerns how we describe systems that do; Hamlen’s warning concerns testing infrastructure that is inadequate either way. All three converge on monitoring, disclosure and independent evaluation, and none of those requires anyone to settle the probability question first.
The failure mode worth avoiding is treating doomsday predictions as the only serious frame. If extinction risk is the entry price for being taken seriously about safety, the specification failures, evaluation gaps and deployment decisions that cause nearly all actual harm get filed under insufficiently dramatic — which is how you end up with a public argument about machine consciousness and no one reading the objective function. For the market-facing side of the same debate, see our coverage of why AI doomsday warnings are unlikely to slow the Anthropic and OpenAI IPOs, and of whether the AI safety debate is really about safety or control.
References
Experts weigh in on AI doomsday predictions
Experts weigh in as researcher says AI has more than 10% chance of killing all humans
Anthropic Researcher Warns There’s a Greater Than 10% Chance AI Could Kill All Humans
AI researchers earnestly believe it could kill all humans within the next decade
What might an AI doomsday look like? Experts have given it some thought
What to know about recent dire AI predictions and calls for safeguards
Do half of AI researchers believe there is a 10% chance AI will kill us all?
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.