Recursive self-improvement is the term Anthropic has chosen for the moment an AI system can design, build and train its own successor with no human in the loop. It is the stated subject of a document called When AI builds itself, published by the Anthropic Institute, and the press has reported it as a warning from Anthropic researchers about what future models might do. That framing is fair. It is also worth checking against the document itself, because the text is doing something more specific than warning.
We counted it. The body runs 4,944 words. Forty of its sentences carry a number, and every one of those forty sits in the passages describing how much faster AI development has become. The section headed “What should we do?” — the part holding the actual proposal — runs 465 words and contains no figures whatsoever. The word “risk” appears once in the whole body, in the introduction, as a hyperlink pointing at a different essay on a different website.
None of that makes the warning wrong. Anthropic is describing a real and steep set of curves — its own timeline already places autonomous agents in the present tense — and it says plainly that it may be describing them from inside the acceleration. But a reader deciding what to do on Monday morning should know which half of this document is measured and which half is not. This article separates the two, using only Anthropic’s own published figures and arithmetic performed on them. For more coverage of frontier model releases and claims, see our AI models and tools hub.
Table of contents
- What Anthropic Actually Said About Recursive Self-Improvement
- The Recursive Self-Improvement Evidence Is Quantified to the Decimal
- Counting the Warning in a Recursive Self-Improvement Document
- The Recursive Self-Improvement Prescription Contains No Numbers
- Three Futures, and Which Recursive Self-Improvement Scenario Anthropic Picks
- Research Taste Is the Last Gap Before Recursive Self-Improvement
- What Recursive Self-Improvement Means for Buyers of AI Today
- How to Read a Recursive Self-Improvement Announcement
- Frequently Asked Questions About Recursive Self-Improvement
- References and Further Reading
What Anthropic Actually Said About Recursive Self-Improvement
The document opens by defining its own terms, and the definition is narrower than the headlines suggest. Anthropic is not claiming that recursive self-improvement has arrived. It writes that “we are not there yet, and recursive self-improvement is not inevitable”, then immediately adds that “it could come sooner than most institutions are prepared for”. The claim under examination is about a trend line and a lead time, not about a capability that exists today.
What the piece does assert as fact is that AI is already accelerating the building of AI. That is the evidenced part, and Anthropic supports it with benchmark data plus what it describes as “previously unreported data from within Anthropic” — internal figures on code, review, experiments and research judgement that no outside party could have produced.
The Five Stages That End in a Question Mark
The article lays its history out as five stages. Four of them are dated. The fifth is not.
| Stage | Period | What Anthropic says happened |
|---|---|---|
| Building the first Claude | 2021–2023 | People writing code and docs on laptops |
| Chatbots | 2023–2025 | Short code snippets copied into editors by hand |
| Coding agents | 2025–2026 | Agents write and edit whole files on their own |
| Autonomous agents | Today | Agents run code and delegate hours of work to other agents |
| Closing the loop | 20XX? | Claude could be continuously improved by Claude |
That final row is the honest part of the presentation. Anthropic puts a question mark where a date would go, and the whole recursive self-improvement thesis rests on a gap it declines to size. Four measured stages lead to one unmeasured one.
Where the Recursive Self-Improvement Document Sits
The piece is published under the Anthropic Institute, the company’s policy and research arm, and is co-authored by Marina Favaro and Jack Clark, with a credited list of roughly twenty colleagues who gave feedback. It is not a model card and not a peer-reviewed paper. It is a position document with internal telemetry attached, which is an unusual and genuinely useful thing for a frontier lab to publish.
The Recursive Self-Improvement Evidence Is Quantified to the Decimal
This is the part of the document that earns its keep. Anthropic supplies specific, checkable, dated figures, and several of them are figures no competitor has published.
Benchmarks: Four Minutes to Twelve Hours
The external evidence rests on METR’s time-horizon measure, which asks how long a task a model can reliably finish on its own. Anthropic reports that the doubling time for that horizon has itself shortened, from roughly every seven months to roughly every four — the acceleration is in the acceleration.
The worked examples are stark. In March 2024, Claude Opus 3 handled software tasks that take a human about four minutes. A year later, Claude Sonnet 3.7 managed about an hour and a half. A year after that, Claude Opus 4.6 managed 12-hour tasks. Four minutes to twelve hours is a 180-fold increase in two years, and METR separately found Claude Mythos Preview could work for “at least” 16 hours — “at the upper end of what [METR] can measure without new tasks”.
Inside Anthropic: 80% of Merged Code
The internal numbers are the reason to read the document. As of May 2026, Anthropic says more than 80% of the code merged into its own codebase was authored by Claude, against “low single digits” before Claude Code launched in February 2025. Lines merged per engineer per day held flat across 2021–2024, then rose twice, and by the second quarter of 2026 the typical engineer merged 8x as much code per day as in 2024.
Anthropic caveats this itself, and the caveat is worth quoting: lines of code “measures quantity over quality”, so the 8x figure is “almost certainly an overstatement of the true productivity gain”. A March 2026 poll of 130 research-team employees put the median self-estimated uplift at about 4x, which the company also marks down.
The 52x Speedup and the Human Yardstick
The cleanest measurement in the document is a fixed test Anthropic runs at every model release: take code that trains a small model, make it run as fast as possible while still passing the same correctness checks. In May 2025 Claude Opus 4 averaged a ~3x speedup. By April 2026 Claude Mythos Preview reached ~52x. A skilled human researcher, Anthropic says, needs four to eight hours to reach 4x.
So on this one narrow task the model improved 17.3x in eleven months, and its result is 13x what a skilled human reaches in a working day. Anthropic is careful to say the absolute multiple should not be read as a real-world training speedup — the like-for-like comparison is the point, not the number.
| Measure | Earlier | Later | Change |
|---|---|---|---|
| Task length handled | 4 min (Mar 2024) | 12 hr (2026) | 180x |
| Training-code speedup | ~3x (May 2025) | ~52x (Apr 2026) | 17.3x |
| Share of merged code by Claude | Low single digits (Feb 2025) | >80% (May 2026) | Not stated precisely |
| Lines merged per engineer per day | Baseline (2024) | Q2 2026 | 8x |
| Success on most open-ended tasks | 26% (Nov 2025) | 76% (May 2026) | +50pp / 2.9x |
| Beats human next-step choice | 51% (Nov 2025) | 64% (Apr 2026) | +13pp |
| CORE-Bench reproduction | ~20% (2024) | Saturated | 15 months |
Counting the Warning in a Recursive Self-Improvement Document
Here is where the reading diverges from the reporting. If this document is a warning about what future models with recursive self-improvement capabilities might do, the vocabulary of warning should be somewhere in it. It is almost entirely absent.
One Mention of Risk in 4,944 Words
Across the 4,944-word body, the word “risk” appears once. “Danger” appears zero times. “Catastrophic” appears zero times. “Extinction” appears zero times. “Harm” appears once, and “misalignment” once. Meanwhile “faster” appears eight times, “slow” twelve, “compute” eight and “verify” or “verification” eight.
The Single Risk Mention Is a Link Out
The one appearance is this sentence: “But full recursive self-improvement also might increase the risks of humans losing control over AI systems.” Two things about it. It is hedged twice — “might” and “increase the risks of” rather than a stated risk. And the word itself is a hyperlink, pointing away from the document to a separate essay hosted on Dario Amodei’s personal site. The document does not analyse the risk; it links to somewhere that does.
Recursive Self-Improvement Beats Risk Seven to One
The subject term is used seven times across the page. The risk term is used once. In a piece whose entire news value is the pairing of the two, the acceleration half outnumbers the caution half seven to one. Widen the sample to all four Anthropic Institute pages — the landing page, the launch announcement, the research agenda and this article, 10,309 words together — and “risk” appears 8 times, or once every 1,289 words.
The Recursive Self-Improvement Prescription Contains No Numbers
The document’s final section is called “What should we do?”. It is the part a policymaker or a board would turn to first, and it is by some distance the least quantified thing in the file.
465 Words, Zero Figures
The section runs 465 words. It contains no number, no date, no threshold, no named signatory and no budget. Compare that with the 3,244-word evidence portion, which carries 40 numeric sentences. The document measures the problem to one decimal place and describes the response in prose.
The Conditional That Carries the Recursive Self-Improvement Plan
The proposal itself is a single conditional sentence, and it is worth reading closely: Anthropic says it would slow or temporarily pause “if other developers at or near the frontier also did so in a verifiable manner”. Every load-bearing element there is unresolved. Which developers count as “at or near the frontier” is not defined. What “verifiable” means in practice is the open research problem the Institute has just been created to work on. And the document is candid that a unilateral pause “would change who the front-runner is, but it would not create the wider deliberative process that is currently missing”.
Anthropic also states the hard part honestly: detecting a concealed training run is “much more challenging” than verifying missile silos, because “training runs are far easier to conceal”, “their inputs are general-purpose”, and “the incentive to defect quietly is enormous”. It cites the Intermediate-Range Nuclear Forces Treaty as precedent, then notes those regimes “took decades to build” — and that “we don’t have that long”.
Three Futures, and Which Recursive Self-Improvement Scenario Anthropic Picks
The document sets out three possible futures, and it does not treat them as equally likely.
| Scenario | Anthropic’s stated view | Quantified? |
|---|---|---|
| Trend stalls, today’s AI diffuses widely | “We don’t believe it’s likely” | Partly — Project Glasswing’s 10,000+ vulnerabilities |
| Compounding efficiency, humans still set direction | “We’re likely heading into this scenario” | No figures |
| Full recursive self-improvement | “Plausible” if trends continue | No figures, no date |
The Scenario Anthropic Says Is Most Likely
Note which one it picks. The scenario Anthropic explicitly says the evidence points to is the middle one — substantial automation with humans still choosing the direction — not full recursive self-improvement. The headline capability is presented as plausible and undated, while the near-term forecast is the more modest one. Anthropic’s stated worry about the middle scenario is not loss of control but concentration of capability: it names “authoritarian surveillance of whole populations” and tailored influence operations “at a scale no human team could match”.
It also flags a limit that cuts against its own excitement. Speeding one stage of a process moves the bottleneck rather than removing it — Amdahl’s law — and Anthropic reports hitting exactly that internally, where human code review has become the new constraint on shipping.
Research Taste Is the Last Gap Before Recursive Self-Improvement
The document is explicit that one capability separates today’s systems from a system that could build its successor: research judgement. Choosing which problem matters, which result to trust, when an approach is dead.
51% to 64% Is the Only Judgment Trendline
Anthropic’s measurement here is the most interesting and the most caveated in the file. Researchers took 129 real Claude Code sessions where a human had taken a wrong turn, showed models only the work before the detour, and had a separate model judge whose next step was better. Opus 4.5 in November 2025 beat the human choice 51% of the time; Mythos Preview in April 2026 reached 64%.
Anthropic states the bias plainly: the moments were chosen because the human’s choice had room for improvement, so this “isn’t a like-for-like comparison”. As a control, they ran 127 moments where the human’s move was already strong, and there the models won only about 20% of the time. That control is the single most useful number in the document for anyone trying to calibrate how close recursive self-improvement actually is, and it points the opposite way from the headline.
The Agent Experiment That Beat Two Researchers
The other judgement datapoint is a published experiment in which Claude-powered agents attacked an open AI safety problem end to end. Two human researchers recovered roughly 23% of the available gap in about a week. The agents recovered 97% over 800 cumulative hours and about $18,000 of compute — 4.2x the result, for roughly 10x the labour-hours at about $22.50 per agent-hour. Anthropic notes the result “didn’t transfer cleanly to production-scale models”, and that humans still chose the problem and wrote the scoring rubric.
What Recursive Self-Improvement Means for Buyers of AI Today
Strip out the forecasting and the document still tells a buyer three concrete things, all of which are actionable now regardless of whether the loop ever closes.
What the Recursive Self-Improvement Document Answers for Buyers
Capability on well-specified work is no longer the constraint. Anthropic’s own framing is that “the doing now costs almost nothing in human time”, and its measured bottleneck moved to human review. If you are automating anything, plan capacity for review rather than for production — the same shift we covered when OpenAI said it had built an automated research intern.
Second, verification is becoming the scarce skill, inside labs and outside them. Anthropic reports that an automated reviewer would have caught roughly a third of the bugs behind past production incidents on claude.ai — code written by engineers it describes as “among the best in the world at building these systems”.
Third, diffusion is already the near-term story. Anthropic’s own most-likely scenario needs no new capability at all: a 100-person company doing the work of a much larger one, because each person “will sit atop a pyramid of agents”. That is a staffing and autonomous agent design question available today.
What It Does Not Answer
It gives no date for the closed loop, no threshold that would trigger a pause, no definition of “near the frontier”, and no independent verification of any internal figure. Every internal number is self-reported and unaudited, which Anthropic does not hide but also cannot fix alone. Governance readers should note how thin the accountability layer is by comparison — a gap we looked at when OpenAI’s preparedness team was reportedly disbanded.
How to Read a Recursive Self-Improvement Announcement
The counting method used in this article is reusable, and it is cheap.
Three Questions to Ask of Any Recursive Self-Improvement Claim
First, where do the numbers live? If the evidence section is quantified and the mitigation section is not, the document is a capability report with a safety preface, whatever the headline says. Second, is the alarming word load-bearing or decorative? A term used once, hedged twice and hyperlinked away is not an argument. Third, which scenario does the author actually endorse? Here the endorsed forecast is explicitly the middle one, not the dramatic one.
Applied to When AI builds itself, the result is a document that is unusually honest about its evidence, unusually vague about its remedy, and considerably less alarmed than the coverage of it. Anthropic never uses the abbreviation “RSI” anywhere on the page — that compression belongs to the reporting, not the source.
Frequently Asked Questions About Recursive Self-Improvement
Has Anthropic said recursive self-improvement has been achieved?
No. The document states plainly that “we are not there yet” and that it “is not inevitable”. What Anthropic claims is that AI is already measurably accelerating AI development, and that the gap to a closed loop is research judgement.
What is the strongest single number in the document?
Probably the internal code share: more than 80% of merged code authored by Claude as of May 2026, up from low single digits in February 2025. Anthropic notes leadership has publicly estimated 90% or more including scripts, and that its own 80% figure is the more conservative production measure.
Does the document propose a pause?
Not unconditionally. It says a pause would be good to have as an option, and that Anthropic would slow down if other frontier developers did so verifiably. The verification machinery that conditional depends on does not yet exist; building it is the Institute’s stated agenda.
Is the 52x speedup a real-world result?
No, and Anthropic says so. It is a fixed internal benchmark, and the company warns the absolute multiple “should not be read as a real-world training speedup”. The useful comparison is like-for-like: ~3x to ~52x across models, against ~4x for a skilled human in four to eight hours.
References and Further Reading
Anthropic — When AI builds itself
Anthropic — Announcing the Anthropic Institute
Anthropic Institute — Research Agenda
Anthropic Alignment Science — Automated Weak-to-Strong Researcher
Anthropic — Automated Researchers Can Help Mitigate Alignment Failures
METR — Measuring AI Ability to Complete Long Tasks
METR — Measuring the Impact of Early-2025 AI on Developer Productivity
CORE-Bench — Computational Reproducibility Agent Benchmark
Anthropic — Project Glasswing
Dario Amodei — The Adolescence of Technology
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.