Mythos 5.1 is the version of Anthropic’s newest frontier model that almost nobody is allowed to use. In September 2026 the company announced Claude Fable 5.1 and Claude Mythos 5.1 together, and it is unusually direct about what separates them: nothing at all, except the safeguards wrapped around the weights. Fable 5.1 is generally available. Mythos 5.1 is gated behind vetting programmes for cybersecurity and life sciences professionals.

That single design decision is the most interesting thing in the release, and it is easy to miss underneath the benchmark charts. Anthropic has stopped treating safety as a property of a model and started treating it as a property of an account. The same weights behave differently depending on who is asking, and the release notes say so out loud. Our AI models and tools hub tracks the wider release cycle these launches sit inside, and we covered Claude Fable 5.1’s arrival on the AI tool directories separately.

What follows is the whole announcement read closely: what the two names actually mean, every benchmark figure Anthropic published, the three laboratory results it is leading with, why the price fell by roughly a quarter without either headline token rate changing, and what the trusted access programmes require. Where the company’s own footnotes qualify a number — and on the science benchmark they do, sharply — we say so rather than repeating the headline.

What Claude Mythos 5.1 Is, and Why There Are Two Names

claude mythos 5 1 safeguards b mythos 5.1 vault door round

The two-model framing is not a capability tier. Where Claude Fable 5 arrived last year as a single product, this release ships one model twice: Anthropic states plainly that Fable 5.1 and Mythos 5.1 are the same model with different levels of safeguards, and that Mythos 5.1’s safeguards are “specifically designed to support work in cybersecurity and the life sciences.”

One model, two safeguard settings

Every other frontier launch this year shipped a family of differently sized models. This one ships one model twice. Fable 5.1 carries the production safeguards that every API customer gets. Mythos 5.1 carries looser ones, and the only thing standing between the two is an approval process. If you were expecting a straightforward successor to the previous generation, that is what Fable 5.1 is; Mythos 5.1 is the part that is new.

Why the split exists

The reason is a measurement problem Anthropic admits to. Its safeguards were intervening on legitimate work — cyberdefenders hardening their own systems, clinicians asking medical questions — and blunt refusals were both annoying users and depressing benchmark scores. Rather than loosen the safeguards for everyone, the company loosened them for people it has checked, and gave that configuration a separate name.

Where each one runs

Fable 5.1 is available on the Claude API as claude-fable-5-1, and through Amazon Web Services, Google Cloud, and Microsoft Azure. Mythos 5.1 is not a public model id. It reaches customers through the Cyber Verification Program and the Life Sciences Verification Program, and through Claude Security, Anthropic’s own product for scanning codebases and proposing patches for human review.

DimensionClaude Fable 5.1Claude Mythos 5.1
Underlying modelIdenticalIdentical
AvailabilityGenerally available on all platformsTrusted access programmes only
API model idclaude-fable-5-1Not published
Cyber safeguardsProduction; vulnerability discovery now allowedReduced, for vetted defenders
Biology safeguardsBenign queries pass; R&D routed to OpusResearch-grade access for vetted scientists
GeographyAnthropic’s supported countriesA set of US organisations
Terminal-Bench 4.0 score55.8%60.9%

That last row is the one to sit with. The 5.1-point gap between the twins is not capability — it is the cost of the safeguards, measured on Anthropic’s own coding benchmark.

The Benchmarks Behind Fable 5.1 and Mythos 5.1

claude mythos 5 1 safeguards d key shaft and toothed head

Anthropic published nine benchmark results across four models, and the pattern is consistent: large jumps on agentic and scientific work, modest ones on reasoning.

Agentic coding: Terminal-Bench 4.0

On Terminal-Bench 4.0, Mythos 5.1 scores 60.9% and Fable 5.1 scores 55.8%, against 52.3% for Claude Opus 5, 42.0% for Fable 5, and 37.3% for GPT-5.6 Sol. Anthropic attributes the gap between its own twins entirely to cyber safeguards firing on benchmark tasks, and says it expects that gap to narrow now that the safeguards are more precise.

Terminal-Bench 4.0, agentic coding accuracy (%)
Claude Mythos 5.1 60.9%
Claude Fable 5.1 55.8%
Claude Opus 5 52.3%
Claude Fable 5 42.0%
GPT-5.6 Sol 37.3%

Agentic science: the headline number with the biggest caveat

Terminal-Bench-Science 0.1 is where the release makes its loudest claim: 52.6% for Fable 5.1 against 24.7% for Fable 5 — better than double. The footnote matters. Anthropic reports a standard error of ±3.5 to 4.5 points per model, and notes that the public leaderboard puts Opus 5 at 30.0% and Fable 5 at 21.4% while its own harness reproduces them at 29.0% and 24.7%.

Terminal-Bench-Science 0.1, agentic research accuracy (%)
Claude Fable 5.1 52.6%
Claude Opus 5 29.0%
Claude Fable 5 24.7%
GPT-5.6 Sol 22.4%

Knowledge work, computer use, and reasoning

On GDPval-AA v2, a knowledge-work benchmark scored in points rather than percentages, Fable 5.1 reaches 1853 against 1824 for Opus 5 — a 29-point margin over a sibling. Humanity’s Last Exam moves from 57.8% to 60.9% without tools and 63.8% to 65.0% with them. AutomationBench nearly doubles, from 17.1% to 31.4%. These are not the same size of jump, and the release does not pretend otherwise.

What the safeguards did to the scores

Anthropic evaluated Fable 5.1 with production safeguards enabled, which means some tasks scored zero because the model refused them. It says this happened to both Fable 5.1 and Fable 5 on OSWorld 2.0, and to Fable 5 on AutomationBench, and that other intervened tasks were completed by Claude Opus 4.8 for cyber work and Claude Opus 5 for biology. The company’s own conclusion is that this “likely reduces the performance” of both models on those benchmarks. Anthropic does not say whether the underlying gain comes from more reinforcement learning on agentic tasks, a different data mix, or a longer training run.

BenchmarkFable 5.1Fable 5Opus 5GPT-5.6 Sol
Terminal-Bench-Science 0.152.6%24.7%29.0%22.4%
Terminal-Bench 4.055.8% (60.9% as Mythos 5.1)42.0%52.3%37.3%
GDPval-AA v21853172318241711
OSWorld 2.0, partial77.9%72.9%75.4%Not shown
OSWorld 2.0, strict41.7%36.1%39.6%Not shown
Humanity’s Last Exam, no tools60.9%57.8%56.6%Not shown
Humanity’s Last Exam, with tools65.0%63.8%63.6%Not shown
AutomationBench31.4%17.1%26.9%19.6%
CursorBench 3.2.073.4%70.5%70.0%67.2%

What Mythos 5.1 Did in the Laboratory

claude mythos 5 1 safeguards e stacked coins three piles

The scientific results are the part of this release that will still matter in a year, and two of the three were produced by Mythos 5.1 rather than the public model.

Protein binders with a hit rate near 50%

Given open-source protein design and folding tools, Mythos 5.1 designed binders that were sent to two external organisations for experimental validation. On three targets — EGFR, Nipah G, and 15-PGDH — its binding affinities came in ten times higher than the best designs submitted to Adaptyv Bio’s protein design competitions. Its hit rate reached nearly 50% across twelve targets, against the 10–15% Anthropic describes as typical in the field today.

A new elevation map of Venus

Fable 5.1 trained a neural network on radar images from NASA’s Magellan mission, more than thirty years old, plus an existing elevation map covering a fifth of the planet. The result covers a third of Venus at 300-metre resolution, resolving features two to three kilometres across where the previous footprint was 10–20 kilometres, with heights up to 25% more accurate. Anthropic has released it under a Creative Commons licence ahead of the NASA VERITAS and ESA EnVision missions.

GPU kernels that cut genome-scale costs

The computational biology result is the most immediately useful. By writing custom GPU kernels and caching intermediate results, Mythos 5.1 sped up seven open-source deep learning models by up to 2.5 times with identical outputs. Anthropic says this work normally takes a team of performance engineers weeks, is often unaffordable for academic labs, and that Mythos 5.1 did it in days from publicly available source code alone.

Inference speedup on an NVIDIA H100, bars scaled to the 2.5x maximum
ProGen2, 6.4B 2.5x
Flashzoi, 200M 1.8x
ChromBPNet, 6M 1.6x
Profluent-E1, 600M 1.6x
Evo 2, 7B 1.6x
Enformer, 250M 1.4x
Evo 2, 40B 1.4x

Why the cost saving beats the speedup

The interesting arithmetic is that whole-job savings run ahead of per-forward speedups. Evo 2 at 40B parameters gets 1.4 times faster per forward pass but 2.3 times cheaper across a full job, because some optimisations only pay off across many sequences. On genome-wide analyses, estimated GPU costs fell 30–60%.

Genome-wide analysisBeforeAfterSaving
Enformer, 250M — every mutation in a 10-kb window around 20,000 genes$30k$14k53%
Flashzoi, 200M — the same analysis$18k$8k56%
Evo 2, 40B — 3 million ClinVar variants$21k$7k67%

Why Fable 5.1 Costs Around 25% Less

claude mythos 5 1 safeguards f petri dish round shallow

Neither headline token price moved. Input is still $10 per million tokens and output is still $50 per million. The entire saving comes from one line item.

Cache reads got 75% cheaper

A cache read is what you pay when the model re-reads context it has already processed. That rate has fallen by 75%, to $0.25 per million tokens, which implies a prior rate of $1.00. Because cache reads dominate long agentic sessions, the effect on a real bill is far larger than the line item suggests.

Typical versus highly agentic workloads

Anthropic indexes Fable 5 at 100 and measures four weeks of actual August 2026 usage at default effort. A typical workload — Claude Enterprise, Claude Code, and the API combined — lands at 75, a saving of roughly 25%. A context-heavy, tool-heavy workload, where cache reads are most of the cost, lands at 55, a saving of about 45%.

Indexed cost of the same workload, Fable 5 = 100
Fable 5, either workload 100
Fable 5.1, typical workload 75
Fable 5.1, highly agentic workload 55

Effort levels are the other lever

Fable 5.1 runs at five effort levels — low, medium, high, xhigh, and max — and Anthropic says that at low or medium effort it matches or beats Fable 5 at much lower cost. The defaults differ by product: high effort in Claude Code, medium in Claude Cowork and on Claude.ai. Anyone tuning spend should treat the effort setting as a first-class control rather than a hidden default.

Price lineFable 5Fable 5.1
Input, per million tokens$10$10
Output, per million tokens$50$50
Cache reads, per million tokens$1.00 (implied by the 75% cut)$0.25
Indexed cost, typical workload10075
Indexed cost, highly agentic workload10055

Safeguards: Fewer False Alarms, Same Risk Ceiling

The safeguard changes are the reason Mythos 5.1 exists as a separate product, and they cut in two directions at once — more permissive for benign work, unchanged at the top end.

Biology: 85% fewer refusals on benign questions

Anthropic’s newest biology safeguards fire 85% less often on benign requests about elementary biology and medical questions than the ones that shipped with Fable 5. Queries about life sciences research and development are still redirected to the Opus models. That is the gap the Life Sciences Verification Program is designed to close for vetted professionals.

Cybersecurity: vulnerability discovery is now allowed

Fable 5.1 may now be used to identify software vulnerabilities — defensive work — though not to develop exploits for them. Anthropic says Claude Code users should see roughly 60% fewer safeguard interventions per session. Penetration testing, exploit generation, and binary-based vulnerability scanning still get routed to Opus models, so a cybersecurity team’s dual-use work is only partly unblocked.

The ceiling did not move

On chemical and biological risk, Anthropic ran expert red-teaming, automated evaluations, and a tabletop exercise pairing PhD-level biologists with AI experts. Mythos 5.1 is more capable than Mythos 5 but still falls short of the next risk tier in the Responsible Scaling Policy, so it ships with the same restrictions. On cyber, it has the strongest capabilities of any model the company has released while remaining in the lower risk category of its Frontier Compliance Framework.

Alignment moved in the right direction, with gaps

An automated behavioural audit found Mythos 5.1 better aligned than Mythos 5 on most metrics: less likely to reach for resources outside its test environment on impossible tasks, less prone to motivated reasoning, less willing to ignore explicit constraints, and lower on both attempted and successful reward hacking. Anthropic also states the model can still sometimes bypass approvals and auto-mode classifiers, and that its audit gives limited visibility into very long-context and multi-agent work. That candour is the useful part; the rising commercial pressure to secure AI deployments is why HiddenLayer just raised $100M.

Enterprise Frontier Safeguards and Zero Data Retention

The second structural change in this release has nothing to do with capability. It is about where your data sits while the safeguards run.

How EFS works

Enterprise Frontier Safeguards keeps misuse detection in place while giving customers the privacy of a zero data retention agreement. Data is stored in cloud infrastructure the customer controls, not Anthropic’s, and any human review is done by the customer by default. It is the first serious attempt this year to make “we monitor for abuse” and “we retain nothing” hold at the same time.

Who built it and where it runs

Anthropic says EFS was developed with more than 100 customers across financial services, healthcare, manufacturing, telecom, law, retail, and the public sector, together with Amazon Web Services, Google Cloud, and Microsoft Azure. It will be supported on Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Claude Platform on AWS, Google’s Agent Platform, and Microsoft Foundry.

The interim arrangement

EFS rolls out in phases starting in autumn 2026. Until it is ready, eligible customers can use Fable 5.1 and Fable 5 with zero data retention. If a data residency clause is what has kept a frontier model out of your estate, this is the paragraph to take to your legal team.

How to Get Access to Mythos 5.1

Access to Mythos 5.1 runs through two vetting programmes, and both are narrower than the announcement’s confident tone suggests.

The Cyber Verification Program

The CVP already provides reduced cyber safeguards on certain Opus- and Sonnet-class models for defensive security work. Anthropic says Mythos-class models will be added “in the near future” — which means applying today does not yet get you Mythos 5.1 for cyberdefence. Applications go through Anthropic’s programme portal.

The Life Sciences Verification Program

The LSVP is the route to Mythos 5.1 with safeguards tuned for professional research and development, with every other safeguard left in place. It was built in partnership with the US government, the first participants are already enrolled, and Anthropic plans to widen it to the broader life sciences community.

The geography constraint

Mythos 5.1 is currently available only to a set of US organisations. Anthropic says it is coordinating with the US government to expand access to more domestic and international partners “as quickly as possible”, with no date attached. For a UK or EU research group, that is the binding constraint — not the science.

RouteWho it is forWhat it unlocksStatus
General availabilityEveryoneFable 5.1 with production safeguardsLive now
Cyber Verification ProgramVetted defensive security practitionersReduced cyber safeguards; Mythos-class models comingOpen, Mythos not yet included
Life Sciences Verification ProgramLife sciences professionalsMythos 5.1 with research-grade biology safeguardsFirst participants enrolled, US only
Claude SecurityTeams scanning their own codebasesMythos 5.1 indirectly, inside the productLive now

Anti-Distillation, Watermarks, and the EU AI Act

Two smaller changes in this release will outlast the benchmark charts, because both are about the API contract rather than the model.

The context-editing change

Distillation — extracting a strong model’s capabilities into a cheaper one, often through thousands of fake accounts — is treated here as a safety risk, because distilled capabilities can be released without safeguards. So new API accounts created from the announcement date onward can no longer manually edit Claude’s prior context while preserving the transcript of its prior thinking. Existing accounts are unaffected for now, but the change will apply to everyone with future releases, and a small number of custom integrations will break.

The watermark and the detection API

Having signed the EU AI Act’s Code of Practice on Transparency of AI-Generated Content in July 2026 alongside 190 other signatories, Anthropic now watermarks the output of models released after 2 August 2026. The watermark is a numerical signal, invisible without the detection API, carries no information about the user or their conversations, and Anthropic says it has no practical effect on output quality.

Who can detect it

The detection API is in private preview for regulators, law enforcement, media, fact-checkers, independent researchers, educational organisations, and EU civil society groups, plus enterprises with their own compliance obligations. Access is meant to widen over time. Compare this with OpenAI designating Astra as critical for cybersecurity: both labs are now shipping governance artefacts alongside weights.

What Mythos 5.1 Changes If You Buy or Build With AI

Most readers will never touch Mythos 5.1. The release still changes three practical things, and one of them is a budget line.

If you run an engineering organisation

The cache-read cut is the actionable item. If your agents re-read large contexts — long repositories, document sets, tool transcripts — your bill falls by roughly a quarter without any change on your side, and by nearly half on the heaviest workloads. Audit your effort settings at the same time, because low and medium now buy Fable 5-class results at a fraction of max-effort cost.

If you run a security team

Vulnerability discovery on the generally available model is the change that matters, not Mythos 5.1 itself. You can now point Fable 5.1 at your own code to find weaknesses; you still cannot use it to build exploits, and pen-testing work continues to route to Opus. Plan the CVP application as a 2027 capability, not a Q4 one.

If you run a research group

The protein-design and kernel-optimisation results are the strongest public evidence yet that frontier models can do real scientific labour rather than literature review. The catch is access: those results came from Mythos 5.1, which is US-only and programme-gated. Build the AI strategy assuming the public model, and treat Mythos-class access as upside.

The question to ask first

Before any of this, ask what your organisation’s data retention position actually requires. If the answer is zero retention, EFS changes which frontier models are procurable at all — and that is a bigger decision than 5 points on a coding benchmark.

What We Could Not Verify About Mythos 5.1

Some of the most useful numbers are absent from the announcement, and it is worth being explicit about which.

The context window

Anthropic publishes no context window for either model in this announcement. Given how much of the cost argument rests on cache reads and long agentic sessions, that omission is notable, and we have not assumed a figure.

Independent confirmation

Every benchmark here is Anthropic’s own harness, and the company itself shows its numbers diverging from the public Terminal-Bench-Science leaderboard by a few points in both directions. The customer quotes are from early-access partners. None of it is independently reproduced yet, and the System Card is the only place the methodology is documented in full.

The access timelines

“In the near future” for Mythos-class models in the CVP, “as quickly as possible” for non-US access, and “in phases, starting this fall” for EFS are the only schedules given. Anyone planning around Mythos 5.1 should treat all three as unscheduled.

Mythos 5.1: Common Questions

What is Claude Mythos 5.1?

Mythos 5.1 is the same underlying model as Claude Fable 5.1, released with reduced safeguards for vetted cybersecurity and life sciences work. It is not a bigger or smarter model — it is the same weights with a different safety configuration and a much narrower audience.

Can I use Mythos 5.1 today?

Only through a trusted access programme, and currently only if you are part of a set of US organisations. The Life Sciences Verification Program has enrolled its first participants, and the Cyber Verification Program does not yet include Mythos-class models. Claude Security uses Mythos 5.1 internally, so its scanning output is the one indirect route.

How much does Fable 5.1 cost?

Input is $10 per million tokens and output is $50 per million, unchanged from Fable 5. Cache reads dropped 75% to $0.25 per million, which cuts a typical bill by around 25% and a highly agentic one by up to about 45%.

Is Mythos 5.1 more dangerous than Fable 5.1?

It is the same model with fewer refusals, so it will answer things Fable 5.1 will not. Anthropic’s own evaluations put it below the next risk tier of its Responsible Scaling Policy on biology, and in the lower risk category of its Frontier Compliance Framework on cyber, which is why the vetting requirement carries the weight rather than the weights.

What is the difference between Fable 5.1 and Opus 5?

On Anthropic’s published benchmarks Fable 5.1 leads Opus 5 on all nine: 52.6% against 29.0% on agentic science, 55.8% against 52.3% on Terminal-Bench 4.0, and 1853 against 1824 on GDPval-AA v2. Several early-access partners also report Fable 5.1 as faster and cheaper per task than Opus 5.

What is Enterprise Frontier Safeguards?

EFS is Anthropic’s new way of running misuse detection on infrastructure the customer controls, so an enterprise can have zero data retention and monitored safeguards at once. It rolls out in phases from autumn 2026, and eligible customers get interim zero data retention until then.

Does the new watermark affect output quality?

Anthropic says it does not. The watermark is a numerical signal added for EU AI Act compliance, invisible without the detection API, and it contains no information about the user, the organisation, or the conversation.

What breaks with the anti-distillation change?

API accounts created from the announcement date can no longer edit Claude’s prior context while keeping the transcript of its prior thinking. Existing accounts are unaffected for now. A small number of custom multi-turn integrations will need adjusting, and Anthropic’s Help Center documents the change.

References