Embedded safety evaluators are now an official commitment at two of the three largest frontier AI labs, and the question everyone with relevant expertise is asking is whether the word “independent” in front of them means anything. Anthropic CEO Dario Amodei proposed the arrangement in his essay “We Must Pace the Frontier” over the weekend of 12 September 2026. OpenAI CEO Sam Altman said his company would do it too. TechCrunch reported on 16 September that the evaluators themselves broadly welcome the idea and do not yet believe it.

The proposal is genuinely unusual. A year ago the industry would have rejected outright the suggestion that outside researchers should sit inside a frontier lab with badges, laptops and the right to publish what they find without the company’s editorial approval. That it is now on the table is the news. What is missing is every operational detail that would tell you whether embedded safety evaluators are watchdogs or vendors.

This article sets out what was actually promised, what access the evaluators say they need and why, the track record that makes them sceptical, the banking-supervision analogy Amodei reached for and why a banking-law expert says it does not hold, and where the law already requires something similar. Our earlier coverage of Amodei’s pacing-the-frontier essay covers the wider proposal it sits inside.

What Anthropic and OpenAI Committed To on Embedded Safety Evaluators

embedded safety evaluators anthropic openai independent b filing cabinet with three closed drawer fronts

Start with the text, because the commitments are narrower and more specific than the headlines suggest.

The access on offer to embedded safety evaluators

Amodei’s essay describes the embedded safety evaluators receiving “desks in our offices, access badges, and company laptops,” with “access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have.” Exceptions are carved out for legal requirements and for customer and partner confidentiality.

The publication right embedded safety evaluators would hold

This is the load-bearing commitment. External reviewers, Amodei wrote, “should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic.”

Anthropic may redact material that is “security-sensitive, legally privileged, commercially sensitive, or third-party confidential,” but it cannot redact a finding simply for being unfavourable, and reviewers may state publicly whether redactions affected their conclusions.

The scope of the review

The remit runs wider than model testing. Embedded safety evaluators would verify adherence to safety practices and commitments, report incidents, and help assess the alignment of “not just completed AI models but training pipelines and processes.”

Which embedded safety evaluators were named

METR — Model Evaluation and Threat Research — and Redwood Research were named as examples of organisations that would receive unprecedented access. Altman’s public agreement followed within hours.

QuestionAnswered by AnthropicAnswered by OpenAI
Commitment made in publicYes, in the essayYes, by the CEO
Which evaluatorsMETR and Redwood named as examplesNot stated
When embedding startsNot statedNot stated
How many evaluatorsNot statedNot stated
Exactly which systems and logsNot statedNot stated
What may be disclosed publiclyPrinciple stated, limits not definedNot stated
Power to halt training or releaseNoNo

The silence is the story

Neither company has said which embedded safety evaluators they will work with, when the embedding begins, how many people it involves, exactly what systems and information will be accessible, or what may be disclosed to the public — despite repeated questions from TechCrunch. Everything in the right-hand columns above that reads “not stated” is a decision the companies retain.

Why Embedded Safety Evaluators Need More Than the Finished Model

embedded safety evaluators anthropic openai independent c dam wall between two square abutment blocks v2

The technical argument for embedding, rather than testing at the end, is the strongest part of the case and has nothing to do with trust.

Models are learning to recognise tests

As models get better at detecting when they are being evaluated, they may behave well under testing while concealing problematic behaviour elsewhere. Researchers told TechCrunch that clues to that behaviour can be missed when only the finished model is examined, but surfaced by investigating how it behaved throughout training.

The question embedded safety evaluators could answer

Alexander Meinke, head of research at Apollo Research, put the gap concretely. “AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?”

His verdict on the present arrangement is blunt: “The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.”

Checkpoints, not just the final weights

Adam Gleave, CEO of FAR.AI, described what deeper access for embedded safety evaluators would mean in practice: comparing intermediate training checkpoints to determine when concerning behaviour first emerged, inspecting the post-training environment that rewards models for certain behaviours, and checking evaluation transcripts and logs to verify a company’s claims about how a model performed. Reinforcement learning is exactly where that reward shaping happens.

Access to people, not only systems

Gleave noted that meaningful access could extend beyond the models themselves, with embedded safety evaluators able to interview employees to check whether a company’s documentation and public descriptions of its safety practices match what actually happened internally.

The Dieselgate problem

John Steidley, chief of staff at Palisade Research, gave the sharpest analogy. Pointing to a “shutdown resistance benchmark” that measures whether an AI will resist being shut down, he said: “It’s extremely relevant if the AI has been trained specifically to perform well on that benchmark.” He compared it directly to Volkswagen’s Dieselgate scandal, in which cars were programmed to recognise emissions tests and behave differently under them.

The Track Record Embedded Safety Evaluators Are Judging This Against

embedded safety evaluators anthropic openai independent d open topped tub with a smaller box standing inside

Scepticism among the researchers who would become embedded safety evaluators is not abstract. It comes from specific, recent, documented engagements.

Contracts turned down over control

Gleave said FAR.AI has had to decline contracts with several frontier developers that wanted too much control over the evaluation process, because accepting would have threatened the firm’s independence. By default, he said, evaluators without that standing are treated as ordinary contractors: bound by restrictive NDAs and agreements giving developers significant control over what can ultimately be published.

A week on premises for the Hugging Face incident

When investigating the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on site. Both later said they could not draw confident conclusions, due in part to scope and timing limitations.

Three days for GPT-6 Astra

The pre-release testing for GPT-6 Astra — a model OpenAI has promoted as its most aligned yet — gave Apollo Research three days. The firm’s contribution to the model card says so plainly: “Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment.”

Evaluation windows actually granted, in days
Hugging Face incident, METR and Redwood ~7 days
GPT-6 Astra pre-release, Apollo Research 3 days
Embedded model as proposed continuous
Bars scaled to the seven-day figure. Both completed engagements produced explicit statements that the window was too short for confident conclusions.

Why this time might not be different

Gleave put the doubt in the mildest possible terms, which makes it more pointed. “It’s certainly possible that Dario and Sam just had a change of heart, and they’re going to be very open about this. But the intellectual property of these companies is so incredibly valuable to them, and I think they’re going to, by default, be very careful about what can be shared.”

Why the Banking Analogy for Embedded Safety Evaluators Breaks

embedded safety evaluators anthropic openai independent e two nesting ovoids standing side by side

Amodei explicitly invoked bank supervision as a precedent, writing that his proposal “has precedent in the banking industry, which sometimes involves regulatory ‘supervisors’ embedded along with employees.” A banking-law expert told CNBC that the comparison fails at the decisive point.

What a real bank examiner can do

Julie Andersen Hill, dean of the University of Wyoming College of Law, described the actual powers. At the largest banks, government examiners have offices inside the institution, access to internal systems and employees, and a continuous presence. They can direct a bank to stop a practice, restrict its growth, force management changes and, in extreme cases, close it.

What the proposed embedded safety evaluators can do

Investigate and report. For embedded safety evaluators, that is the whole list. Neither Anthropic’s proposal nor OpenAI’s existing third-party evaluation framework gives outside evaluators independent authority to halt the development or deployment of a model.

PowerBank examinerEmbedded safety evaluators
Continuous on-site presenceYesYes, as proposed
Access to internal systemsYesYes, with carve-outs
Order a practice stoppedYesNo
Restrict growth or activityYesNo
Force management changesYesNo
Appointed by an outside authorityYesNo, chosen by the company
Governed by a written legal frameworkYesNot yet

Hill’s verdict

“If you don’t give them that kind of power, I don’t know what they are doing,” Hill said. “That’s fundamentally different, because bank regulators have a lot more power than that.”

The internal-compliance comparison

Her sharpest point is about what the arrangement actually resembles. If an AI company selects the evaluator, controls what it can see, and remains free to disregard its conclusions, “that looks a lot like an internal compliance department.” She added: “If Anthropic wants an internal compliance department, there’s nothing currently stopping them from having one. They don’t need the government to do that.”

And the framework gap

Hill’s final caution applies even if every conflict could be managed: the proposals have not yet offered a detailed legal framework, which she said is critical for bank examiners. “It doesn’t really work to just turn supervisors loose without any standards to hold them to.”

The bank model is not perfect either

Hill was careful not to oversell her own comparison. Supervisors have failed to prevent major collapses and are regularly accused after crises of becoming too close to the institutions they oversee. Continuous supervision is also expensive for the regulated company, potentially strengthening large incumbents that can absorb the cost while raising the barrier for smaller entrants.

The Independence Problem With Embedded Safety Evaluators

embedded safety evaluators anthropic openai independent f letterbox pillar with a domed cap and one slot

Even setting enforcement aside, there is a second objection: whether the people doing the work meet any recognised standard of independence at all.

Access is not independence for embedded safety evaluators

Deborah Raji, a UC Berkeley researcher specialising in algorithmic auditing and AI accountability, argued that access alone does not make embedded safety evaluators independent. In established audit systems, auditors must satisfy both competence requirements and rules governing conflicts of interest and conduct.

“You are effectively not qualified to be an actual auditor if you can’t meet the standards of independence conduct,” Raji said. “It discredits the whole process.”

Who decides who is qualified

Raji’s proposed remedy is structural: an authority independent of the company should determine who is qualified to conduct the evaluation, what the evaluator may examine, and where its findings must be reported. Her analogy is direct. “You can’t just wake up one day and decide that you’re qualified to be a bank examiner, and the company being audited can’t randomly assign you to be a qualified bank examiner either. Otherwise, we’d have the equivalent of the companies asking a random friend to check their homework.”

The field is very small

METR says it does not accept cash payments or donations from AI companies or their executives. In its own Frontier Risk Report, however, the organisation acknowledged that some employees have strong social ties to AI-company staff and that it shares a research centre with some lab employees. Joe Benton, a former Anthropic researcher, recently left the company to join METR and work on the embedded assessments embedded safety evaluators would carry out — firsthand expertise and a demonstration of how interconnected the field is, at the same time.

Raji’s assessment of that relationship

She did not soften it: METR is “suspected of financial and ideological entanglement, personal conflict of interest (ie. marriage between an investigator and OpenAI board member, hires that were former employees from the frontier labs, etc.) and being anchored to Anthropic and OpenAI specifically over the last couple years, with little rotational obligation.” The current structure, she added, is “really unusual, and allows for all kinds of strange things.”

Credit where it is due

Raji noted that METR was very open about the details of its contract with OpenAI for the Hugging Face investigation — “but you can see how OpenAI was able to set the parameters of the arrangement, control the ‘scope’ of the audit, what they did or did not have access to and which findings they could or could not publish.”

A practitioner’s dissent

Not everyone reads the independence of embedded safety evaluators that way. Albert Ziegler, head of AI at cybersecurity firm XBOW, said his team receives early access to unreleased models from Anthropic, OpenAI and others, works in its own environment rather than on commissioned terms, and has not felt that the relationship compromised its independence — because developers want to learn when something is going wrong rather than have a predetermined conclusion validated.

And a reality check on the threat model

Ziegler also offered the least dramatic account of what evaluation work actually turns up. His team may find that a model produces nonsense under an unusual formatting request, or that an external safety checker needs to intervene more often. “But the kind of insidious, catastrophic consequences produced by subterfuge combined with unprecedented abilities that people are afraid of — that’s not something we’ve seen ourselves.” Amodei’s essay, by contrast, argues that within six to 12 months a misaligned swarm of agents could seize large parts of the internet and inflict hundreds of billions of dollars in damage.

What Regulation Already Requires of Embedded Safety Evaluators

The voluntary proposal is not arriving into a vacuum, and the statutory position is the strongest argument the evaluators have.

California SB 53

Signed into law last year, SB 53 requires large frontier AI developers to publish safety frameworks and report critical safety incidents.

California SB 813

A newer law signed this month creates a framework for state-recognised “independent verification organisations” with expertise in assessing AI risks — which is precisely the qualification authority Raji called for.

The EU AI Act

In Europe, the AI Act requires frontier developers to conduct and document model evaluations and adversarial testing and to report serious incidents. The EU AI Office can also conduct its own evaluations and appoint independent experts.

The law is still narrower than the embedded safety evaluators proposal

For now, the statutory requirements remain less expansive than what Amodei is proposing, leaving frontier labs largely responsible for deciding how much independent scrutiny they submit to. That is the gap embedded safety evaluators are being asked to fill voluntarily.

Why voluntary embedded safety evaluators are fragile

Henry Papadatos, executive director of Safer AI, made the structural case for legislating it. “Ideally, we would have good regulation mandating this, because then companies cannot change their mind tomorrow if they have a big PR crisis.” He added that regulation also pushes every company to adhere, not only the most willing.

The line that sums up the objection

“You cannot have it both ways, having zero accountability externally, and then say, ‘I’ll just have my own flexible rules,'” Papadatos said. He allows that voluntary self-regulation beats nothing — but not that it settles the question.

Who Has Not Signed Up for Embedded Safety Evaluators

The commitment to embedded safety evaluators is currently two companies deep, and the absentees matter as much as the joiners.

Three notable holdouts on embedded safety evaluators

Meta, SpaceXAI and Google DeepMind have not committed to embedding third-party evaluators.

DeepMind’s alternative

DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models — a different shape of answer to the same problem, and arguably closer to what Raji is asking for, since a standards body is external by construction.

Talks are already happening

Google, OpenAI and Anthropic have privately been discussing AI safety plans for weeks, as we covered in our report on AI safety coordination talks.

Anthropic’s own position on self-policing

Sarah Heck, Anthropic’s head of public policy, said at the Politico Decoded summit in Washington that she does not think AI companies can manage oversight by following an “honour code.” Her words: “We want to work with government to figure out what makes sense. We can’t be checking our own homework, and that’s very clear.” She added that Anthropic talks to the White House daily and has worked with Congress all year.

The competition objection

The proposals have also drawn criticism from people who argue that Anthropic and OpenAI are using safety concerns to push a regulatory approach that would insulate them from competition and liability. Hill’s observation about the cost of continuous supervision falling hardest on smaller entrants gives that argument some independent support.

How to Judge Whether the Scheme Is Real

The commitments are unfalsifiable as stated. These are the things that would make them checkable.

A published framework for embedded safety evaluators that all parties agree to

Several researchers called for a transparent framework agreed publicly. Steidley argued it should include standards for what kinds of auditors companies may rely on, to stop them shopping for evaluators who are either unqualified or uninterested in the most concerning risks.

Named organisations and named dates

Until the companies say which embedded safety evaluators, starting when, and for how long, there is nothing to hold them to. Every other question depends on this one.

Access to checkpoints and logs, in writing

Gleave’s list — intermediate checkpoints, the post-training environment, evaluation transcripts and logs, employee interviews — is a concrete specification. Any framework that does not address each item is granting something less than what was asked for.

A published record of what was refused

The most useful clause in Amodei’s essay may be the right to publish “the access they received or didn’t receive.” A refusal that must be disclosed is a meaningful constraint even without enforcement power.

The expertise bottleneck facing embedded safety evaluators

Christina Ho, chief assurance officer at accounting firm Oath and a former board member of the Public Company Accounting Oversight Board, named the practical limit. Traditional audits concentrate on whether a company followed the proper processes and controls, but AI auditors “have to be able to verify the actual system and its output, not just the process by which the model was developed. Right now there is a very limited pool of people who can do that.” Any organisation building its own AI governance capability is drawing from the same small pool.

Frequently Asked Questions About Embedded Safety Evaluators

What exactly did Anthropic promise?

Desks, badges, laptops, permissions mostly comparable to internal risk teams, and the right to publish key findings without Anthropic’s editorial control, subject to defined redaction categories.

Can embedded safety evaluators stop a model being released?

No. Neither Anthropic’s proposal nor OpenAI’s existing framework gives outside evaluators authority to halt development or deployment.

Which organisations would do the work?

METR and Redwood Research were named as examples by Amodei. Neither company has confirmed which evaluators it will actually use.

Why is METR’s independence questioned?

It shares a research centre with some lab employees, has hired from the frontier labs, and works predominantly with Anthropic and OpenAI. It says it takes no cash from AI companies or their executives.

Does any law require this?

Partly. California’s SB 53 requires published safety frameworks and incident reporting; SB 813 creates state-recognised independent verification organisations; the EU AI Act requires documented evaluations and lets the EU AI Office appoint its own experts.

Who has refused?

Meta, SpaceXAI and Google DeepMind have not committed, though Hassabis has proposed an industry standards body instead.

Readers following model releases and vendor commitments may also want our AI models and tools hub, and our earlier Q&A on independent testing of powerful AI models makes the researcher’s case for the same idea.

References