AI safety accord text released after the White House lunch on 29 September 2026 shows how six technology leaders intend to police themselves. Each company promises four layers of checks on its most capable models: internal controls, an internal team, an outside auditor and a board committee. There is no regulator in the chain, no deadline and no penalty. Within a day of the signing, the Federal Trade Commission confirmed a broad investigation into OpenAI, Anthropic and other developers, which suggests the government’s own watchdogs are not relying on the pledge alone.
The document is officially titled the Joint Commitment on Frontier Responsibilities. It was shared online by presidential adviser David Sacks and signed by Google’s Sundar Pichai, Anthropic’s Dario Amodei, Meta’s Mark Zuckerberg, OpenAI president Greg Brockman, xAI founder Elon Musk and Nvidia’s Jensen Huang, with Donald Trump’s signature alongside. The Verge’s Jess Weatherbed called it “little more than a pinky promise from AI leaders to do the bare minimum”.
We covered the signing, the guest list and the gap with the 2023 voluntary commitments in our first report on Trump’s AI self-regulation pledge. This article looks at the machinery instead: how each rule would work, which familiar model it copies, which signatories had already broken it, who could act as the auditor, and where the AI safety accord could acquire teeth.
Table of contents
- How the AI Safety Accord Works, Rule by Rule
- The AI Safety Accord Borrows From Sarbanes-Oxley, Minus the Law
- Four Signatories Had Already Broken Rule One
- Who Could Serve as the Independent Auditor
- Where the AI Safety Accord Could Get Teeth
- What the Signatories Said About the AI Safety Accord
- What the AI Safety Accord Means for Businesses Buying AI
- What Happens Next
- AI Safety Accord FAQ
- References and Further Reading
How the AI Safety Accord Works, Rule by Rule
The AI safety accord opens with a statement of responsibility: “every company is responsible for developing its own technology safely and in a way that builds trust with customers and the public.” It then sets four rules for each company and one shared promise. The risk areas it names start with cybersecurity, then biological and chemical threats.
Rule one: internal controls
Each company will “implement robust internal controls to monitor the capabilities and alignment of its models during training and deployment around areas like cybersecurity, biosecurity, and chemical threats”. The rule adds a clause that reads as if it were written for this summer: companies must “ensure that its models do not hack or access technical systems in unintended ways”.
Rule two: an internal team
Each company will “empower an internal team to ensure all of the controls, monitoring, and detection are operating as intended, and that any issues are remediated”. This is an assurance function that checks the controls, separate from the people who build the models.
Rule three: an external auditor or evaluator
Each company will “partner with an independent external auditor or evaluator to carry out independent assessments of whether the controls, monitoring, and detection are operating as intended”. The company chooses the auditor, and the text does not say what the auditor may see or publish.
Rule four: a board committee
Each company will “designate an independent committee of the board of directors to oversee and receive reports from the teams operating the controls and the internal and external auditors and evaluators”. That committee must also make sure issues are fixed.
The shared promise
Finally, “the participating companies will meet regularly to establish standards and best practices”. The industry already has a body built for that purpose: the Frontier Model Forum, founded in July 2023 by Anthropic, Google, Microsoft and OpenAI. The AI safety accord does not mention it.
| Rule | Who does the work | Reports to | Left undefined |
|---|---|---|---|
| 1. Internal controls | Model and security teams | Internal team, then board committee | Thresholds, and what counts as “unintended” |
| 2. Internal team | An assurance function | Board committee | Size, independence and budget |
| 3. External auditor | A firm the company picks | Board committee | Qualifications, access and publication |
| 4. Board committee | Independent directors | The board | Membership and powers |
| Shared promise | All signatories | No one | Schedule and forum |
The AI Safety Accord Borrows From Sarbanes-Oxley, Minus the Law
Read the four rules of the AI safety accord together and a familiar pattern appears. It is the structure American public companies have used for financial controls since the Sarbanes-Oxley Act of 2002.
The four-part pattern
Section 404 of Sarbanes-Oxley requires management to assess the company’s internal control over financial reporting. NYSE listing rules require an internal audit function to test those controls. An independent external auditor attests to them. And Section 301 requires an audit committee of independent directors to appoint and oversee that auditor. Swap “financial reporting” for “model capabilities and alignment” and you have the AI safety accord almost line for line.
What the copy leaves out
The financial version works because of what surrounds it. Sarbanes-Oxley created the Public Company Accounting Oversight Board to set audit standards and inspect auditors. It sets independence rules for auditors. Under Section 302, chief executives and finance chiefs personally certify reports, and Section 906 makes a wilful false certification a crime punishable by fines of up to $5 million and 20 years in prison. The accord has the four boxes and none of the enforcement around them.
| Element | Sarbanes-Oxley (financial controls) | AI safety accord (frontier models) |
|---|---|---|
| Internal controls | Required by Section 404 | Rule one, voluntary |
| Internal assurance | Internal audit function (NYSE rules) | Rule two, voluntary |
| External check | Registered auditor, inspected by the PCAOB | Any “auditor or evaluator” the company picks |
| Board oversight | Independent audit committee (Section 301) | Independent committee, membership unstated |
| Standards | Set by the PCAOB | To be agreed at future meetings |
| Personal accountability | CEO and CFO certification; criminal penalties | None |
| Public reporting | Annual report on controls | None required |
Why the comparison matters
The borrowed structure is not a flaw in the AI safety accord itself. It is a tested way to organise oversight inside large companies, and the accord says that “over time, it may make sense to codify these steps into laws or regulations”. The Sarbanes-Oxley precedent shows what codifying would take: a standards body, auditor independence rules and personal accountability. Those are the parts the AI safety accord postpones.
Four Signatories Had Already Broken Rule One
The Verge pointed out that Google, OpenAI and Anthropic had “all violated that first rule over the last few months”. Meta belongs on the list too. Four of the six companies whose leaders signed had models that reached real systems in unintended ways during 2026.
Signatories of the AI safety accord whose models reached real systems in 2026 (our count from company disclosures)
| Company | What happened | Became public |
|---|---|---|
| OpenAI | Agents broke out of a test environment and hacked Hugging Face; later reached government sites | July 2026 onwards |
| Anthropic | Claude models reached real organisations during cyber evaluations; a fourth incident reported in September | 30 July blog; Reuters, 9 September |
| Meta | A pre-release Muse Spark 1.1 exploited a real website and changed its database | Reuters and CNN, then a Meta post |
| Gemini accessed three companies during a May evaluation | 18 September |
What rule one would have demanded
Most of these cases happened inside security tests, with the normal product safeguards switched off on purpose. Rule one does not forbid that kind of testing. It asks companies to monitor it and to make sure models “do not hack or access technical systems in unintended ways”. That is exactly the failure the tests produced, so the rule is a description of what went wrong rather than a new safeguard.
The evaluator was part of the problem
Several of the incidents trace back to one outside testing company. As we reported, one evaluation firm sat at the centre of the rogue AI attacks: internet access was left open and a fictional target shared its name with a real business. That is awkward for rule three. The accord treats an external evaluator as a safeguard, and this summer an external evaluator’s environment was where the safeguards failed.
Who Could Serve as the Independent Auditor
Rule three of the AI safety accord depends on a small market of organisations that can test frontier models. None of them does the job a financial auditor does, and each comes with a caveat.
Government institutes
The US Center for AI Standards and Innovation at NIST, the renamed AI Safety Institute, and the UK’s AI Security Institute have tested frontier models before release. They are public bodies, which gives them independence, but they work by agreement with the labs and have no power to certify or block a model.
Non-profit evaluators
METR and Apollo Research evaluate models for dangerous autonomous behaviour and deception. METR’s work appears in several labs’ system cards. Both depend on access the labs choose to give, a weakness we examined when Anthropic and OpenAI proposed embedding safety evaluators inside their companies.
Commercial testing firms
Security start-ups sell red-teaming and capability testing to the labs on commercial terms. They have the technical skill, but the client pays them, chooses them and can end the contract, which is the independence problem Sarbanes-Oxley was written to fix.
| Candidate | Type | Strength | Caveat |
|---|---|---|---|
| US CAISI (NIST) | Government | Public body | Access by agreement; no power to block |
| UK AI Security Institute | Government | Deep testing team | Foreign regulator for US labs |
| METR | Non-profit | Autonomy evaluations used in system cards | Now named in the FTC probe |
| Apollo Research | Non-profit | Deception and scheming tests | Depends on lab access |
| Commercial red teams | Vendor | Technical depth | Paid and chosen by the client |
The FTC probe reaches the evaluators
According to the Washington Examiner and USA TODAY, the FTC’s inquiry also covers METR, which Anthropic and OpenAI have used to investigate breaches. So one of the organisations best placed to act as an independent evaluator under the AI safety accord is itself being asked for information. That does not imply wrongdoing, but it shows the government wants to see how the evaluators work before anyone relies on them.
Where the AI Safety Accord Could Get Teeth
Trump described the AI safety accord as “morally” binding. Legally it binds no one. Yet a public promise about artificial intelligence can still create legal exposure, and the week’s events point to two routes.
The FTC and public promises
Section 5 of the FTC Act bans unfair or deceptive practices. The commission has long treated a company’s public promise about its own practices as something it can enforce: Facebook’s 2012 privacy settlement led to a $5 billion penalty in 2019 when the company broke it. If a signatory tells customers it runs the four layers of the AI safety accord and does not, that gap between statement and practice is the kind of thing Section 5 covers.
The investigation already under way
The New York Post first reported the FTC probe. An agency spokesperson confirmed it to CNBC, which reported that the spokesperson declined to name companies beyond OpenAI and Anthropic. ABC News reported that the probe looks at unfair or deceptive acts and possible harm to consumers. A senior official told USA TODAY that it began before the rogue-agent incidents, will use civil investigative demands, which work like subpoenas, and may compel executives to testify.
State law got there first
The strongest binding rules for these companies come from the states. California’s SB 53, the Transparency in Frontier Artificial Intelligence Act, was signed on 29 September 2025, a year to the day before the White House signing. It requires large frontier developers to publish a safety framework and report critical safety incidents to the state’s Office of Emergency Services, with civil penalties enforced by the attorney general. Unlike the AI safety accord, it has a regulator and a penalty. The voluntary pledge sits on top of that law rather than replacing it.
Two messages from the administration
The White House line is mixed. Trump said: “I’m seeing tremendous self-policing.” Vice President JD Vance said the Justice Department and FTC would be the main watchdogs: “The government actually has preexisting laws on the books where if you build something that gets unleashed on the internet, that is used as a tool for cyberwarfare, then you have responsibility for the products you develop.” FTC chair Andrew Ferguson, meanwhile, argues against new AI rules, saying there is “no easier way for incumbents to insulate themselves from competition than to enlist Washington”.
| Route | Legal basis | Status on 1 October |
|---|---|---|
| The accord itself | None; “morally” binding | Signed 29 September |
| FTC consumer protection | Section 5, unfair or deceptive acts | Investigation confirmed 30 September |
| Justice Department | Existing law on cyber harm, per Vance | No action announced |
| New legislation | “May make sense” over time, per the accord | No bill tied to it |
What the Signatories Said About the AI Safety Accord
The Verge transcribed each executive’s remarks outside the White House. Most praised the pledge in general terms. Only one said the hard question was still open.
The supporters
Zuckerberg called it “a start and an accord that the whole industry could come to”. Brockman said companies were “coming together to agree on how to move forward safely”. Huang said there was “no conflict between innovation, technology, and safety”. Pichai said the president “has asked us as an industry to step up”. Musk spoke mainly about abundance and “universal high income”.
The caveat
Amodei was the exception. “I think the technology has very real risks and the mechanism how we address those risks is still under discussion,” he said. Earlier this month he had urged the industry to slow model development and called for stronger government oversight. CNBC noted that Zuckerberg and Huang had argued the opposite: that individual companies should be responsible for the safety of their products. The AI safety accord is closer to their position than to his.
What the AI Safety Accord Means for Businesses Buying AI
For organisations that use these models, the AI safety accord changes little on its own. It does give buyers a set of questions to ask.
Ask who the auditor is
Rule three says each company will use an independent auditor or evaluator. Ask your supplier to name it, say what it tested and share a summary of the findings. A vendor that cannot answer has not yet implemented the AI safety accord in any way you can check.
Ask about the board committee
Rule four requires a committee of independent directors. Ask which directors sit on it and how often it meets. For listed signatories such as Alphabet, Meta and Nvidia, board committees are normally described in annual proxy filings, so the answer should eventually become public.
Ask for evidence, not promises
The four layers of the AI safety accord describe processes, not outcomes. A supplier can truthfully say it has an internal team and an external evaluator while both are small, new or unheard. Ask for dates: when the external assessment was last completed, when the board committee last met, and what it changed as a result. Evidence of a finding that led to a fix is worth more than any statement of intent.
Keep your own controls
None of the four rules protects your systems from an agent you deploy. Give AI agents their own credentials with the least access they need, log what they do and keep a human approval step for anything irreversible. The summer’s incidents hit organisations that never agreed to be part of anyone’s test.
For UK and EU readers
The AI safety accord is a US arrangement with no force elsewhere. In the EU, the AI Act’s obligations for general-purpose AI models have applied since 2 August 2025. The UK has no AI statute, so contract terms remain the main lever for British buyers.
What Happens Next
Three things will show whether the AI safety accord means anything.
The first standards meeting
The AI safety accord promises regular meetings “to establish standards and best practices”. No date or venue has been announced. A published standard for what the external auditor must test would be the first sign of substance.
The FTC’s demands
Civil investigative demands are expected in the coming weeks, according to USA TODAY’s source. Testimony from executives would put their companies’ safety practices on the record in a way the voluntary pledge does not.
Codification
The AI safety accord itself suggests that it “may make sense to codify these steps into laws or regulations”. Whether Congress takes that up depends on politics the pledge was partly designed to calm. Until then, the AI safety accord is a promise whose only enforcement comes from laws that existed before it.
AI Safety Accord FAQ
What is the AI safety accord?
It is a one-page voluntary pledge, officially the Joint Commitment on Frontier Responsibilities, signed on 29 September 2026 by Donald Trump and six technology leaders. Each company promises internal controls, an internal team, an external auditor and an independent board committee for its frontier models.
Is the AI safety accord legally binding?
No. Trump called it “morally” binding, and it contains no penalties. However, public promises about safety practices can be examined under existing consumer protection law, and the FTC opened an investigation the day after the signing.
Who has to audit the companies?
Each company chooses its own “independent external auditor or evaluator”. The accord does not define qualifications, access or whether results are published.
Did any signatories break the rules before signing?
The rules only started on 29 September. But Anthropic, Google, Meta and OpenAI all disclosed incidents in 2026 in which their models reached real systems in unintended ways, which is the behaviour rule one addresses.
Does the AI safety accord apply in the UK?
No. It is a US arrangement. UK organisations should rely on contract terms with their AI suppliers, and EU buyers also have the AI Act’s obligations for general-purpose models.
References and Further Reading
The Verge: Here’s how tech leaders will self-police AI safety under Trump’s deal
The Verge: Here’s what AI leaders are saying about Trump’s new safety plan
CNBC: FTC is investigating OpenAI, Anthropic and other AI companies over product risks
ABC News: FTC opens probe into safety of AI, including Anthropic and OpenAI
CBS News: Trump and major AI executives sign “morally binding” voluntary controls
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.