Auto-review, the OpenAI feature that lets a second AI agent approve or block a coding agent’s risky actions, is now free for everyone who signs in to Codex with a ChatGPT account. Thibault Sottiaux, who leads Codex and ChatGPT work at OpenAI, announced the change on X early on Tuesday 6 October. His post said the feature “is now free and does not draw usage from your plan”.

That last clause is the news. The feature already existed and was already available on every paid plan. What changed is that the reviewer agent’s work no longer counts against a user’s usage limits, so there is no longer a cost to turning on the safer setting.

This article explains what the feature does, what OpenAI’s own evaluation numbers show, what it will not protect against, how businesses can configure it, and why making it free matters more than it first appears.

What OpenAI Changed With Auto-review on 6 October

Auto-review - auto review free chatgpt codex all users b ship in a bottle on a cradle stand

Sottiaux’s post went up at 07:13 UTC and had passed 840,000 views within hours. It was labelled “Day 2.1”, part of a programme he set out two days earlier.

The announcement

The post reads: “We have made Auto-review free for all users signed in through a ChatGPT account.” It describes the default sandbox setting, which asks the user to approve everything, as “prone to decision fatigue unless you spend a lot of time configuring specific rules”. Auto-review, it says, lets you “run long tasks while having a second agent review all actions taken by the primary agent”.

Its stated goal is narrow: “to prevent high-risk actions from being taken and to protect against unwanted actions that are not aligned with the original user intent.”

Where to switch Auto-review on

A follow-up post gave the path: Settings, then General, then Permissions, then the reviewer option. Another follow-up made the pitch even more directly: “Use auto-review instead of full-access. Not a trade-off anymore and you’ll sleep more peacefully.”

Part of a 28-day push

On 4 October Sottiaux wrote that “over the next 28 days, each day we’ll either ship one thing that is a clear improvement and relevant for most codex/work users or ship a full reset”. Day 1, on 5 October, made the default speed about 50% faster for GPT-6 Astra and GPT-6.1 Sol, moving from 30 to 50 tokens per second. Free Auto-review was the second release in that run.

How Auto-review Works

auto review free chatgpt codex all users c railway buffer stop at the end of a track

The reviewer sits at the edge of the sandbox that Codex runs in. OpenAI’s documentation describes it as replacing “manual approval at the sandbox boundary with a separate reviewer agent”.

The sandbox boundary

By default, Codex can read files, edit files inside a permitted folder and run local commands inside that boundary. Anything beyond it, such as a network request or an edit outside the project, needs approval. Without the reviewer, that approval comes from you, one prompt at a time.

With Auto-review enabled, the main agent works exactly as before, under the same limits. When it asks to cross the boundary, the request goes to a separate reviewer agent instead of to you. The reviewer decides whether the action should run and returns a reason.

What the reviewer sees

The reviewer is itself a Codex agent with a narrower job. According to the docs, it sees “a compact transcript plus the exact approval request”, including user messages, visible assistant updates, relevant tool calls and their outputs. It can run read-only checks to gather context, though rarely.

It does not see the main agent’s hidden reasoning. The documentation is clear: the reviewer “sees retained chat items and tool evidence, not private chain-of-thought”.

What triggers a review

Reviews only happen for actions that would otherwise stop for a person. The docs list five kinds:

  • shell commands that ask for escalated sandbox permissions
  • network requests blocked by the current sandbox or policy
  • file edits outside the allowed folders
  • MCP or app tool calls that require approval
  • Computer Use access to a new website or domain

Routine actions already allowed inside the sandbox are never reviewed. That is why most actions run without any check at all.

What Auto-review is designed to block

OpenAI lists four categories: sending private data, secrets or credentials to untrusted places; probing for credentials, tokens, cookies or session data; broad or lasting weakening of security; and destructive actions with a significant risk of irreversible damage. The full policy is published in the open-source Codex repository.

Auto-review by the Numbers

auto review free chatgpt codex all users d child safety gate at the foot of a staircase

OpenAI’s alignment team published an evaluation of Auto-review on 30 April 2026, a week after the feature first shipped. Its headline claim is that sessions “stop for human approval roughly 200x less often than in manual approval mode, while still catching many of the actions humans would want stopped”.

MeasureResultWhat it tests
Approval rate on escalated actions99.1%Usability on real internal traffic
Effective approval rate on all actions99.93%Including actions that never left the sandbox
Overeagerness recall90.3%Risky actions an agent takes while trying too hard
Prompt injection recall99.3% (90.2% across all categories)Code execution, secret theft and external upload attacks
MonitoringBench (hard) recall96.1%Agents covertly attempting malicious tasks

A snapshot of 10,000 actions

The paper includes an illustrative snapshot from OpenAI’s internal use. Of 10,000 actions, 9,280 stayed inside the sandbox and needed no review. The other 720 went to the reviewer, which approved 713 and denied 7. Of those 7, 4 continued by a safer route and 3 stopped to ask the user.

The arithmetic is worth stating. Only 7.2% of actions (720 of 10,000) reached the reviewer, and only 0.97% of those (7 of 720) were denied. Adding the 9,280 sandbox actions to the 713 approvals gives 9,993 of 10,000 actions running unattended, which is the 99.93% in the table.

Interruptions per 10,000 Codex actions in OpenAI’s snapshot (manual approval = 100%)
Manual approval: every out-of-sandbox action stops for a person 720 (100%)
With the reviewer on: actions denied 7 (0.97%)
With the reviewer on: denials that stopped to ask the user 3 (0.42%)

In this snapshot the person was interrupted 3 times instead of 720, a 240-fold reduction (720 divided by 3). OpenAI rounds the general figure to “roughly 200x” and notes that ratios depend on the environment and sandbox settings.

Inside OpenAI

The paper also says that “a majority of Codex Desktop token usage within OpenAI comes from Auto-review mode, and that share is growing”. At the time it was written, the reviewer ran on GPT-5.4 Thinking at a low reasoning setting, and OpenAI noted that “more powerful models are better at both risk detection and user intention discernment”.

Three ways to run Codex

The snapshot makes it easier to compare the three permission modes side by side. The interruption counts below use OpenAI’s illustrative 10,000-action snapshot, so treat them as an example rather than a forecast for your own projects.

ModeWho approves boundary crossingsInterruptions per 10,000 actionsWhat still blocks a harmful action
Manual approval (default)You, one prompt at a time720The sandbox and your judgement on every prompt
Reviewer mode (“Approve for me”)A separate reviewer agent3The sandbox, the reviewer policy and the circuit breaker
Full accessNobody0Only the model’s own training

The middle row is the point of the whole feature. It keeps the sandbox in place while removing almost all of the interruptions that push people towards the bottom row.

Why Approval Fatigue Is a Security Problem

auto review free chatgpt codex all users e robot lawnmower inside a boundary wire

It is tempting to see the manual mode as the safe option and an AI reviewer as a compromise. OpenAI argues the opposite, and its reasoning is persuasive for anyone who has watched developers use coding agents.

People switch to full access

When an agent stops every few minutes to ask permission for a mundane network call, many users give up and switch to Full Access mode. In that mode nothing at the environment level blocks a harmful action. The paper lists “malicious code execution, data exfiltration, credential probing” as things the sandbox would no longer stop.

People write rules that are too broad

Others write permission rules that let whole families of commands through. OpenAI says that in its internal traffic it found “a sizable minority of users who allow all commands that begin with the word python”. Its “favorite discovery” was a configuration that always allowed the agent to run with its fully permissive flag.

People approve without reading

The third failure is the most human. After the fiftieth prompt, people click approve without understanding what they are approving. A reviewer that never gets tired is, in that sense, a better gatekeeper than a person on their fiftieth click. Our earlier report on a coding agent that deleted 48,000 files in 103 seconds shows how much damage one approved command can do.

What Auto-review Will Not Do

auto review free chatgpt codex all users f hamster inside a clear exercise ball

OpenAI is unusually direct about the limits of its reviewer, and anyone enabling it should read them.

It is not a security guarantee

The documentation says the feature “is not a deterministic security guarantee”. The alignment paper goes further: during automated and human red-teaming, OpenAI “identified cases where Auto-review could be misled into approving commands without user approval”. It does not expect this kind of system “to become a source of deterministic guarantees”.

It does not stop a scheming model

The paper also warns that the reviewer “should not be understood as a mechanism for protecting against model scheming”. A model hiding its intent could produce a request that looks harmless, or act inside the sandbox where the reviewer never looks. That needs other tools, such as monitoring the model’s reasoning.

It only sees boundary crossings

The reviewer only judges actions that ask to leave the sandbox. If your sandbox allows too much, the reviewer never sees the dangerous action at all. A network address on an allowlist, for example, does not trigger a review by itself.

Some modes skip it entirely

With approval set to never, with full access, or with the fully permissive command-line flag, actions can avoid creating the approval request that review depends on. App approvals for Computer Use also still go straight to the user. Turning on Auto-review does not protect a session configured to bypass approvals.

Denials, Circuit Breakers and Overrides

A denial from the reviewer is not just a refusal. It shapes what the main agent does next.

What happens after a denial

The reviewer’s reason is passed back to the main agent with a firm instruction: do not pursue the same outcome through a workaround, continue only with a “materially safer alternative”, or stop and ask the user. OpenAI says that internally, Codex finds an acceptable route after a denial in more than half of cases.

The circuit breaker

To stop an agent from repeatedly probing the reviewer, Codex has a per-turn circuit breaker. In the current open-source code it interrupts the turn after 3 denials in a row, or after 10 denials within a rolling window of the last 50 reviews. Any approval resets the consecutive count. Timeouts are reported separately and are not treated as proof that an action was unsafe.

Overriding a denial

If the reviewer blocks something you actually want, the terminal interface offers an /approve command. It opens a list of recent denials, up to 10 per task, and lets you approve one exact action for one retry. That retry still goes through Auto-review, which sees your override as context but can still deny it if policy forbids that class of action.

Configuring Auto-review for a Business

For an individual developer, switching the reviewer on in settings is enough. For a business, the configuration options matter more.

Turning it on

In configuration files, the reviewer needs an interactive approval policy such as “on-request” and the setting approvals_reviewer = "auto_review". It has nothing to review if approvals are set to never.

Custom policies

The default reviewer policy lives in the open-source Codex repository. Individuals can replace it locally with an [auto_review].policy block, and organisations can set guardian_policy_config in managed requirements, which takes precedence. OpenAI’s advice is to copy the complete default policy first and then add rules, because a custom policy replaces rather than merges with the existing one.

Locking it in

Managed requirements can restrict users to approved settings, for example allowing only the “on-request” policy, only the AI reviewer and only read-only or workspace-write sandboxes. Full-access profiles can be left out so they are denied. For regulated teams, this is the part to involve your cybersecurity lead in.

Reducing review volume safely

If too many routine actions reach the reviewer, OpenAI’s advice is to fix the boundary rather than loosen the reviewer. Add narrow writable folders and precise command rules, such as allowing “cargo test” rather than every Python command. Auto-review transcripts are kept under ~/.codex/sessions, so you can analyse past traffic before changing anything.

Why Making Auto-review Free Matters

The feature itself has not changed. The economics have, and that is what will change behaviour.

The hidden cost of a second agent

Every review is work done by a second model. Before this change, that work drew on the same usage allowance as the main agent. A careful developer who turned on the reviewer was, in effect, spending part of their plan on safety. A careless one who chose full access spent nothing.

Sottiaux’s post removes that trade-off for users signed in with a ChatGPT account: the review “does not draw usage from your plan”. Choosing the safer setting no longer costs anything in usage, and it saves the time that manual approval takes.

Who is covered

The free change applies to people “signed in through a ChatGPT account”. OpenAI’s pricing page lists Auto-review as available on Plus, Pro, Business, Enterprise and API access, but the post does not say that API-key usage is free. Teams using Codex through an API key should check how reviews are billed before assuming they cost nothing.

A nudge towards a safer default

OpenAI’s own desktop app already switches the permission control to the reviewer mode, labelled “Approve for me”, when a user selects one of its approved security models. Making the reviewer free extends that nudge to everyone. If the result is fewer developers running agents in full access mode, the security benefit could be substantial.

Questions to Settle Before Switching a Team Over

Turning on an AI reviewer for one developer is a settings change. Turning it on across a team is a policy decision, and a few questions are worth answering first.

Which projects are in scope?

Not every repository carries the same risk. A prototype with no credentials and no production access is a good place to start. A repository that can deploy infrastructure or touch customer data deserves a tighter sandbox and more human sign-off, even with a reviewer in place. List your repositories and decide which ones move first.

Who owns the reviewer policy?

If you customise the policy, someone has to own it. That person should keep a copy of the default wording, record every change and the reason for it, and review it when OpenAI updates the default. A policy that nobody owns tends to grow exceptions until it no longer protects anything.

How will you handle overrides?

The override command lets a developer approve one denied action once. Decide whether overrides need to be logged or explained, and who reviews them. A pattern of repeated overrides on the same kind of action is a sign that either the sandbox or the policy needs attention.

What evidence will you keep?

Session transcripts are stored locally by default. If your organisation needs an audit trail, decide where transcripts and denial records should be kept, for how long, and who can read them. That decision belongs with whoever owns your IT governance.

What Auto-review Means for UK Development Teams

For UK businesses using AI coding agents, this is a low-cost improvement worth acting on this week.

Make Auto-review the default

Ask developers to switch from manual approval or full access to the reviewer mode, and record that choice in your development policy. For managed deployments, enforce it through requirements files.

Tighten the sandbox first

Review which folders and commands your agents are allowed to use without approval. Remove broad rules, especially ones that allow whole interpreters or network tools.

Keep humans on the high-risk work

The reviewer is designed for routine boundary crossings. Deployments to production, changes to access controls and anything touching customer data should still need a named person’s approval, whatever the reviewer says. Our IT security team recommends writing those exceptions down.

Log and review denials

Denials are useful signals. A cluster of denied actions on one project can reveal a misconfigured agent, a risky instruction or an attempted prompt injection. Build a habit of reading them.

Auto-review: Frequently Asked Questions

What is Auto-review in ChatGPT and Codex?

It is a setting in Codex that sends an agent’s requests to cross its sandbox boundary to a separate reviewer agent instead of to you. The reviewer approves or denies each request and explains why.

Is Auto-review free?

Yes, for users signed in through a ChatGPT account. OpenAI’s Thibault Sottiaux said on 6 October that it “does not draw usage from your plan”.

How do I turn on Auto-review?

Open Settings, then General, then Permissions, and choose the reviewer option.

Is Auto-review safe enough to replace human approval?

Not entirely. OpenAI says it is not a deterministic guarantee, red-teaming found ways to mislead it, and it does not protect against a model that is hiding its intent. Keep human sign-off for high-risk actions.

What happens if Auto-review keeps saying no?

After 3 denials in a row, or 10 within the last 50 reviews in one turn, Codex stops the turn. You can approve a specific denied action once using the /approve command.

References