Strands Box is a new open-source sandbox from Amazon Web Services that puts hard limits on what AI agents can do on a developer’s machine, and it remembers what an agent has already done when it decides what the agent may do next. AWS released Strands Box in developer preview on Wednesday 7 October 2026 under the Apache 2.0 licence, and InfoWorld framed the launch as AWS taking aim at runaway AI agent behaviour.
The pitch is simple to state. Coding assistants and home-grown agents increasingly approve their own actions, a setup AWS calls “YOLO mode”. Strands Box wraps the agent in operating-system isolation, then routes its shell commands, Python scripts, network requests and Model Context Protocol (MCP) tool calls through a policy engine that can say no. A rule can let an agent post to Slack, but only three times every ten minutes. Another can block outbound web requests once the agent has read a file from a customer-data folder.
This article explains what launched, how Strands Box works, what the Dogwood policy language underneath it adds, and where the design makes trade-offs. It also compares Strands Box with other agent sandboxes and sets out what a team should do before trusting it. For the wider picture of how agents go wrong, see our report on the five main ways OpenAI says rogue AI agents are misbehaving online.
Table of contents
- What AWS Launched With Strands Box
- The Problem Strands Box Is Built to Solve
- How Strands Box Works: Containment Plus Policy
- Temporal Rules: How Strands Box Remembers What an Agent Did
- How Strands Box Keeps Credentials Away From Agents
- Dogwood: The Policy Language Underneath Strands Box
- The AWS Agent-Safety Stack Around Strands Box
- Strands Box Compared With Other Agent Sandboxes
- The Trade-Offs in the Strands Box Design
- What Analysts Say About Strands Box
- The Strands Box Roadmap: Linux, Windows and the Cloud
- Should Your Team Try Strands Box Now?
- Strands Box FAQ
- References
What AWS Launched With Strands Box
Strands Box is a local sandbox for running AI agents on a Mac. You choose the agent program to run, describe its environment in one configuration file, write its rules in a second file, and start it with a single command. Everything the agent tries to do that crosses the sandbox boundary passes through Strands Box first, and the policy engine either permits or forbids it.
The AWS launch post, written by Principal Engineer Fernando Dingler, describes Strands Box as combining “operating-system isolation with fine-grained policies for governing agent actions”. The code is on GitHub at strands-agents/box, written in Rust, and the first tagged release, v0.1.0, was published on 7 October at 17:24 UTC.
The developer preview at a glance
| Item | Strands Box developer preview | Source |
|---|---|---|
| Released | 7 October 2026, developer preview, release v0.1.0 | AWS, GitHub |
| Licence and price | Apache 2.0, free and open source | AWS |
| Platform today | Macs with Apple silicon on macOS 15 or later | InfoWorld, GitHub README |
| Language | Rust, built as a Cargo workspace | GitHub |
| Configuration | box.toml for the environment, policy.dw for the rules | AWS |
| Policy language | Dogwood, evaluated by the embedded Dogwood Local Engine | AWS |
| Enforcement points | Egress gateway, Shell interpreter, Python interpreter, MCP broker | AWS |
| Default stance | Operations checked by the engine are denied unless a rule permits them | GitHub README |
| Decision records | OTLP JSON, written to a records.jsonl file inside the box directory | GitHub README |
| Agent support | Harness-agnostic: any agent program or binary | AWS, The Register |
Who built it and how it fits the Strands family
Strands Box sits under the Strands brand, which AWS uses for its open-source agent tooling, including the Strands Agents SDK it open-sourced in May 2025. It is not tied to that SDK, though. Marc Brooker, an AWS vice president and distinguished engineer who worked on Dogwood and Strands Box, wrote in an accompanying post that “you can use Box with any agent framework or harness”, adding that AWS would still love people to try the Strands harness.
The getting-started guide in the README runs the Strands CLI inside Strands Box. To follow it you need a Mac with Apple silicon, macOS 15 or later, Node.js 22.21 or later from Homebrew, and access to Claude Opus 5 on Amazon Bedrock in the us-west-2 region. Those are requirements of the guide’s example agent, not of Strands Box itself, which runs whatever program you point it at.
Why the headline says “runaway”
InfoWorld’s headline speaks of runaway AI agent behaviour, and The Register went further, calling Strands Box a way to “prevent YOLO mode disasters”. The worry is not that agents are malicious. It is that they act fast, repeat themselves and take literal instructions in directions nobody expected. We documented one case in the Claude Code session that deleted 48,000 files in 103 seconds. Strands Box is AWS’s answer to the question of how to keep that kind of momentum inside fixed lines.
The Problem Strands Box Is Built to Solve
AWS opens its launch post with a plain description of the risk. Developers delegate more work to AI agents, so the agents need broader access. “They read files, run shell commands, execute generated code, and call APIs,” the post says. “That is what makes agents useful, but that same autonomy is the source of the biggest risks.”
YOLO mode is now normal
“Coding assistants, home-grown harnesses, and off-the-shelf agents increasingly run in ‘YOLO mode,’ approving every action without human review,” AWS writes. Agents “can wander out of their working directory, run a dangerous command, or reach a credential they were not meant to see.” Approval prompts slow people down, so people switch them off, and the agent’s judgement becomes the only check left.
Why a plain sandbox is not enough
The usual fix is a sandbox that limits what an agent can reach. AWS’s argument is that reach is only half the job. “An agent investigating a production incident might need to read logs or inspect infrastructure, without being allowed to change it,” the post says. “Containers and microVMs provide strong isolation, but isolation alone doesn’t enforce contextual rules. Once an agent can reach a tool, we still need to define what it can do with it.”
That is the gap Strands Box targets. A container can stop an agent reaching a database host. It cannot easily say “read from this database, never write to it, and stop querying after fifty calls an hour”, because by the time traffic leaves a container it is usually encrypted and stripped of meaning.
Why harness permissions fall short
Many agent tools ship their own permission prompts and allow lists. AWS says these help but have two weaknesses. First, they run inside the agent’s own process and “judge the tool call, not its effect”. A harness sees “run this shell command”, not which files the command will touch or which hosts it will reach. Second, every harness uses a different rule format, so a policy written for one does not carry to another.
Strands Box enforces from outside the agent process, at the operating-system and network boundary. The same box.toml and policy.dw files can govern different agent applications, which is the consistency point analysts picked up on.
How Strands Box Works: Containment Plus Policy
Strands Box relies on two layers. Containment sets the hard boundary. Policy decides what happens at each crossing point. Brooker’s description is that the sandbox mechanisms “make sure that the policy layer is the only way out of the box”.
Layer one: operating-system containment
On macOS, Strands Box uses the operating system’s own isolation, such as Seatbelt, to fix which files, programs and network routes the agent can reach at all. Brooker calls it “a kernel-enforced sandbox, using the same OS-provided mechanisms as many other sandbox offerings”. The difference is what the sandbox is used for. Instead of writing all the rules at the kernel layer, Strands Box uses containment to funnel every interaction towards a policy layer in user space.
Layer two: Dogwood policy
Inside that boundary, rules written in Dogwood decide which actions the agent may take. Dogwood uses the syntax of Cedar, the authorisation language AWS created earlier: each rule is a permit or a forbid over a principal, an action and a resource, with conditions. In Strands Box the principal is always the agent and the resource is fixed, so rules match on the action and on details such as the host, the file path or the program being run. A forbid always overrides a permit.
The four enforcement points in Strands Box
| Enforcement point | What it sees | Example event | Example rule |
|---|---|---|---|
| Egress gateway | Host, port, HTTP method and path of every outbound request | http:request | Allow HTTPS only to *.us-west-2.amazonaws.com |
| Strands Shell | Commands, programs launched and the files each command touches | shell:spawn, fs:delete | Forbid file deletion outside the workspace |
| Monty for Python | File access and network calls made by agent-written Python | fs:read, http:request | Block uploads after a sensitive read |
| MCP broker | Which MCP tools the agent calls and with what arguments | mcp:call | Cap calls to the AWS MCP server at 60 an hour |
The two interpreters are worth explaining. Strands Shell is AWS’s own Rust shell for agents, and Monty is a minimal Python interpreter in Rust from Pydantic. Strands Box runs both in its own process, outside the sandbox, and the agent reaches them as ordinary bash and python3 commands over a local socket. Because the interpreter itself resolves each command, the policy sees meaning rather than raw system calls.
AWS gives a concrete case. If the agent runs rm -rf build/, Strands Shell “raises an fs:delete decision for each file it would remove, so a forbid rule on fs:delete stops it before anything is deleted, while a normal edit inside the workspace still goes through.” That is a rule written in terms a developer actually cares about.
One event history across every tool in Strands Box
All four enforcement points share one history of events, and they report actions the same way whichever tool performed them. A file read through a shell command and a file read through a Python script are both an fs:read event. An HTTP request from curl and one from Python are both an http:request event.
That shared record is what makes cross-tool rules possible. AWS’s example is “After the agent reads a file from the customer-data directory, block further outbound HTTP requests”, written without saying which tool did the read. An agent cannot dodge the rule by switching from shell to Python halfway through a task.
The star counts come from the GitHub API on 9 October 2026, two days after launch, and bar widths are each count divided by Dogwood’s 437. Pydantic’s Monty is left out because it is a separate, older project with 8,604 stars that would flatten the rest of the chart. The point is scale: Strands Box had more than 200 stars two days after launch, but every piece of this stack is still young.
Temporal Rules: How Strands Box Remembers What an Agent Did
Most permission systems judge each request on its own. Strands Box can judge a request against what the agent has already done, when and in what order. AWS says rules can depend on “earlier actions, their order and limits accumulated over time”. This is the feature AWS is most keen to stress, and it is what separates Strands Box from a plain allow list.
The Slack rate-limit example
AWS’s worked example is an on-call agent investigating a production incident. It should post progress to the incident channel in Slack, but “not to flood the channel and bury the updates from humans”. One rule permits posts to Slack’s chat.postMessage endpoint. A second rule forbids a post once Strands Box has recorded three successful (HTTP 200) responses from that endpoint in the last ten minutes. The agent keeps pulling logs the whole time.
| Time | Agent action | Strands Box decision |
|---|---|---|
| 10:00 | Post an update | Allow |
| 10:03 | Post an update | Allow |
| 10:05 | Post an update | Allow |
| 10:06 | Post a fourth update within 10 minutes | Deny |
| 10:07 | Retrieve more logs | Allow |
| 10:11 | Post an update | Allow |
The timeline comes straight from AWS’s post. The chart below adds one step of arithmetic: at each posting attempt, how many successful posts already sit inside the trailing ten-minute window. At 10:11 the 10:00 post has aged out, leaving only the 10:03 and 10:05 posts, so the fourth post goes through.
Refusals the agent can read
At 10:06 the agent does not hit a silent wall. The gateway returns an HTTP 403 whose body says “policy denied this operation”, names the rule by its identifier and includes the rule’s description. AWS’s sample description reads “Slack posts are capped at three every 10 minutes. Wait before you post again.” A clear description lets the agent adjust, for example by waiting rather than retrying in a loop.
The rule counts successful responses, not attempts, so refusals do not count towards the three. AWS notes that binding the rule to the request event instead would count attempts, and a refused post would then use up part of the allowance. That choice has a catch we return to in the Dogwood section below.
Other rules AWS has shown for Strands Box
The launch material includes several more examples, each a single rule:
- Tested pushes only. Brooker shows a policy that forbids
git pushunless a successfulnpm testran in the project within the last 15 minutes, with nogit addsince. He notes it is not a complete policy, because a branch switch could still break the promise. - A spending ceiling. “Allow this agent to use the payments API, but only up to a total of $100 per day.”
- A call budget. Limit the agent to 60 calls an hour to the AWS MCP server, which Brooker says is useful for paid APIs, shared quotas, “or simply to prevent unexpected bills”.
- Read-then-quiet. Block outbound HTTP once the agent has read from a customer-data directory.
An AWS spokesperson told The Register that Strands Box could also control when an agent may run a Git push, or cap API calls that, in The Register’s words, “could end up costing a small fortune”.
How Strands Box Keeps Credentials Away From Agents
The egress gateway does one more job that matters for security teams. Every outbound request goes through it by default, and because it sits in that path it can attach credentials on the way out. For configured routes, the agent receives a placeholder token. The gateway swaps in the real secret only when policy permits the request. “The real secret never enters the agent’s environment,” AWS says.
Five authentication methods
| Method | What the Strands Box gateway does |
|---|---|
| Bearer token | Replaces the placeholder with the real token in the Authorization header |
| Custom header | Injects the credential into a configured header, such as x-api-key |
| HTTP Basic | Sets the Authorization header from a configured username-and-password value |
| Query parameter | Places the credential in a named query parameter |
| AWS SigV4 | Signs the request with AWS credentials obtained outside the agent |
For cybersecurity teams, this closes a common leak path. An agent that can read its own environment variables can paste an API key into a log, a commit or a chat message. An agent that only ever holds a placeholder has nothing worth leaking. It mirrors the thinking in Opal Zero’s approach to risky AI agent access permissions, where the aim is to stop agents holding standing access in the first place.
A worked on-call example
AWS’s sample configuration gives the on-call agent three things: the Anthropic API for its model, Slack for updates, and the AWS CLI. In box.toml, the AWS CLI is declared as a tool by its absolute path, because Strands Box does not inherit your PATH. Outbound traffic is limited to AWS endpoints in us-west-2, api.anthropic.com and slack.com, each with a credential binding.
Neither the agent nor the CLI ever sees real AWS credentials. The gateway reads session credentials from an “oncall” AWS profile on the host and signs each permitted request. The IAM role behind that profile only allows reading logs and inspecting infrastructure, so even a fully permitted request cannot change production. A credential binding grants no network reach by itself, so each destination also needs an explicit permit rule in policy.dw. The agent is then started with box run --config box.toml.
Dogwood: The Policy Language Underneath Strands Box
Strands Box builds directly on two open-source releases AWS made earlier this year. Dogwood is the policy language. The Dogwood Local Engine evaluates Dogwood rules against an agent’s requested action and recorded history. Strands Box embeds that engine and enforces its allow or deny decisions.
From Cedar to Dogwood
Cedar, the language behind AWS’s AgentCore Policy service, judges one request at a time. As InfoQ explained when Dogwood launched in August, feed Cedar the same request twice and you get the same answer regardless of what happened before. That predictability is what allows Cedar policies to be formally analysed, but it means Cedar cannot describe a sequence of actions.
Dogwood adds a second clause type, “when temporal”, that reads the agent’s event history. Any valid Cedar policy is also a valid Dogwood policy, so existing rules keep working. InfoQ reported four standard operators: one for whether something happened within a time window, one for how many times, one for how many distinct values, and one for a running total.
The concurrency trap, and what it means for Strands Box rules
InfoQ highlighted a correctness trap from AWS’s own Dogwood material. Suppose a rule caps money transfers at $5,000, and three $2,000 transfers arrive at once, before any of them settles. A policy that sums completed responses sees nothing finished yet and lets all three through. A policy that sums requests denies the third.
The arithmetic is three times $2,000, or $6,000, which is $1,000 over the cap, against two times $2,000, or $4,000, when the third request is refused. This matters for Strands Box because AWS’s own Slack example counts responses so that refused posts do not use up the allowance. That is sensible for a gentle rate limit on chat messages. For anything involving money, deletions or data leaving the building, an agent firing parallel calls could slip past a response-based cap. Teams writing their first Strands Box policies should decide, rule by rule, whether they are counting what was attempted or what completed.
What InfoQ said about Dogwood’s limits
AWS was candid about Dogwood’s costs. Temporal evaluation needs stateful tracking, and evaluation time can grow with the length of the event log. Temporal conditions also lose Cedar’s automated reasoning tools, so a policy set that uses them can no longer be formally analysed. InfoQ added that AWS described the reference interpreter as a tool for exploring and testing the language, not for running authorisation in production, and that AWS was not yet accepting outside contributions.
The Dogwood Local Engine is the piece meant to address the runtime side. Brooker describes it as “a durable, fast, local runtime for Dogwood”, open-sourced the week before Strands Box. The full language guide is published at the Dogwood documentation site.
The AWS Agent-Safety Stack Around Strands Box
Brooker frames Strands Box as one step in a longer programme. AgentCore Runtime, launched in summer 2025, gives each agent session its own Firecracker microVM in the cloud. AgentCore Policy followed with Cedar-based control over how cloud-hosted agents use tools. Dogwood and its local engine extended that model to sequences of actions. Strands Box now brings the same kind of policy to a developer’s own machine.
| Component | When | Role |
|---|---|---|
| AgentCore Runtime | Summer 2025 | Cloud sandbox, one Firecracker microVM per agent session |
| AgentCore Policy | re:Invent 2025 | Deterministic tool-call control in the cloud, written in Cedar |
| Strands Shell | Repository created 27 February 2026 | A shell for agents that can be governed by policy |
| Dogwood | Repository created 27 July 2026, reported 16 August 2026 | Cedar plus temporal rules over an agent’s history |
| Dogwood Local Engine | Open-sourced the week before Strands Box | Local runtime that evaluates Dogwood rules |
| Strands Box | 7 October 2026 | Local sandbox that enforces Dogwood decisions around an agent |
The dates are GitHub repository creation dates from the GitHub API, plus InfoQ’s publication date, each subtracted from 7 October 2026. A repository can be created privately before it is made public, so these are upper bounds on how long each piece has been visible. The pattern is still clear: AWS built the policy language first, then the evaluator, then the sandbox that uses both.
Strands Box Compared With Other Agent Sandboxes
Strands Box is not the only way to contain an agent, and AWS says so. The useful comparison is with the built-in sandbox in coding tools such as Claude Code, with Nvidia’s open-source OpenShell, and with a dedicated microVM per session.
| Approach | Isolation | What policy can see | History-aware rules | Platforms |
|---|---|---|---|---|
| Strands Box | macOS Seatbelt, policy layer as the only route out | Shell and Python operations, HTTP method and path, MCP tool calls | Yes, via Dogwood | macOS on Apple silicon; Linux in development |
| Claude Code Bash sandbox | Seatbelt on macOS, bubblewrap on Linux and WSL2 | Shell commands only; file tools and MCP servers run outside it | No | macOS, Linux, WSL2; off by default |
| Nvidia OpenShell | Landlock and seccomp on Linux, supervisor decides policy | Network endpoints and file access; credentials added for approved endpoints | Not described | Linux |
| MicroVM per session | Hardware virtualisation, such as Firecracker | Mostly encrypted network traffic and raw disk blocks | Not at this layer | Cloud, for example AgentCore Runtime |
The Claude Code row comes from Anthropic’s sandboxing documentation, which says the sandbox covers shell commands only, allows writes to the working directory and a temp directory by default, and routes network traffic through a local proxy that checks each host against an allow list that starts empty. It is a good boundary, but it judges hosts and paths, not sequences. The OpenShell row draws on our coverage of Nvidia’s open-source answer to rogue agents, which takes a similar “the sandbox never decides policy” line on Linux.
Why AWS did not just use a microVM for Strands Box
Brooker’s post walks through the layers where a sandbox can apply rules. At the virtual hardware layer, a microVM has a small trusted base and strong security, but almost everything it sees is TLS ciphertext and “read N bytes at offset X”. That is a great place for blunt rules such as “all traffic goes to this gateway”, and a poor place for rich ones. At the system-call layer, used by containers, you see “open this file” or “connect to this host”, but network traffic is still encrypted and the system-call surface keeps growing.
AWS also wanted agents to work inside a developer’s existing environment. A separate container or VM “introduces another environment to provision and maintain, along with decisions about how to expose local files and tools to it”, the launch post says. Strands Box keeps the agent on the host, behind kernel isolation, and gives it exactly one way out.
Where Strands Box pairs with a microVM
For cloud use, Brooker does not argue for Strands Box alone. “We recommend for multi-tenant deployments in the cloud that you pair it with a dedicated per-session MicroVM,” he writes. The simplest route is running Strands Box inside an AgentCore agent, though the same design works on EC2, Lambda or Fargate. That gives fine-grained local policy from Strands Box, tool-call policy from AgentCore Policy, and a hard virtualisation wall between sessions.
The Trade-Offs in the Strands Box Design
AWS and Brooker are both open about where the design gives ground. Three trade-offs stand out, and the analysts InfoWorld spoke to added two more.
A wider trusted computing base
The interpreters and protocol interceptors that enforce policy run outside the sandbox, in Strands Box’s own trusted process. “Running the interpreters outside of the sandbox is a deliberate trade-off,” AWS writes. “They are the code that enforces policy, which means the box’s trusted computing base is wider.” Brooker puts it more bluntly: “more code means more opportunity for bugs, and in this space, bugs could mean bypassed policy.” AWS says it has written that code in Rust and invested in testing, fuzzing and validation.
Policy does not see everything yet
Paths granted to the agent directly in box.toml, such as a project folder that the harness’s own file tool opens, are bounded by containment only. They do not produce individual Dogwood decisions and do not appear in the policy history. InfoWorld made the same point: files accessed through an agent harness’s built-in tools are subject to operating-system restrictions but not evaluated by Dogwood. To apply a policy to file operations, the README says, use the interpreters and keep those files out of the direct grants. AWS says it wants policy to cover more over time, helped by the Endpoint Security additions in macOS 27.
Rules are hard to write today
Brooker concedes that Dogwood rules may look “difficult to understand, and difficult to write”, and calls this “a major area of investment” for the coming months. For now, AWS has released a policy-authoring agent skill so a coding agent can draft Dogwood rules and install them in a box. Brooker’s view is that “frontier coding agents are already great at reading and writing Dogwood”. That is convenient, but it also means an AI may be writing the rules meant to contain an AI, so human review of each rule still matters.
Overhead and badly tuned policy
Pareekh Jain, chief executive of Pareekh Consulting, told InfoWorld the extra controls “could increase processing overhead and introduce new components that might themselves contain vulnerabilities”. Tulika Sheel, senior vice president at Kadence International, warned that “poorly designed policies could block legitimate agent actions or create operational complexity, while overly permissive policies could still leave gaps”.
What Strands Box cannot stop
“It cannot prevent every harmful decision an agent makes within its allowed permissions,” Jain said. “Enterprises will still need IAM, monitoring, and human oversight.” Brooker made the same point to The Register: even with correct permissions, an agentic action can still produce an unwanted result, and “developers remain responsible for deciding what access to grant and where human review is needed.” Monitoring remains a separate layer, as our guide to AI agent monitoring in production sets out.
| Limitation | Who raised it | Practical response |
|---|---|---|
| Wider trusted computing base | AWS, Brooker | Pair Strands Box with a per-session microVM for multi-tenant cloud use |
| Direct file grants bypass Dogwood | AWS, InfoWorld | Keep sensitive folders out of box.toml grants; reach them through the interpreters |
| Response-based caps can be raced | InfoQ on Dogwood | Count requests, not responses, for money, deletions and data egress |
| Processing overhead, new code paths | Pareekh Jain | Measure latency on real tasks before rolling out widely |
| Policies too tight or too loose | Tulika Sheel | Start default-deny, read the decision records, loosen one rule at a time |
| Harm inside allowed permissions | Jain, Brooker | Keep least-privilege IAM, monitoring and human review |
What Analysts Say About Strands Box
The two analysts InfoWorld quoted were broadly positive about the idea and cautious about the evidence.
A real gap, filled with familiar parts
“Strands Box addresses a real security gap, although its underlying technologies are not new,” Jain said. “Its main advantage is making security easier to enforce consistently across different AI agent frameworks.” That matches AWS’s own pitch: OS sandboxes, proxies and policy languages all existed already, and the new part is wiring them into one place that every agent has to pass through.
Portability is the real test
Jain said a common policy layer could let developers focus on building agents while security teams maintain shared rules, but that adoption would depend on broader platform support and how much overhead enforcement adds. Sheel said the open-source approach could help adoption, but enterprises would need evidence that the controls work reliably in production. Her longer view was more ambitious: “As agents become more autonomous, behavioral controls could become as fundamental to AI infrastructure as identity and access management are today.”
The Strands Box Roadmap: Linux, Windows and the Cloud
AWS’s launch post lists four next steps, and The Register added detail on timing. None of them has a date.
More operating systems
Expanding beyond macOS is “one of our next priorities”, AWS says, with the aim of letting developers carry the same box.toml and Dogwood policies to the environments where they already work. AWS told The Register that Linux support is in development and that a Windows client is “on our radar”, with no planned release date for either. Each operating system offers different isolation tools, so the work includes proving that those tools preserve the guarantees Strands Box relies on.
An easier setup
AWS is building a command-line tool that will detect the agent harnesses already installed on a machine and generate a box.toml and a baseline Dogwood policy with recommended defaults. That would replace the current step of writing both files by hand.
Strands Box for deployed agents
AWS wants developers building custom agents with the Strands Harness SDK, LangChain or the Claude Agent SDK to package their agent with Strands Box and deploy it to Amazon Bedrock AgentCore, Amazon ECS or Kubernetes, with the policies travelling alongside the agent. This is the portability test the analysts described.
Liveness rules
Today’s Dogwood rules describe what must not happen. AWS wants to add liveness, meaning rules about what must eventually happen, so developers can detect when an agent fails to meet an obligation. An example might be “after changing a configuration, the agent must run the health check”.
Should Your Team Try Strands Box Now?
For most organisations the honest answer is: try it, but do not depend on it yet. Strands Box is a version 0.1 developer preview on one operating system, built on a policy engine AWS itself calls young. The design is sound and the documentation is unusually frank, which makes it worth learning now.
Who Strands Box suits today
It suits developer teams on Apple silicon Macs who already let coding agents run commands without approval, and who want something firmer than the agent’s own settings. It also suits platform and security teams who want to prototype agent policies now, so they are ready if AWS delivers the Linux and AgentCore support it has promised. It does not yet suit Windows estates, Linux build servers or production cloud agents.
A practical first-week plan
- Pick one agent and one task. A coding agent fixing tests in a single repository is ideal.
- Start from default deny. Permit only the model API, your package registry and the folders the task needs.
- Keep sensitive files out of direct grants. Route access through Strands Shell or Python so Dogwood can see and log it.
- Add one temporal rule. A tested-push rule or an API call budget shows the value quickly.
- Read the decision records. The OTLP JSON log shows every allow and deny, which is how you tune rules without guessing.
- Test refusals deliberately. Ask the agent to delete a folder or reach a blocked host and confirm the 403 explains itself.
- Write down what is still unprotected. Built-in file tools on direct grants, harm inside allowed actions, and anything the agent does outside the box.
Before any of this reaches production, apply the disciplines in our guide to testing AI agents before production. A sandbox narrows what can go wrong. It does not prove an agent behaves well.
Questions to put to any agent sandbox
Whether you choose Strands Box or a rival, ask the same questions. Can policy see the effect of a command, or only its name? Do rules carry across tools, or can an agent switch from shell to Python to escape one? Where do credentials live, and can the agent read them? Does the policy count attempts or completions? Who writes the rules, and who reviews them? And what is the plan for the operating systems your developers actually use? Strands Box has good answers to most of these on a Mac today. The last one is still open.
Strands Box FAQ
What is Strands Box?
Strands Box is an open-source sandbox from AWS that runs AI agents with operating-system isolation and checks their shell commands, Python code, network requests and MCP tool calls against rules written in the Dogwood policy language. Rules can depend on what the agent has already done.
Is Strands Box free?
Yes. Strands Box is released under the Apache 2.0 licence and is free to download from GitHub. Running an agent inside it still costs whatever the agent’s model and APIs cost.
Does Strands Box work on Windows or Linux?
Not yet. The developer preview supports Macs with Apple silicon on macOS 15 or later. AWS says Linux support is in development and a Windows client is on its radar, without dates.
Does Strands Box only work with Strands agents?
No. AWS describes Strands Box as harness-agnostic. You can run any agent program or framework inside it, and AWS names LangChain and the Claude Agent SDK among the frameworks it wants to support for deployed agents.
Does Strands Box replace IAM and human review?
No. AWS and the analysts agree that Strands Box enforces the rules you write, deterministically, but cannot stop harmful choices made inside the permissions you grant. Least-privilege IAM roles, monitoring and human review are still needed.
References
InfoWorld: AWS takes aim at runaway AI agent behavior with Strands Box
AWS Open Source Blog: Introducing Strands Box, AI agent sandboxes powered by Dogwood
Strands Agents: Strands Box, the big picture (Marc Brooker)
The Register: AWS launches open-source AI agent sandbox to prevent YOLO mode disasters
InfoQ: AWS Open-Sources Dogwood, Extending Cedar to Govern Sequences of Agent Tool Calls
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.