亚马逊Score B (65)
Introducing Strands Box: AI agent sandboxes powered by Dogwood - AWS
4 小时前2 viewsSource: aws.amazon.com
Introducing Strands Box: AI agent sandboxes powered by Dogwood As developers increasingly delegate work to AI agents, these agents require broader access to systems and resources. They read files, run shell commands, execute generated code, and call APIs. That is what makes agents useful, but that same autonomy is the source of the biggest risks. Coding assistants, home-grown harnesses, and off-the-shelf agents increasingly run in “YOLO mode,” approving every action without human review. Agents can wander out of their working directory, run a dangerous command, or reach a credential they were not meant to see. The usual solution to this problem is a sandbox, which puts a boundary around what an agent can access. But access is only part of what we want to control. An agent investigating a production incident might need to read logs or inspect infrastructure, without being allowed to change it. Containers and microVMs provide strong isolation, but isolation alone doesn’t enforce contextual rules. Once an agent can reach a tool, we still need to define what it can do with it. Today, we’re launching Strands Box in developer preview, an open source sandbox licensed under Apache 2.0 that combines operating-system isolation with fine-grained policies for governing agent actions. Box builds on our earlier releases of Dogwood , an open source policy language, and the Dogwood Local Engine , which evaluates those policies against an agent’s requested action and recorded history. Box embeds that engine and enforces its allow or deny decisions around the running agent. For the bigger picture behind Strands Box and AWS’s investment in AI agent safety, read Marc Brooker’s accompanying post . By supporting policies written in Dogwood , Box lets us express rules that not only account for what the agent is trying to do, but also what it has already done. Permissions can depend on earlier actions, their order and limits accumulated over time. Consider an agent investigating a production incident. We want it to post progress updates to the incident channel in Slack as it finds things, but not to flood the channel and bury the updates from humans. A policy can let the agent post, but no more than three times every 10 minutes. The agent can keep investigating, while Box enforces the posting limit without relying on the agent to remember it. Inside Strands Box Strands Box relies on two layers: containment and policy. Containment uses OS-level isolation, such as macOS Seatbelt, to define what the agent can reach, on the host machine and on the network. It establishes the hard boundary. Within that boundary, policy governs which actions the agent can perform. Policy is applied at multiple enforcement points: network egress, a Python interpreter, a Shell interpreter and a broker for Model Context Protocol (MCP) servers. These enforcement points govern concrete operations: which files the agent’s shell commands and Python scripts can read or modify, which commands it can run, which HTTP methods and paths it can access, and which MCP tools it can call. These enforcement points use the embedded Dogwood Local Engine and share a history of events, allowing a rule to connect an earlier action through one tool with a later action through another. The egress gateway is a proxy that all of the box’s outbound traffic goes through, and the Shell interpreter, the Python interpreter, and the MCP broker run in Box’s own process, outside the sandbox, where the agent reaches them as ordinary bash, python3, and MCP commands over a local socket. Every enforcement point reports an action the same way, whichever interpreter performed it. A file read through a shell command or a Python script is an fs:read event; an HTTP request from curl or from a Python script is an http:request event. That is what lets you write a rule such as “After the agent reads a file from the customer-data directory, block further outbound HTTP requests” without specifying which tool did the read. Policy sees the operations that pass through these enforcement points. Paths you grant to the agent directly in box.toml , such as a project directory that the harness’s built-in file-reading tool opens, are bounded by containment instead and do not appear in the policy history. Over time, we want policy to cover more of what the agent does, until it is the one place you govern an agent. New operating system features will help us get there, such as the Endpoint Security additions in macOS 27, which we are eager to adopt. Why interpreters? Agents increasingly write code to get work done, and that work is often done in Python and Shell. For example, an agent doing data analysis could write a program that fetches logs from several places, groups errors and returns a summary. The intermediate results stay in the execution environment, reducing the data passed through the model and therefore reducing token spend. Because that code can read files and make network requests, code interpreters are a great place to enforce policy. This is why we chose to embed two code interpreters into Strands Box: Strands Shell ( https://github.com/strands-agents/shell ) and Monty for Python ( https://github.com/pydantic/monty ). These interpreters allow us to intercept operations such as filesystem access or network calls and run them through the policy engine for authorization. Here’s what that looks like in practice. If the agent runs rm -rf build/ , the Shell resolves the command and raises an fs:delete decision for each file it would remove, so a forbid rule on fs:delete stops it before anything is deleted, while a normal edit inside the workspace still goes through. The policy sees the file and the operation instead of an opaque syscall. This lets you write the rule in terms you actually care about. Running the interpreters outside of the sandbox is a deliberate trade-off. They are the code that enforces policy, which means the box’s trusted computing base is wider. Why not containers or microVMs? MicroVMs offer a strong isolation boundary and are a reasonable choice for running untrusted code. For Strands Box, we wanted the agents you run to operate within a developer’s existing environment, without having a separate guest operating system. A separate container or VM introduces another environment to provision and maintain, along with decisions about how to expose local files and tools to it. But the choice of isolation mechanism only answers part of the problem. We still have to solve the problem of policy: Once we run an agent and give it some tools, we still need to govern what it does with them. We believe that the sweet spot is OS-level containment to establish a hard boundary, paired with policy enforcement in various interception points where the agent’s actions attempt to cross the boundary. Why not the harness’s own permissions? Many off-the-shelf custom harnesses ship their own permission prompts and allow lists, so a fair question is why not rely on those. They help, but they run inside the agent’s own process and they judge the tool call, not its effect. A harness sees “run this shell command”. It does not see the files the command will touch or the hosts it will reach. Different harnesses also have different formats for rules, so a policy written for one may not carry to another. Box enforces from outside the process, at the OS and network boundary, and the same box.toml and policy.dw can be used for consistent configuration and policy enforcement across agent applications. By default, every outbound request from the box goes through the egress gateway. The gateway is the interception point for the network: it sees the host, port, method and path of each request, raises an http:request decision, and forwards only what policy permits. Because it sits in that path, it can also do something the agent cannot do safely on its own: attach credentials. For configured API-key routes, the agent receives a placeholder token, which the gateway replaces with the real secret before forwarding a permitted request. The real secret never enters the agent’s environment; the gateway attaches it on the way out. Strands Box supports these authentication methods today: Authentication method What the gateway does Bearer token Replaces the placeholder with the real token in the Authorization header Custom header Injects the credential into the configured header, such as x-api-key HTTP Basic Sets the Authorization header using the configured username-and-password value Query parameter Places the credential in a named query parameter AWS SigV4 Signs the request using AWS credentials obtained outside the agent Configuring a Box A Box is configured through two files: box.toml and policy.dw . The first describes the environment the agent runs in: its command, working directory, direct filesystem access, available tools, MCP servers, and credential bindings. The second contains the Dogwood rules. We start with rules that judge one request at a time, then add a rule that depends on what the agent has already done. Consider an on-call agent using the AWS CLI to investigate errors in a production service. We want it to retrieve CloudWatch logs using temporary credentials from an oncall AWS profile on the host. The IAM role behind those credentials limits what they can do, so the agent can read logs or inspect infrastructure, without being allowed to change it. In this example, we are using Strands harness as the harness we are running in the box, though Box itself is harness agnostic. The following snippets illustrate the configuration and policies. For the complete runnable Strands Harness setup, follow the getting-started guide . The box configuration and policy below limit which tools, hosts, and credentials the agent can reach at all: name = "oncall" box_dir = "/Users/me/.box/oncall" policy = "policy.dw" [agent] command = ["strands", "--builtin-tools", "shell"] workspace = "/Users/me/oncall" [agent.filesystem] read = ["/Users/me/oncall"] write = ["/Users/me/oncall"] [tool.aws] command = ["/usr/local/bin/aws"] [tool.aws.env] AWS_REGION = "us-west-2" [egress.aws] destinations = ["*.us-west-2.amazonaws.com"] secret.ref = "aws://oncall" [egress.model] destinations = ["api.anthropic.com"] secret.ref = "env://ANTHROPIC_API_KEY" [egress.slack] destinations = ["slack.com"] secret.ref = "env://SLACK_BOT_TOKEN" secret.inject = "always" The agent can use three things: the Anthropic API for its model, Slack for posting updates, and the AWS CLI declared under [tool.aws]. Commands in box.toml are absolute paths, because the box does not inherit your PATH; use the output of which aws and which strands for your machine. The agent reads and writes its workspace directly, and Box prints that grant when it starts. The agent gets placeholders in place of ANTHROPIC_API_KEY , and the gateway swaps in the real values on the way out. For Slack, secret.inject = "always" injects the bot token into permitted requests without requiring the caller to supply a placeholder. Neither the agent nor the CLI ever sees real AWS credentials: the gateway reads session credentials from the oncall AWS profile on the host and signs each permitted request. Now, in policy.dw, we define which actions we allow the agent to do. Dogwood uses Cedar’s syntax: each rule is a permit or a forbid over a principal, an action, and a resource, with conditions in when . In Box, the principal is always the agent and the resource is fixed, so the action is what a rule matches on, and the details of each request, such as the host or the program path, are in context.input : @id("allow_aws_cli") permit (principal, action == Box::Action::"shell:spawn", resource) when { context.input.program_path == "/usr/local/bin/aws" }; @id("allow_https_to_aws") permit (principal, action == Box::Action::"http:request", resource) when { context.input.host like "*.us-west-2.amazonaws.com" && context.input.port == 443 }; @id("allow_model_api") permit (principal, action == Box::Action::"http:request", resource) when { context.input.host == "api.anthropic.com" && context.input.port == 443 }; The allow_aws_cli rule above permits the AWS CLI execution, the allow_https_to_aws rule permits HTTPS requests to AWS endpoint services in the us-west-2 region, and allow_model_api lets the agent reach its model. A credential binding in box.toml grants no network reach on its own, so each destination needs a permit. With the configuration in place, we can start the agent inside Box: box run --config box.toml Temporal Policies So far, these rules evaluate each operation independently. But Dogwood shines with its temporal operators to make authorization decisions depend on what the agent has already done. These rules apply to any action Box observes, whether an HTTP call to any host, a file read, a shell command, or an MCP tool call. Back to the incident channel from the introduction. An agent that posts on every finding, or retries a post it believes failed, floods the channel. We want to let the agent post, but no more than three times every 10 minutes. @id("allow_slack_post") permit (principal, action == Box::Action::"http:request", resource) when { context.input.host == "slack.com" && context.input.port == 443 && context.input.method == "POST" && context.input.path == "/api/chat.postMessage" }; @id("rate_limit_slack_posts") @description("Slack posts are capped at three every 10 minutes. Wait before you post again.") forbid (principal, action == Box::Action::"http:request", resource) when { context.input.host == "slack.com" && context.input.port == 443 && context.input.method == "POST" && context.input.path == "/api/chat.postMessage" } when temporal { (count for (t: Timepoint). where ( formerly within 10m ( Box::Action::"http:request"::response{ input.host: "slack.com", input.port: 443, input.method: "POST", input.path: "/api/chat.postMessage", output.status: 200 } && tp(t) ) )) >= 3 }; The first rule permits the post. The second refuses another one once the Box has recorded three 200 responses from that endpoint within the last ten minutes. Read the temporal clause from the inside out: the ::response{…} pattern matches a recorded Slack post that returned 200, formerly within 10m keeps only those from the last 10 minutes, and count for (t: Timepoint) counts the moments, bound by tp(t) , at which one occurred. The agent keeps pulling logs and metrics under the AWS permits from earlier while the cap is in effect. Three per ten minutes is an example. The same rule shape expresses any count, cooldown, or ordering constraint. Let’s review an example timeline: Time Agent action Policy decision 10:00 Post an update Allow 10:03 Post an update Allow 10:05 Post an update Allow 10:06 Post a fourth update within 10 minutes Deny 10:07 Retrieve more logs Allow 10:11 Post an update Allow At 10:06, the agent does not get a silent failure. The gateway returns an HTTP 403 whose body reads policy denied this operation , names the rule by its @id , and carries its @description . Every enforcement point reports a refusal with the rule’s identifier and description, so a clear description lets the agent adjust, for example by waiting instead of retrying. The rule counts matching requests that returned HTTP 200, so refusals are excluded. This is because the rule binds on the ::response event of an http:request action. If the rule bound on the ::request event, then it would count attempts, and a refused post would count toward the three. To learn more about the Dogwood Language, you can read the full guide here: https://dogwood-policy.github.io/dogwood/ . However, we understand most people will prefer to use an LLM to write policies, so we are also launching a Policy Authoring agent skill to teach your favorite agent how to write Dogwood policies and install them in your Box. Download the skill from here: authoring-box-policy . What’s next This is another step in AWS’s approach to AI agent safety: giving developers greater control as agents take on more responsibility. By open sourcing the policy language, evaluation engine, and sandbox, we’re making these controls available for developers to inspect, test, and build on. Expand OS support . We’re starting with macOS, and expanding support to other operating systems is one of our next priorities. Our goal is to let developers bring their Box configuration and Dogwood policies to the environments where they already work. Each OS provides different isolation mechanisms, so that work includes making sure those mechanisms preserve the containment guarantees Box relies on. An easier setup . We’re building a CLI that will detect the agent harnesses you already have installed and generate a box.toml and a baseline Dogwood policy with recommended defaults. That will give developers a working starting point they can inspect and adjust, without having to write both files from scratch. Bring Box to deployed agents . We also want developers building custom agents with frameworks such as the Strands Harness SDK, LangChain, or the Claude Agent SDK to package their agent with Box and deploy it to platforms such as Amazon Bedrock AgentCore, ECS, or Kubernetes. The goal is to carry the agent’s policies with it, so the rules governing its actions remain part of the deployment wherever it runs. Increase Dogwood capabilities . We’re investing in Dogwood to express more of the rules developers need to govern their agents. One direction is liveness . Safety rules describe what must not happen; liveness describes what must eventually happen. We want to extend Dogwood so developers can express those obligations and detect when agents fail to meet them, and much more. Get started To get started, check out our GitHub repository at https://github.com/strands-agents/box , and follow the getting started guide to run your first agent in a box. Strands Box is in its early days, and we want to build it with you. Report bugs and request features in GitHub Issues , ask questions on Discord , and read CONTRIBUTING.md before you open a pull request.
Read the full original article:
aws.amazon.com