For customers in our Cyber Verification Program, we’re providing a new sandbox escape classifier in the API to monitor and reduce undesired behavior. This article provides an overview of the guidance available to you, with links to each resource.
This classifier is in private beta.
This guidance is provided as is, for information only. It describes recommendations, not requirements. These practices should reduce risk but may not prevent every out-of-scope action. You remain responsible for choosing, testing, and operating the controls that fit your environment.
Overview of best practices
Run every autonomous agent inside a hardware-virtualized sandbox. The reference design uses one microVM per agent (Kata Containers with Firecracker) on an internal network (see Sandboxing).
Interactive use, with a human approving every action or supervising the agent in Claude Code's auto mode, carries less risk, but we still recommend a sandbox. Auto mode on its own is a best-effort check, not a substitute for sandboxing. If you run interactive, file-only work without a sandbox, review every write and don’t run anything the agent wrote without reading it.
Never mount credential-bearing paths into the agent's environment. Keep the model-API credential out of the agent's environment too: have a separate proxy add it to the agent's API requests (see the credential proxy section in Sandboxing).
Run Claude Code agents in auto permission mode as an extra check inside the sandbox (see Auto mode). Its classifier may deny some legitimate security work; that article explains what to expect and how to tune it.
Keep agent transcripts and proxy logs for at least 30 days, review them during and after runs (see Offline monitoring of agent transcripts), and tell your Anthropic account team if you see an agent act outside its scope.
Overview of articles
The guidance covers five areas:
Sandboxing. Why autonomous agents need strong isolation, how the reference design isolates them, and how to scope and supervise engagements that need network access.
Agent permissions: auto mode. Why agents should run under a permission classifier, what it can and cannot be relied on for, and what to tell the classifier.
Prompt injection. How text from a target could steer an agent, and how to limit the effect.
Offline monitoring of agent transcripts. What to look for in agent transcripts and proxy logs, how to do it, and how long to keep them.
Online monitoring: sandbox-escape scores in API responses. For agents you run against the Messages API directly: the per-request score from the sandbox escape classifier, how to request it, and how to read it.
How this guidance maps to the best practices
Best practice | Where it's covered | What you do |
Sandbox and network isolation | Run each agent in its own VM with a hardened host, deny egress by default, and keep the model-API credential in a proxy outside the sandbox. | |
Pre-engagement validation | Review the configuration and run a supervised escape test on each new host, after each sandbox change, and with each new model. Confirm the target builds and the agent has the tools it needs. | |
Transcript retention and review | Write a transcript for every agent stage, scan transcripts on a schedule, read what the scans flag, and keep transcripts and logs for at least 30 days. | |
Sandbox escape classifier | Request the score from agents you run against the Messages API directly, and use it to decide which transcripts to read first. It is not available to agents that run in Claude Code, or on Amazon Bedrock or Google Cloud's Agent Platform (Vertex AI). | |
Auto mode | Run Claude Code agents in auto mode, and describe your environment to the classifier. | |
Network scoping | Allowlist only the in-scope targets, and prefer a staging or replica environment. | |
Explicit scope-setting | State the scope in the instructions you give each agent. | |
Brokered tool use | Put an observable proxy or tool between the agent and the target. | |
Human supervision | Watch runs that have network access, limit every agent's turns and run time, and know how to stop an agent. | |
High-consequence targets | Do not point autonomous agents at live systems whose failure could endanger safety or critical services. |
Which checks apply depends on how your agent runs:
If your agent runs | Auto mode | API score | Transcript review |
in Claude Code | Yes | No. Claude Code does not surface the score | Yes |
against the Messages API directly | No. Auto mode is a Claude Code feature | Yes, on the first-party Claude API | Yes |
The monitoring in this guidance (transcript review and the API score) looks for attempts to leave the sandbox or to widen network access. It doesn’t detect an agent that stays inside its sandbox but exceeds the engagement on hosts it’s permitted to reach.
Prompt injection isn’t one of the best practices, but it matters whenever an agent reads content that someone outside your organization wrote.
Reporting out-of-scope behavior
If an agent probes its sandbox or network isolation, or acts outside the engagement, notify your Anthropic account team, in addition to whatever your own incident process requires. What to do on a hit describes how to stop the agent and what to collect first.