Skip to main content

Agent containment best practices: getting started (private beta)

For customers in our Cyber Verification Program, we’re providing a new sandbox escape classifier in the API to monitor and reduce undesired behavior. This article provides an overview of the guidance available to you, with links to each resource.

This classifier is in private beta.

This guidance is provided as is, for information only. It describes recommendations, not requirements. These practices should reduce risk but may not prevent every out-of-scope action. You remain responsible for choosing, testing, and operating the controls that fit your environment.

Overview of best practices

  • Run every autonomous agent inside a hardware-virtualized sandbox. The reference design uses one microVM per agent (Kata Containers with Firecracker) on an internal network (see Sandboxing).

  • Interactive use, with a human approving every action or supervising the agent in Claude Code's auto mode, carries less risk, but we still recommend a sandbox. Auto mode on its own is a best-effort check, not a substitute for sandboxing. If you run interactive, file-only work without a sandbox, review every write and don’t run anything the agent wrote without reading it.

  • Never mount credential-bearing paths into the agent's environment. Keep the model-API credential out of the agent's environment too: have a separate proxy add it to the agent's API requests (see the credential proxy section in Sandboxing).

  • Run Claude Code agents in auto permission mode as an extra check inside the sandbox (see Auto mode). Its classifier may deny some legitimate security work; that article explains what to expect and how to tune it.

  • Keep agent transcripts and proxy logs for at least 30 days, review them during and after runs (see Offline monitoring of agent transcripts), and tell your Anthropic account team if you see an agent act outside its scope.

Overview of articles

The guidance covers five areas:

  • Sandboxing. Why autonomous agents need strong isolation, how the reference design isolates them, and how to scope and supervise engagements that need network access.

  • Agent permissions: auto mode. Why agents should run under a permission classifier, what it can and cannot be relied on for, and what to tell the classifier.

  • Prompt injection. How text from a target could steer an agent, and how to limit the effect.

  • Offline monitoring of agent transcripts. What to look for in agent transcripts and proxy logs, how to do it, and how long to keep them.

  • Online monitoring: sandbox-escape scores in API responses. For agents you run against the Messages API directly: the per-request score from the sandbox escape classifier, how to request it, and how to read it.

How this guidance maps to the best practices

Best practice

Where it's covered

What you do

Sandbox and network isolation

Run each agent in its own VM with a hardened host, deny egress by default, and keep the model-API credential in a proxy outside the sandbox.

Pre-engagement validation

Review the configuration and run a supervised escape test on each new host, after each sandbox change, and with each new model. Confirm the target builds and the agent has the tools it needs.

Transcript retention and review

Write a transcript for every agent stage, scan transcripts on a schedule, read what the scans flag, and keep transcripts and logs for at least 30 days.

Sandbox escape classifier

Request the score from agents you run against the Messages API directly, and use it to decide which transcripts to read first. It is not available to agents that run in Claude Code, or on Amazon Bedrock or Google Cloud's Agent Platform (Vertex AI).

Auto mode

Run Claude Code agents in auto mode, and describe your environment to the classifier.

Network scoping

Allowlist only the in-scope targets, and prefer a staging or replica environment.

Explicit scope-setting

State the scope in the instructions you give each agent.

Brokered tool use

Put an observable proxy or tool between the agent and the target.

Human supervision

Watch runs that have network access, limit every agent's turns and run time, and know how to stop an agent.

High-consequence targets

Do not point autonomous agents at live systems whose failure could endanger safety or critical services.

Which checks apply depends on how your agent runs:

If your agent runs

Auto mode

API score

Transcript review

in Claude Code

Yes

No. Claude Code does not surface the score

Yes

against the Messages API directly

No. Auto mode is a Claude Code feature

Yes, on the first-party Claude API

Yes

The monitoring in this guidance (transcript review and the API score) looks for attempts to leave the sandbox or to widen network access. It doesn’t detect an agent that stays inside its sandbox but exceeds the engagement on hosts it’s permitted to reach.

Prompt injection isn’t one of the best practices, but it matters whenever an agent reads content that someone outside your organization wrote.

Reporting out-of-scope behavior

If an agent probes its sandbox or network isolation, or acts outside the engagement, notify your Anthropic account team, in addition to whatever your own incident process requires. What to do on a hit describes how to stop the agent and what to collect first.

Did this answer your question?