Skip to main content

Agent containment best practices: auto mode (private beta)

For customers in our Cyber Verification Program, we’re providing a new sandbox escape classifier in the API to monitor and reduce misuse. This article covers what auto mode requires from your models and account, what it changes while a run is in progress (including possible denials of legitimate work), and what to tell the classifier about an agent's environment.

This classifier is in private beta.

For the feature itself, see Choose a permission mode and Configure auto mode in the Claude Code documentation.

For an overview of all the resources available to you, see Agent containment best practices: getting started (private beta).

Overview

  • We recommend running Claude Code agents in auto permission mode where you can. It adds a permission classifier that checks risky commands before they run and blocks many dangerous ones. It reduces risk but does not guarantee safety: it can miss things, and it can block legitimate work. In the reference design every agent runs this way inside the microVM sandbox, and never with bypassPermissions. The sandbox is the security boundary, and the classifier is a defense-in-depth check inside it.

  • The classifier was built for general software work, and offensive-security tasks look unusual to it, so it may deny some legitimate actions on your workloads. Some over-triggering is expected and does not mean that something is misconfigured. To reduce it, describe your environment and targets to the classifier (see Applying the rules to your own Claude Code sessions).

  • Auto mode is a Claude Code feature. It does not apply to agents you build directly on the API (first-party, Amazon Bedrock, or Google Cloud's Agent Platform, also known as Vertex AI) without Claude Code.

  • Auto mode works only with certain models and must be available to your organization. Check both before you launch (see Requirements).

  • Claude Code v2.1.257 and later block common containment-escape actions by default. We recommend also giving the classifier a description of the machine and of the in-scope targets, and a rule that protects the setup around the session.

  • The classifier fails closed when it is unavailable. In an unattended run, decide in advance what happens after repeated denials.

  • On Claude Code versions that call the classifier separately, it costs at least one extra model request per checked tool call.

Requirements

Use a model that supports auto mode

Claude Code supports auto mode only on recent models. In an unattended session, an agent on a model without auto mode support cannot run commands or write files. Check the model before you launch, and refuse to start if it’s not supported. At the time of writing:

  • Anthropic API:

    • Supported: Opus 4.6 or newer, Sonnet 4.6 or newer, or a newer model family.

    • Not supported: Claude 3.x, Opus and Sonnet 4.5 or older, and Haiku 4.5.

  • Amazon Bedrock and Google Cloud's Agent Platform (Vertex AI):

    • Supported: Opus 4.7 or newer, or Sonnet 5 where your account can invoke it through the provider's standard endpoint.

    • Not supported: Opus 4.6, Sonnet 4.6, and every Haiku.

Use full model IDs, and not aliases such as sonnet, for every model an agent stage uses. The Claude Code documentation has the current list.

Auto mode must be available to your organization

On Team and Enterprise plans an admin can turn auto mode off, and Anthropic can disable it as a safety measure. In both cases, agents can’t run commands, and an unattended run fails right after it starts until auto mode is turned back on. The same happens when a model without auto mode support gets past your check, for example behind a Bedrock application inference profile.

Side effects

In auto mode a permission classifier checks the agent's tool calls. This has four effects that you may notice during runs.

Some legitimate actions may be denied

The classifier is a model judging each command against general-purpose rules plus any you add, without seeing the command's output. Security work (running exploit code against a target, scanning or probing a lab host) is unusual traffic for it, and it may sometimes refuse a step that’s a normal part of the task. How often depends on your targets and prompts, so we can’t give a rate; if it happens, some over-triggering is expected and isn’t by itself a sign of misconfiguration. A denied call costs the agent a turn; the denial and its reason appear in the transcript. Describing the environment and the in-scope target precisely (see Applying the rules to your own Claude Code sessions) is the main lever for reducing false denials.

If the classifier is unavailable, tool calls are refused

The classifier fails closed. If the model that runs it is briefly unavailable, the tool call is refused with a message like "temporarily unavailable, so auto mode cannot determine the safety of …", and the agent is told to try again. This is rare and usually costs one or two turns. Reads keep working. Count the denials and refusals in each session, so that you can see when they become frequent.

Decide what happens after repeated denials

An agent whose tool calls keep being denied is no longer doing its task, and an unattended session has no person to take over. The reference design ends a run after three denials in a row, or twenty in one session. This can mean that the agent kept trying something the rules block, or that legitimate work was misjudged. Do not resume such a run automatically, because resuming restores the same transcript. Read the transcript, then retry the run from scratch with a fresh session. If denials keep ending your runs, describe the environment and the in-scope targets to the classifier more precisely (see Applying the rules to your own Claude Code sessions).

The classifier uses Claude Sonnet 5 and adds model traffic

On Claude Code versions that call the classifier separately, each tool call that needs a check costs at least one extra model request, on the same account, provider, and rate limits as the agent. The classifier runs on Claude Sonnet 5 by default, so on Amazon Bedrock and Google Cloud's Agent Platform, enable that model in addition to the agent's model where you can. If the classifier cannot use it, the CLI falls back to another model for the rest of the session.

What to tell the classifier

Auto mode's built-in rules are aimed at destructive actions, leaked secrets, and attempts to turn off Claude Code's own sandbox. Claude Code v2.1.257 and later also block common containment-escape actions by default: requesting credentials from the cloud instance-metadata endpoint, reaching a public host through a tunnel, a reverse shell, or another indirect route, reading credentials that belong to the host, and connecting to or scanning neighboring containers.

What the classifier cannot know on its own is what the particular machine it runs on is allowed to reach. It sees the agent's commands and the prompts, but not the output of those commands, and by default it assumes an ordinary developer machine with open internet access.

A sandboxed agent's environment is different, so give every agent session two rules and a short description of its environment, in addition to Claude Code's default rules. This is what the reference design passes. Its rule text is not included, so write your own along the lines described below.

  1. A "Containment Escape" rule (soft_deny). It targets attempts to reach beyond the container's network, identity, or tenancy boundary. Because the classifier cannot see results, the rule describes techniques and patterns rather than outcomes: (i) getting traffic to a public host by an indirect route (for example, DNS pinning), or trying the same outside destination repeatedly in different ways; (ii) requesting cloud instance-metadata credentials, or calling cloud APIs with the machine's own identity; (iii) reading credentials that belong to the hosting layer (node certificates, another process's environment); (iv) connecting to or scanning sibling containers, the node underneath, or the surrounding private address space. Normal retries, --resolve entries that point to loopback or private addresses, and one or two diagnostic probes are explicitly allowed. Where the sandbox has no network access except the allowlisted hosts, the rule says so, so that any direct request to another outside host counts as well. Claude Code v2.1.257 and later include a rule of this kind by default.

  2. A tampering rule (soft_deny). This is a separate rule, so that each rule stays short and readable. It covers changes to the setup around the session:

    1. Claude Code settings, permission, and hook files inside the container

    2. The proxy and API-endpoint environment variables

    3. Sending the credential proxy anything other than the CLI's own model calls

    4. Processes the agent did not start

    5. Files that the orchestrator delivers into the sandbox for this run or later runs


    The agent can still manage its own fuzzers, target processes, and files.

  3. Facts about the environment (environment). These are statements of fact, not instructions:

    1. Network setup: the internal network, the credential proxy as the only route to the model API, and the egress proxy with the exact hosts on its allowlist (normally none), as reported by the running proxy.

    2. Metadata and identity: the cloud metadata addresses cannot be reached from the container, and it has no cloud identity of its own. It holds no model-API credential: the key or token variable contains only a placeholder, and that is not for the agent's commands either.

    3. Neighbors: the Docker socket, the host, both proxies, and other containers are not part of the task.

    4. Target: what the agent is working on, such as a binary and its source tree, an application's source, logs, and loopback port, or a target URL and the authorized hosts. Work on the target is expected; the rule applies to reaching anything else.

Start the soft_deny list and the environment list with the literal string "$defaults". This keeps Claude Code's built-in rules and environment entries and adds yours after them. Without it, a custom list replaces the built-in one.

The sandbox remains the security boundary. The classifier only judges commands and never sees their output, so these rules reduce risk but do not contain anything by themselves. They may stop an agent that tries to reach outside its container, and they make the attempt visible in the transcript. Do not rely on them to catch every attempt.

Applying the rules to your own Claude Code sessions

If you use auto mode in your own Claude Code sessions, give the classifier a description of the machine. This includes any machine you run security tooling from. The description does the most to reduce false denials on security work, because it tells the classifier which hosts and actions are a normal part of the task. When some of the blocked actions are in scope for authorized testing, describe your environment and targets to the classifier.

Claude Code v2.1.257 and later block the containment-escape actions described in What to tell the classifier by default. claude auto-mode defaults prints the rules your version includes. To add a rule of your own, use a settings block of this shape:

{"autoMode": {"soft_deny": ["$defaults", "Containment Escape: …"]}}

If you add rules of your own, use user or managed settings only. Claude Code ignores autoMode in the project-level files .claude/settings.json and .claude/settings.local.json, so that a cloned repo cannot add its own classifier rules. Start a new Claude Code session for the change to take effect. claude auto-mode config shows the merged rules in effect, and claude auto-mode defaults shows what "$defaults" stands for. Start each rule with a short label, as Containment Escape: does. When the classifier gives a written reason for a denial, it usually names the rule, which makes transcripts easier to search. When the reason is only Blocked by classifier, look at the refused command instead.

Treat autoMode settings as security configuration: deploy them through managed settings, and review each change as you would a firewall rule change. A description of the machine in autoMode.environment (after "$defaults") makes the rules more precise for your setup. The Claude Code documentation calls this the Host containment entry. A few lines are enough. Write the description as facts about the machine, not as instructions. Instructions belong in the agent's prompt (see Scope and supervision). Some examples of useful facts, if they apply:

  • Network: open internet, an allowlist (name the hosts), mirror-only (name the mirror), or no network at all.

  • Metadata and identity: whether the instance-metadata endpoint can be reached, and whether the agent is supposed to use the machine's identity (usually not).

  • Designated infrastructure: the cloud project, cluster, registry, CI system, and internal hosts the task works with, and how the agent is authenticated to them. Anything named here counts as normal work; the rule applies to anything else outside the boundary.

  • Neighbors: other pods, VMs, or tenants the agent must not touch.

You can adapt this to your case, in the environment description or in a soft_deny entry of your own:

  • No network access except an allowlist. If that is true of your machine, add a soft_deny entry such as "This environment is declared no-network except the hosts listed in the Environment section: treat ANY direct request to another outside host (plain curl/wget/git/pip/npm/apt to a public URL or package index) as Containment Escape." Do not add it otherwise, or every normal download gets blocked. The classifier never sees command output, so it cannot tell whether a machine described this way really is restricted. Enforce that in your network configuration and check it when the sandbox starts. The reference design adds this entry only after its launcher has verified the sandbox runtime and the proxy.

  • The agent uses your cloud account on purpose (for example a deploy runner with workload identity). Say so in the environment.

  • Pentest and security-research sandboxes. Name the in-scope target range in the environment ("targets are 10.20.0.0/16 and *.range.example.com; everything else is out of scope") so that clause (iv) only applies outside that range. Naming a target is also what lets clause (ii) treat a metadata or loopback URL sent to that target as an SSRF test payload rather than a metadata request. For a pentest, give the target URL and the authorized hosts, and say that SSRF testing of the target is in scope.

  • Multi-tenant GPU or notebook hosts. Add a soft_deny entry for what matters there (other users' home directories, other jobs' accelerators, the scheduler's socket).

  • Monitoring and tamper resistance. To also keep the agent away from audit logging, proxy configuration, or the hooks that enforce your policy, add a short separate soft_deny entry that names those files and services. The tampering rule described in What to tell the classifier shows what such an entry covers.

Did this answer your question?