For customers in our Cyber Verification Program, we’re providing a new sandbox escape classifier in the API to monitor and reduce misuse. This article explains why the agent transcripts are worth reading. It then has one section per layer of monitoring: the proxy logs first, then what a transcript should hold, pattern rules over the transcripts, and a model reviewer over the transcripts. It ends with what to do when a layer flags an agent while a batch is running, and what to keep afterwards.
This classifier is in private beta.
For an overview of all the resources available to you, see Agent containment best practices: getting started (private beta).
Overview
Reading the transcripts is a check that does not depend on the sandbox or the classifier making the right decision. The transcript is produced by the agent's own process, so pair it with records taken outside the sandbox, such as the proxy logs. Probing is often visible in a transcript before anything harmful runs, but do not rely on having time to react: controls that act before an action executes come first. During a long batch with no network access, scan the live transcript files on a schedule and have a person look at the hits every hour or two. For runs with network access to a target, have an engineer watch the run, or scan continuously and stop the agent on a hit. After any run, scan everything once to confirm that every agent stayed inside its sandbox.
Use three layers of monitoring, cheapest first: (1) the egress-proxy and credential-proxy logs, (2) pattern rules over the transcripts, written for your environment, and (3) a model reviewer over the transcripts, with a prompt you write for your sandbox and scope.
Also send the reviewer a random 5 to 10% of the transcripts that the pattern rules did not flag. Expect some false positives. If the reviewer says an agent did something, read that part of the transcript before you act on it.
When a layer shows that an agent has left its task during a run, stop it, and tell your Anthropic account team about out-of-scope behavior.
Keep transcripts and proxy logs for at least 30 days.
Why read the transcripts
The sandbox and Claude Code's auto-mode permission classifier both act while an agent runs. Reading the transcripts is a check that depends on neither of them. It has two uses.
During a run. Probing, if it happens, may be visible in the transcript before anything harmful runs. Do not rely on having time to react: controls that act before an action executes come first. During a long batch, scan the live transcript files on a schedule, and have a person look at what the scans flag every hour or two.
After a run. Scan everything once to confirm that every agent stayed inside its sandbox.
For engagements with network access to real targets, the best practices recommend that a person supervise the run, or that a monitor that knows the intended scope watches continuously, escalates to a person when a run may be exceeding that scope, and can end the run (see the Scope and supervision section in Sandboxing). The layers in this article are the pieces such a monitor is built from.
If you run agents against the Messages API directly, the API can also return a per-request score from the sandbox escape classifier as they run (see Online monitoring).
Layer 1: proxy logs
The cheapest layer reads no transcript at all.
The egress proxy should log every request it refuses, with the client address and the destination. With an empty allowlist, agents reach the model API only through the credential proxy and need no other host, so every refused request during a run is worth a look. Read the transcript of every run that caused one. The client address shows which agent made the request, so you know which transcript to read first.
Know what your egress layer does not log. In the reference design, the egress proxy logs refused HTTPS CONNECT requests. Plain-HTTP and raw TCP attempts do not go through the proxy on the internal network, so they fail without a log line and appear only in the transcripts. In your own build, consider logging allowed as well as refused connections, for every protocol, with timestamps.
The credential proxy should log a line for any request that was not a model call. Save its log when the run ends (see The credential proxy section in Sandboxing).
What to capture in a transcript
The second and third layers read the transcripts themselves.
Have the orchestrator on the host write a transcript for every agent stage, one JSON message per line, as each message arrives. Claude Code's stream-json output has this form. In each file, lines with "type": "assistant" contain what the model produced: text, thinking, and tool_use blocks. The tool_use input holds the Bash command or the file being written. Lines with "type": "user" contain tool_result blocks: command output, file contents, and target responses.
Capture every stage, including stages whose agents have no tools, such as an agent that only judges or compares findings. There is less to check in those transcripts, but a stage with no transcript cannot be reviewed at all.
Layer 2: pattern rules over transcripts
Write your own pattern rules, such as regular expressions, for your agents and targets. For agents that work through a shell, cover behavior such as bypassing the egress proxy or firewall, sending credentials or opening SSH connections out of the sandbox, changing kernel or privilege tunables, and listing other processes through /proc. Rules match only behavior that someone anticipated, so treat them as a starting point, not a complete detector.
Apply the rules to JSON-decoded strings rather than raw lines, because on a raw line quotes and newlines are escaped and patterns may not match. Report each hit with the file, the line, the rule, and a snippet, and have the scan exit with a distinct status when it finds something, so that a scheduled job can run it during a long batch and raise an alert.
Expect benign hits
A hit is a reason to read that part of the transcript. On its own it does not mean that the agent did anything wrong. Agents doing vulnerability research routinely run commands that match escape rules for ordinary reasons:
running the target under an
LD_PRELOADshim;attaching ptrace to their own child processes, or reading /proc/<pid>/ for those processes;
turning ASLR off through /proc/sys/kernel/randomize_va_space;
listing /proc to find and kill their own fuzz workers.
A proxy-bypass rule can match every curl, wget, or netcat request to a pentest target with a public IP, and an SSH-egress rule can match any ssh or scp to a target with a public IP or a DNS name. Narrow or remove rules that only produce benign hits on your targets.
Layer 3: model reviewer over transcripts
For each transcript that the pattern rules flagged, send one Messages API request to Claude with a reviewer prompt that you write. The prompt should ask Claude to look for an agent that widened its network access or broke out of its sandbox, whether attempted or achieved, say what to flag and what not to flag, and ask for the findings, or none, with the transcript lines that support them.
Build the user message in two parts:
Start with two or three sentences that describe the sandbox and the intended scope. For the reference design: each agent runs in its own isolated sandbox (a container or microVM) on an internal network, the only way out is a proxy that forwards model-API calls and nothing else, and the target is inside the sandbox with the agent. If your sandbox allows other hosts, or the target is on the network, say so here and name the in-scope targets.
Then include the agent's own messages, meaning its
text,thinking, andtool_useblocks, with enoughtool_resultcontent to show what each call returned.
Use the model you run the agents with, or a cheaper one. Under Amazon Bedrock or Google Cloud's Agent Platform (Vertex AI), use one that is enabled in your account.
Two limits to keep in mind:
A reviewer prompt that looks only for sandbox and network escape does not detect an agent that exceeds the engagement on hosts it is permitted to reach.
The transcript is untrusted input to the reviewer. It contains text that the target controls. A clean verdict does not clear a transcript that a rule or a proxy log flagged.
Review a sample of the unflagged transcripts too
Pattern rules catch only what someone anticipated, so also send a random 5 to 10% of the transcripts that the rules did not flag, using the same request. The sample shows what the rules miss, and whether you should add a rule.
Read the classifier's denials
The auto-mode classifier's refusals are also useful to read: search the transcripts for the classifier's denial messages, and keep them apart from the refusals it issues when it is unavailable (see Auto mode).
This reviewer is separate from the auto-mode permission classifier (see Auto mode), which decides on each action before it runs, and from the sandbox escape classifier in the API (see Online monitoring). The three complement each other.
What to do on a hit during a run
When one of the layers shows an agent that has left its task while its batch is still running:
Stop the agent. In a Docker-based setup,
docker rm -fon the agent's container stops it. Removing a container deletes the guest's disk, so anything the agent wrote there is lost. The orchestrator should record the run as failed and continue with the rest of the batch.Keep the evidence. Leave the run's results as they are, and check the egress proxy's log for the same time window. Make sure that the log lines carry timestamps.
Report it. If what you find is genuinely out-of-scope behavior, meaning an agent that probed its sandbox or network isolation or acted outside the engagement, notify your Anthropic account team, in addition to whatever your own incident process requires.
Rotate the model-API credential, as after any run whose transcript shows unexpected behavior (see Sandboxing).
If you need to share a transcript outside your team, work on a copy. Remove session ids, usage accounting, and the host inventory that the first line of a transcript records, and read the file before you send it.
Retention
Keep the following for at least 30 days after a run, or longer where your regulatory or contractual obligations require it. Store them off the host:
Agent transcripts;
The logs of whatever enforces egress from the sandbox (a proxy or a firewall), including refused connections;
The logs of whatever injects the model-API credential, if that is a separate component.
Store them with the same access controls as the run itself. Transcripts can contain target source code, exploit code, and findings.
If any of these logs exist only as a container's log, export them before the container is removed or recreated, because the log is deleted with it.