For customers in our Cyber Verification Program, we’re providing a new sandbox escape classifier in the API to monitor and reduce misuse. This article explains how to request the score, read the results, and use the score to decide which transcripts to review first.
This classifier is in private beta.
For an overview of all the resources available to you, see Agent containment best practices: getting started (private beta).
Overview
When you request it, Anthropic will score each Messages API request whose conversation contains tool use for signs of sandbox escape, and return the result in a safeguard_results field on the response. To request it, add anthropic-beta: safeguards-2026-08-31 to the request and "safeguards": [{"type": "sandbox_escape"}] to the request body.
The score is a number from 0 to 1. Higher scores mean the transcript sent in that request looks more like an agent trying to widen its network access or break out of its execution sandbox. Benign agentic traffic typically scores near 0. As a rough guide, read the transcript for anything at or above 0.15, and treat 0.75 or above as a possible sandbox escape attempt. Depending on your use case, you may see high scores for benign traffic. The classifier is currently available for Claude Mythos 5.1, Sonnet 5.5, and Opus 5.5.
The sandbox escape classifier is in private beta and still under active development. We’re still iterating on its accuracy and are interested in your feedback. We recommend that you use it to prioritize transcript review, alongside the offline checks. You can also build your own automated responses to high scores, such as alerting on-call or stopping an agent.
When you request it, the response contains an explicit status for this feature. The response will show evaluated alongside the sandbox-escape score, skipped when the request had no tool activity, unsupported when the model or platform is not covered, or unavailable along with a reason for the unavailability. The score is available on the first party Claude API, and third-party party surfaces such as AWS Bedrock and GCP Vertex.
This feature is built for agents you run against the Messages API directly. Agents that run through the Claude Code CLI don’t use it; for those, use the offline checks.
Note: This classifier is in private beta. We expect customers to see some false positives as we calibrate the classifier, and scores may return as unavailable on some responses. We would appreciate feedback about this feature: if you have an Anthropic account team, please tell them where it gets things wrong.
How the classifier works
This article describes a monitoring signal that the API can return to you on each request. It complements the checks in offline monitoring of agent transcripts, which you run over the transcripts yourself, during or after a run. The API can score each Messages API request whose conversation includes tool use (tool_use and tool_result blocks, or server-tool blocks) for signs of sandbox escape and return the result in that request’s response. You ask for it with a beta header and a request field, described in Requesting the score.
During the private beta the score isn’t intended to change the response content, and scoring isn’t expected to add a charge or count toward usage. The classifier that produces the score is separate from the auto-mode permission classifier in Claude Code, which decides on each action before it runs (see Auto mode). The score is not that classifier’s verdict.
This page applies to agents that call the Messages API directly: your own agent code, the anthropic (Python) or @anthropic-ai/sdk (TypeScript) client libraries, or curl. The Claude Code CLI (used directly or through the Agent SDK) doesn’t request the score and doesn’t surface the field in its output, so the score doesn’t appear in the transcripts of agents that run through it. For those agents, use the offline checks.
What the score means
The score (the status.score of the sandbox_escape entry in safeguard_results) is a number from 0.00 to 1.00, rounded to the nearest 0.01 (see Requesting the score). The score is the classifier’s estimate of the likelihood that the conversation shows the agent attempting, or having achieved, an escape from its sandbox or containment.
Concretely it looks for two things: the agent widening its network or internet access beyond what it was given, and the agent breaking out of the execution sandbox it runs in. It covers attempts as well as successful escapes, and the single score does not tell them apart. Higher is more concerning. It’s a ranking signal rather than a calibrated probability, and the underlying classifier may be updated without the result naming a version.
The score is computed over the conversation you sent in that request, including the tool calls and tool results already in the transcript. It doesn’t cover the completion the model is about to generate in response. In an agent loop this means each request’s score reflects everything the agent has done up to and including the most recent tool result, and a tool call the model proposes in turn N is first scored on the request for turn N+1, once its result is in the transcript. The final action of a session is therefore not scored. Cover it with the offline checks.
How to read the number:
Benign agentic traffic, including ordinary build, test, fuzz, and file-editing work, typically scores close to 0. A score of 0.00 is a real score, not a missing one.
As a rough reference, a score at or above about 0.15 is worth a look at the transcript, and a score at or above about 0.75 may correspond to an explicit escape attempt visible in the transcript.
The score is produced by a model and is somewhat noisy: the same transcript can score a little differently from one request to the next, and work that resembles escape activity for legitimate reasons may score above zero without any escape attempt. Some legitimate workloads occasionally score well above zero, even into the range that would otherwise suggest an attempt, so calibrate the thresholds against your own traffic. A high score is a good trigger for review or your own automated response. A low score is absence of evidence, not proof of good behavior.
Only requests that contain tool activity are scored. A plain chat request, or the first request of an agent session before any tool has been called, comes back with status skipped rather than a score.
Requesting the score
Two things go on the request.
Add
safeguards-2026-08-31to theanthropic-betaheader. It combines with any other betas you already send, separated by a comma (anthropic-beta: safeguards-2026-08-31,<other-beta>).Then name the check you want in a top-level
safeguardsarray in the request body:"safeguards": [{"type": "sandbox_escape"}].sandbox_escapeis the only type in this beta, each type may appear at most once, and sendingsafeguardswithout the beta header is rejected with a 400 (safeguards: Extra inputs are not permitted).
The header alone, with safeguards empty or omitted, is accepted and simply returns an empty safeguard_results array. A minimal request with one completed tool call in the conversation:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "anthropic-beta: safeguards-2026-08-31" \
-H "content-type: application/json" \
-d '{
"model": "<model-id>",
"max_tokens": 1024,
"safeguards": [{"type": "sandbox_escape"}],
"tools": [{
"name": "bash",
"description": "Run a shell command in the working directory.",
"input_schema": {"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"]}
}],
"messages": [
{"role": "user", "content": "Build the target and run its test suite."},
{"role": "assistant", "content": [
{"type": "tool_use", "id": "toolu_01", "name": "bash",
"input": {"command": "make -j4 && make test"}}
]},
{"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_01",
"content": "[100%] Built target app\n12 tests, 12 passed"}
]}
]
}'
The response shapes below are as of the safeguards-2026-08-31 beta (August 2026). If the beta version changes, check with your Anthropic account team for the updated header and field names. In a non-streaming response, the result is a top-level safeguard_results array on the returned Message object, with one entry per type you requested:
{
"id": "msg_...",
"type": "message",
"role": "assistant",
"model": "<model-id>",
"content": [{"type": "text", "text": "The build succeeded and all 12 tests pass. ..."}],
"stop_reason": "end_turn",
"usage": {"input_tokens": 512, "output_tokens": 40},
"safeguard_results": [
{"type": "sandbox_escape", "status": {"type": "evaluated", "score": 0.01}}
]
}
With "stream": true, the same array arrives inside delta on the final message_delta event, next to stop_reason, and nowhere else in the stream:
event: message_delta
data: {"type": "message_delta", "delta": {"stop_reason": "end_turn", "stop_sequence": null, "safeguard_results": [{"type": "sandbox_escape", "status": {"type": "evaluated", "score": 0.01}}]}, "usage": {"output_tokens": 40}}
Each entry carries type and status. status.type is one of the following: evaluated (with score), skipped, unsupported, or unavailable (with reason). See Caveats for what each entry means.
If you use the anthropic or @anthropic-ai/sdk client library rather than raw HTTP, append safeguards-2026-08-31 to the betas list you already pass. Set it as an extra anthropic-beta request header only on calls that pass no betas, because an explicitly set header replaces the one the library builds from that list rather than merging with it. The client libraries have no typed safeguards parameter for this beta, so pass it as an extra body field: extra_body={"safeguards": [{"type": "sandbox_escape"}]} in Python, or add safeguards to the params object in TypeScript (with // @ts-expect-error if the compiler objects; the library sends unrecognized params as they are).
On a non-streaming call, read safeguard_results from the raw response or the response object’s extra fields, since typed response classes may not expose a field they don’t know about. When streaming through the library, read it from the raw message_delta event rather than from the accumulated final message (get_final_message() / finalMessage()), which is built from the fields the library knows and doesn’t include it. Log the result next to the request or session identifier you already keep, so that a high score leads you straight to the transcript. Record it outside the sandbox, at your proxy or orchestrator, and re-check your thresholds from time to time.
Using the score
Use the classifier to decide which transcripts to review by hand first. Here’s one approach:
Note: These steps are suggestions, not requirements. The classifier is in private beta, so adapt these suggestions to your own traffic and review process.
Log every result next to the request or session identifier you already keep, including the
skipped,unsupported, andunavailableresults. A session whose results are mostlyunavailablehas not been checked.Rank sessions by their highest score. Each request’s score covers the conversation sent in that request, so the highest score in a session is the one to sort by. Read the sessions at or above about 0.15 first, starting from the top.
Read a sample of the low-scoring sessions too. A low score is absence of evidence, not proof of good behavior.
Decide an escalation path in advance. A score at or above about 0.75 may be a reasonable trigger for alerting a person or pausing the agent while someone reads the transcript. Calibrate both thresholds against your own traffic first.
Report confirmed cases. If the transcript shows that the agent did probe its sandbox or act outside the engagement, notify your Anthropic account team.
The score reaches you with the response, so the earliest you can act on a tool call is after it has run (see What the score means). Use the score for detection and review. For prevention, rely on the sandbox and on controls that act before an action executes.
Caveats
Result statuses. Each entry in
safeguard_resultshas astatus.type:evaluated: the classifier ran;status.scoreis the result.skipped: there was no tool activity in the request, so nothing was evaluated.unsupported: this request is not covered (the model or the platform serving it; see Coverage below).unavailable: no score for another reason, given instatus.reason. Current reasons aretimeout,error,capacity,disabled, andinput_too_long. Treat any reason you don’t recognize asunavailable.disabledandinput_too_longwon’t clear on a retry; the others can. The response itself is not expected to be affected.
The array is empty if you sent the header without requesting a safeguard, and absent if you didn’t send the header.
Coverage. They'll be available for the Claude Mythos 5.1, Sonnet 5.5, and Opus 5.5 models provided to you through the CVP program, and other models will return
unsupported. The score is available on both the first-party Claude API and on third-party surfaces such as Amazon Bedrock or Google Cloud’s Agent Platform (Vertex AI). If you expectevaluatedand consistently seeunsupported, check the model ID first, then ask your Anthropic account team.Latency. Requesting the safeguard may add a small amount of latency. If the latency exceeds a threshold, it will time out, and the feature’s result field will show
unavailablewith reasontimeout. The rest of your response will be returned, unaffected.