For customers in our Cyber Verification Program, we’re providing a new sandbox escape classifier in the API to monitor and reduce misuse. This article explains how text from a target could steer an agent, and how to limit the effect.
This classifier is in private beta.
For an overview of all the resources available to you, see Agent containment best practices: getting started (private beta).
Overview
A prompt injection is an instruction that an adversary writes into text the model reads, and that the model may then follow as if it came from the user. The risk grows with the amount of untrusted content the model reads and with the permissions it has.
Limit what the model can read, limit what it can do, use auto mode in Claude Code, and review output before acting on it.
For vulnerability research on code you wrote, the risk is plausibly low. For penetration tests and red-team engagements, treat everything the agent reads from a target as untrusted: pages, banners, logs, stored user data, and file contents can all carry instructions written to mislead it.
Limit what an injection could achieve in three ways: give the agent no route to an attacker's host and no credential, carry little of what an agent writes into a later run, and run every agent under the auto-mode classifier.
This article first lists general guidelines that apply to any agent. It then describes how the reference design limits the effect of an injection that succeeds.
General guidelines
A prompt injection is an instruction that an adversary writes into some text the model reads (a file, a web page, a program's output), and that the model may then follow as if it came from the user. The risk increases as the model is exposed to untrusted third-party content and given permissions on sensitive data. Four practices reduce it:
Limit what the model can read. Don't add skills, plugins, or MCP servers from the internet unless you trust their source; remove tools the run doesn't need; narrow the inputs that remain to trusted sources if possible (for example, internal email only, or only websites the trusted task requires).
Limit what the model can do. Grant only the permissions the job needs and run the agents in a strong sandbox (see Sandboxing). An injection can only use capabilities the model has.
Use auto mode in Claude Code. Claude models are designed to resist prompt injection; see the system card for the model you use. Claude Code's auto mode uses a recent model to check risky actions before they run, such as shell commands and writes outside the working directory. Reads and edits inside the working directory are not sent to the classifier. The check makes it harder for injected text to turn into a harmful action, though it will not catch everything. Use it where possible. Learn more about auto mode in Claude Code.
Review output before acting on it. Some risk remains whatever you do.
Limiting what an injection can achieve
When you run an agent against code you wrote, the text it reads is perhaps less likely to have been written by an attacker. That changes as soon as the agent reads content that came from outside: third-party source code, request logs, user data that a web target stores, or anything a penetration-test target returns. Treat that content as untrusted, and keep the General guidelines in mind for those runs.
The reference design is built to mitigate part of the residual risk from what an injection can achieve:
Make it hard to exfiltrate data. With the default configuration, an agent's container can reach only (1) the credential proxy, which forwards model calls and nothing else, and (2) the target itself. A target can be inside the agent's own container, or on the network. Every other destination is refused unless the operator added it to the egress allowlist, so a successful injection has no direct route to send data to an attacker's host. An allowlisted host or a networked target is a possible route. There’s also no credential in the container to exfiltrate (see Sandboxing).
Make it hard to propagate attacks to future runs. Let little of what an agent writes be directly carried into a future run against the same target. The reference design has three narrow exceptions, where limited agent-written content crosses a session boundary. The first and third pass through human review, and the second stays inside one batch.
Lessons learned. An agent may propose lessons for later runs. A person sees each proposed lesson next to the transcript excerpt it came from, and accepts or rejects each one.
Shared files within one batch. Later agents read some of what earlier agents produced, such as short crash excerpts and the crash output that a report or patch is based on. None of this is carried into another batch.
Patches. A patch is code written by an agent that has read the target's source and crash output, so it is the output an injection could plausibly try to corrupt. Apply a patch only inside test containers and never to your repository. Treat the diff as a proposal that a person reviews and approves before it goes anywhere.
Run Claude Code agents in auto mode. Every agent in the reference design runs under the auto-mode classifier (see Auto mode), which judges the actions it checks against the task stated in the prompt. An injected instruction often shows up as an action that does not fit the task. The classifier doesn’t see what the agent read.