The following are some recommended best practices to consider when using our most capable models through the Cyber Verification Program. We assess and report on Claude’s alignment with each system card and in our risk reports (see Anthropic’s Transparency Hub) and find these models to be generally aligned, but there are inherently increased risks in using models with reduced safeguards. An agent requested to carry out offensive-security work may take actions beyond the intended scope, and in an unsafeguarded context, this can lead to unintended consequences. These practices are designed to help mitigate these risks when using unsafeguarded models for cyber workflows.
These best practices describe the recommended security infrastructure surrounding the model as well as some guidelines for secure prompting.
We plan to continue iterating on this guidance based on our evolving understanding in consultation with external experts.
For vulnerability discovery and other local work. For jobs that can be run without internet access such as vulnerability finding, patching, or web application testing against a local installation, we recommend:
Sandbox and network isolation. Run testing with isolation inside a virtual machine or dedicated bare-metal host, with egress blocked except for the model API endpoint. The model should reach the API endpoint only through a proxy on localhost, with the API key injected from outside the sandbox. We recommend providing instructions to the model that indicate that it should not use internet access (see Explicit scope-setting below).
We recommend Firecracker on Linux, Hyper-V on Windows, and Hypervisor framework-based VMs on macOS, and to install all tools, packages, and dependencies before the run starts so nothing needs to be fetched mid-run.
Pre-engagement validation. Before live runs:
Whenever using a new sandbox or model, have the model probe the sandbox for vulnerabilities, including via static analysis of the sandbox configuration, then by tasking the model to escape the sandbox from inside of it. Instruct the model to stop and report if it succeeds. Do this under close human supervision, and without internet access.
Whenever making changes to your set-up or running a new task, confirm that all expected tooling is available to the agent, and review for any potentially impossible challenges that merit additional supervision.
Transcript retention and review. Retain model transcripts and sandbox egress logs for at least 30 days. Review the transcripts for any runs where the egress layer logged a denied or unexpected outbound request, such as for probing of the sandbox or network isolation. Please notify your Anthropic account team if you observe out-of-scope behavior. We also recommend automated scanning over a sample of transcripts to confirm the model is staying within the specified scope. Consider using the provided sandbox escape classifier to prioritize transcripts for manual review. For details on how to use this classifier as well as approaches for transcript review, see Agent containment best practices: getting started.
Use Claude Code’s Auto Mode. If you use Claude Code, consider Auto Mode. It reviews each action before it runs, enforces the boundaries you state in the prompt, and by default blocks common containment-escape actions such as requesting instance-metadata credentials, opening tunnels or reverse shells, reading host credentials, or scanning neighboring containers. For authorized testing where some of these are in scope, describe your environment and targets in autoMode.environment rather than disabling the mode. Auto Mode reduces risk but does not guarantee safety.
For pentesting and red teaming. For jobs that require greater network access such as network / AD pentesting and red teaming:
Network scoping. When possible, run the agent harness inside the same VM or bare-metal isolation described above, and enforce scope at the network layer: allow-list egress to the model API endpoint and explicitly in-scope targets. We also recommend pointing engagements at a staging or replica environment that is disconnected from production where one is available.
Explicit scope-setting. Define the intended scope of each engagement in the model’s instructions (in-scope targets, permitted actions, network boundaries, and what is out of scope) before it runs. Phrase constraints as intent (“do not access hosts outside 10.0.3.0/24”) rather than claims about the environment (“you cannot reach the internet”), so the instruction still holds if the environment is misconfigured.
Access proxy/brokered tool use. Where possible, expose access to targeted systems via observable, intermediated systems such as access proxies or tool use instead of allowing the model to craft tooling which has general access to the target. If using a custom harness, we recommend categorizing tools into read-only vs mutating operations, with any non-read only actions subjected to tighter supervision. Consider real-time analysis and denial or escalation to a human of risky commands.
Human supervision. We recommend an engineer supervise each run while it executes, watching tool calls and network activity, with the ability to halt the run immediately. Because agents act at machine speed, live observation should complement, not substitute for, controls that act before an action executes (network allow-lists, brokered tools, and pre-execution review of mutating actions).
Transcript monitoring and retention. For autonomous or long-running harnesses where continuous human attention is impractical, perform continuous monitoring over the model's transcripts, actions, and network activity. Use a monitor (model or rule-based) that is told the intended scope, escalates to a human for review when a run may be exceeding it, and can terminate the engagement. Consider using the provided sandbox escape classifier to prioritize transcripts for manual review. For details on how to use this classifier as well as approaches for transcript review, see Agent containment best practices: getting started. Retain model transcripts for at least 30 days and please notify your Anthropic account team if you observe out-of-scope behavior.
Use Claude Code’s Auto Mode, when possible. We’ve added additional protections that reduce the risk of agents taking misaligned actions and help ensure they stay within the bounds set in the prompt.
High-consequence targets. Do not run autonomous agents against live production systems where an out-of-scope action could endanger the safety or the availability of critical services (for example OT/ICS, medical, or energy systems). This is the existing norm for human-led testing of such systems, and it applies equally here. Test against a replica, testbed, or digital twin, or during a planned outage. Where live access is unavoidable, limit the agent to passive or read-only activity and have a human carry out any state-changing step.