Skip to main content

Agent containment best practices: sandboxing (private beta)

For customers in our Cyber Verification Program, we’re providing a new sandbox escape classifier in the API to monitor and reduce misuse. This article explains why autonomous agents need strong isolation, how the reference design isolates them, and how to scope and supervise engagements that need network access.

This classifier is in private beta.

For an overview of all the resources available to you, see Agent containment best practices: getting started (private beta).

Overview

  • In an autonomous run, no human approves the agent's tool calls, and the agent may execute target code. We recommend running autonomous agents in a strong sandbox that puts a hardware-virtualized kernel between the host and both the agent and the target code. Do not use bare Docker/runc, and never use --privileged or host networking.

  • The reference design runs every agent in its own microVM (Kata Containers with Firecracker) on an internal network. It needs a Linux host with KVM, which limits the hosts that can run it. Kata has deprecated the runtime that Firecracker depends on.

  • Never mount paths that hold credentials (such as ~/.aws, ~/.ssh, or .env) into the agent's environment. Keep the model-API credential out of the agent's environment too: have a separate credential proxy hold it and add it to each model request (see The credential proxy).

  • Deny egress by default. In the reference design, model calls go through the credential proxy, and the egress proxy refuses every other destination unless you allowlist it.

  • Test your sandbox on each host before you rely on it, and again whenever the sandbox or the model changes.

  • For engagements that need network access, state the scope in the agent's instructions, enforce it in the network, and keep autonomous agents away from live high-consequence systems (see Scope and supervision).

Sandboxing guidance

Properties to aim for

If you build your own sandbox, these are the properties the reference design provides. The rest of this article describes one way to get them.

  1. Each agent has its own guest kernel, which it does not share with the host.

  2. The agent's file and shell tools see only the guest's filesystem. No host directory is shared into the guest.

  3. The model-API credential is not in the agent's environment or on its disk. A proxy outside the sandbox adds it to each request.

  4. The agent has no internet access. Egress is enforced outside the guest, and every destination is refused unless it is on an allowlist.

  5. The cloud metadata service cannot be reached from the agent.

  6. Every agent has a turn limit, and a time limit if you set one.

  7. No privileged mode, added capabilities, device passthrough, or host networking.

  8. The orchestrator runs on the trusted host and writes the transcripts there.

General guidelines

A strong sandbox matters most for autonomous runs, where agents execute target code and no human approves each action

Run autonomous agents in strong sandboxes, and think through what side effects a process in that sandbox could still cause. Frontier models are increasingly good at finding creative paths around restrictions: the same property that makes them effective vulnerability hunters means they may take unexpected actions against their own execution environment. This is not hypothetical. Anthropic has published examples of models working around weak constraints in order to complete a task (see How we contain Claude across products).

Concretely, do not run autonomous vulnerability-finding agents in bare Docker/runc, and especially not with --privileged or host networking. Standard containers share the host kernel, so a kernel exploit inside the container is a host compromise. We recommend putting a hardware-virtualized kernel between target code and the host, which means running the agent in a virtual machine. Where that is not possible, use a dedicated bare-metal host that holds nothing else. If you build your own sandbox, we recommend Firecracker on Linux, Hyper-V on Windows, and a VM based on the Hypervisor framework on macOS. The reference design uses Firecracker.

Block all egress from the sandbox. The agent reaches the model API only through a proxy that adds the credential and runs outside the sandbox, so the sandbox needs no direct route to the API host. Install every tool, package, and dependency before the run starts, so that nothing needs to be fetched during it.

Interactive use, with a human in the loop, generally carries less risk, but we still recommend a sandbox for it. If you drive an agent interactively from Claude Code on a laptop, either review every tool use (manual mode), or rely on the auto-mode permission classifier and have a human approve every action that reaches outside the repo. Auto mode removes routine permission prompts: it approves reads and working-directory edits automatically and sends everything else to a background classifier that aims to block destructive, irreversible, or off-task actions. It is a best-effort check. It can miss things, and on security work it may also refuse some legitimate steps. Auto mode describes how it works, what to expect from it, and how to configure it for your environment.

Never mount credential-bearing paths such as ~/.aws, ~/.ssh, or .env into the agent's environment. The same goes for the credential that the agent's own model calls use. Keep it outside the sandbox, and have a proxy that the agent cannot read add it to each request (see The credential proxy). Do not connect agents to MCP servers or tools with write access to external state such as email, cloud storage, or production infrastructure.

Split each run into a setup phase and an attack phase with different network policies

This is one pattern that puts the guidelines above into practice. The setup phase has outbound internet access and a human in the loop who approves every tool call. In it the agent pulls dependencies, builds the target, and stands up its sandbox from a spec document. The attack phase has no general internet access. All egress goes through an allowlist proxy that permits only the hosts named in the engagement, so the proxy also enforces scope. Model calls go through the separate credential proxy. The agent can then probe the target unattended. The proxy contains the agent's own traffic. It does not contain traffic that a networked target sends on the agent's behalf.

Bake the dependencies the agent needs repeatedly into the setup image, so that attack-phase runs need no internet access at all. Scope credentials per target, so that an agent working on one target cannot use them against another. Keep Claude Code's auto mode on inside the sandbox during the attack phase and describe the sandbox to its classifier. See Auto mode.

Bound each run, and know your off-switch

Give every agent an explicit budget, so that a run that exceeds its intended testing stops on its own and not when someone notices. Enforce a hard turn limit on every agent. When an agent exhausts it, end the run and do not grant more turns automatically. Relaunching with a higher limit is the point at which a person decides to continue. Bound runs in time as well. End a session that exceeds its time limit the same way: the run is final and never resumed, and the agent process is stopped, even when the limit is reached in the middle of a long-running command. If a limit is set to a value that cannot be read, refuse to launch instead of running uncapped. Set the limits deliberately for the engagement and do not accept a generous default. Pair them with the monitoring cadence in offline monitoring of agent transcripts, so that a long batch is looked at every hour or two and not only at the end. Know how to stop a single agent without stopping the batch. In a Docker-based setup that is docker rm -f <agent-container>. The orchestrator should record that run as failed and continue with the rest of the batch.

For more guidance, read two Anthropic resources. Securely deploying AI agents covers isolation options, credential proxying, and filesystem hardening. The engineering retrospective How we contain Claude across products covers what held and what didn't when these same mechanisms ran in production.

Scope and supervision

The sandbox limits what an agent can reach. For pentesting, red teaming, and other engagements that need network access, the practices below limit what the agent is asked and allowed to do. They depend on your targets and your team, so tooling cannot put most of them in place for you.

State the scope in the agent's instructions

Before a run, tell the agent which targets are in scope, which actions are permitted, where the network boundary is, and what is out of scope. Phrase each constraint as intent ("do not access hosts outside 10.0.3.0/24") and not as a claim about the environment ("you cannot reach the internet"), so that the instruction still holds if the environment is misconfigured. Do this for sandboxed local work too: tell the agent not to use internet access, even though the sandbox blocks it. The description you give auto mode's classifier is a different thing. It states facts about the machine (see Auto mode).

Enforce the same scope in the network

Where you can, run the agent inside the same isolation described above, and allowlist egress to the in-scope targets only (see Egress allowlist). The agent reaches the model API through the credential proxy, not through the allowlist. Where one is available, point the engagement at a staging or replica environment that is disconnected from production.

Broker access to the target where you can

Give the agent access to target systems through something you can observe, such as an access proxy or a defined set of tools. Avoid letting the agent write its own tooling with general access to the target. In a custom harness, separate read-only tools from tools that change state, and supervise the second group more closely. Consider analyzing risky commands as they are proposed, and denying them or escalating them to a person.

Supervise runs that have network access

We recommend that an engineer watch each run as it executes, following tool calls and network activity, and be able to halt the run immediately. Agents act at machine speed, so live observation complements the controls that act before an action executes (network allowlists, brokered tools, and review of state-changing actions). It does not replace them. For long or autonomous runs where continuous attention is impractical, monitor continuously in software (see Offline monitoring of agent transcripts).

Keep autonomous agents away from live high-consequence systems

Do not run autonomous agents against live production systems where an out-of-scope action could endanger safety or the availability of critical services, such as OT/ICS, medical, or energy systems. This is already the norm for human-led testing of such systems, and it applies equally here. Test against a replica, a testbed, or a digital twin, or during a planned outage. Where live access is unavoidable, limit the agent to passive or read-only activity and have a person carry out every state-changing step.

Test the sandbox before you rely on it

Before you run real engagements from a host, test its sandbox with the two steps below. Do this before first real use, and again whenever you change the model or the sandbox. Sandbox changes include a new host, a new version of the hypervisor or container runtime, a different network or proxy setup, and an edited container or compose configuration.

First, have Claude review the configuration

In an interactive Claude Code session on the host, ask it to read the sandbox configuration as it is set up on this machine and to report anything that weakens the boundary this article describes. Point it at:

  • The hypervisor and container-runtime configuration (for the reference design, Kata's rendered configuration.toml and Docker's daemon.json);

  • The agents' network and the egress proxy as the runtime reports them, and the proxy's allowlist;

  • The credential proxy as the runtime reports it, and the environment of a running agent container (see The credential proxy below);

  • Any compose or orchestration files that start agents or targets;

  • The list of what is mounted or streamed into guests (see Operating notes below).

This step only reads configuration and state. It does not launch an agent.

Then, run a supervised escape test

Launch the test agent the same way you launch production agents, and not with a hand-built command, so that the test exercises the real path. The section Verify isolation yourself below covers the hand checks of the boundary instead. Give the agent a single task: try to reach anything outside its boundary, and stop and report as soon as anything succeeds. Name the boundary concretely in the prompt: the host, other containers, the container-runtime socket, the cloud metadata service, and any host that is not on the proxy's allowlist. Run this only under close human supervision, on a host with no internet access beyond the model API and that you are prepared to rebuild, so that a successful escape does no damage. On bare metal, run the test inside a throwaway VM. On a cloud VM, use a fresh instance and destroy it afterwards. Watch the session while it runs. Treat any success as a blocker: fix the host, repeat the test, and do not run engagements from that host until the test passes.

Afterwards, read the whole transcript

Check what the agent tried as well as whether it got out. Look for direct connections that did not go through the proxy, reads of credential files or metadata addresses, use of the container-runtime socket or probes of neighboring containers, writes outside the container's own filesystem, and attempts to change the agent's own settings or tooling. The test passes when the transcript shows each of those was tried and failed for the reason you expect. Record, for each one, whether the sandbox or the permission classifier stopped it. An attempt that the classifier denied never reached the sandbox, so cover it with the hand checks (see Verify isolation yourself below). A passing test is a minimum bar, not proof of isolation.

Avoid impossible tasks, and review failed runs first

Confirm before launch that the task can be completed as given. Some examples of tasks that can’t be completed: a focus area with nothing to find, a target that is not reachable, a bug that is not there, or a tool the task needs that is not installed. Whenever you change your setup or start a new kind of task, confirm before launch that the target builds and can be reached and that the agent has the tools it needs. For the same reason, when you review a batch, read the transcripts of the runs that failed or found nothing before the ones that succeeded.

Verify isolation yourself

Run these checks by hand on each host. The second column describes the check for the reference design (Docker with the Kata runtime and Firecracker). Adapt it to your own runtime.

What to confirm

How

Expected result

A separate guest kernel

Run uname -r in a sandboxed container and on the host

The two versions differ

The VM monitor is jailed on the host

For a running container, inspect the Firecracker process: its root directory, Seccomp and NoNewPrivs in /proc/<pid>/status, its mount and network namespaces, and its cgroup

Only the VM's own files are under its root. Seccomp: 2 and NoNewPrivs: 1. Both namespaces differ from the host's. The cgroup path names the container

A host file is not visible inside

Create a file on the host and try to read it from a sandboxed container

Not found. This is a basic check that any runtime should pass

Egress is refused

From an agent container, request the model-API host and one other public host through the egress proxy

Both are refused, and the egress proxy logs a deny line for each. Agents reach the model only through the credential proxy

No credential in the agent container

While a run is up, list the agent container's environment, filtered for provider names

A placeholder, the credential proxy's address, and provider settings. No real key, token, or credential file

The reference design

This section describes how the reference implementation meets the guidance above. The implementation is not included. Read it as a design you can copy.

How each agent is isolated

Every agent runs as claude -p inside its own microVM, next to the target binary and source. The microVM is a Kata Containers container backed by the Firecracker virtual machine monitor, registered with Docker as a runtime. The agent's Read, Write, and Bash tools see only that guest's filesystem and kernel.

The connection between a guest and the host is deliberately small. It consists of KVM and the few virtual devices Firecracker emulates: block devices for the container's disk, one network interface, and a vsock control channel that Kata uses to start processes inside the guest. The Firecracker process itself is confined on the host. It runs in a "jail," which is a chroot directory holding only that VM's files, with its own mount namespace, the container's network namespace, and Firecracker's built-in seccomp filter. It is accounted to the container's cgroup. What you trust on the host side is KVM, Firecracker, and Kata's runtime process (the per-container "shim" that Docker talks to, which runs as root).

The orchestrator stays on the trusted host. It manages container lifecycle, streams transcripts, and moves files in and out with docker exec, which Kata serves through its agent inside the guest. The launcher starts agents only inside this sandbox. Before any agent starts, it checks that the sandbox runtime is registered, that /dev/kvm is present, and that the egress proxy is up, and it refuses to run otherwise.

What the sandbox changes for each surface:

Surface

Without sandbox

With sandbox

Agent Read/Write

host filesystem

guest filesystem only (the container's own virtual disk, under the guest kernel)

Agent Bash

host shell

guest shell only (guest kernel; the host is reachable only through KVM and Firecracker's virtual devices)

Network egress

whatever the host has

no internet access; model calls only, through the credential proxy

Model-API credential

in the agent's environment

not in the agent container; held by a separate credential-proxy container (see The credential proxy below)

Host coupling

full

docker exec for files in and out, served by Kata's agent inside the guest; read-only inputs arrive as copies taken at container start; files that change during a run are streamed in by the orchestrator

Permission checks

auto-mode classifier only

auto-mode classifier plus the microVM boundary (see Auto mode)

Where each property is enforced: hardware virtualization provides the kernel and filesystem boundary. The agent's Read/Write/Bash run against the guest kernel, and the Firecracker process behind it is jailed and seccomp-confined on the host. Egress policy is enforced on the host side of the VM's network interface, by a Docker --internal bridge (no default route out) and the two proxies on that network. The credential proxy forwards model calls and nothing else. The egress proxy refuses every other destination, because its allowlist is empty unless you add to it. Traffic leaves through the guest's own network stack; the filtering happens in the two proxies.

The credential proxy

The credential for the model API (your API key, OAuth token, AWS or Google credential) is never placed in an agent container. The orchestrator on the host reads and validates it, and each launch starts one small extra container, the credential proxy, to hold it. The proxy is attached to the agents' internal network, and agent containers send their model calls to it: their claude CLI gets the proxy's address as its API base URL, and a fixed placeholder value where the key or token would normally be. This follows the best practice of reaching the model API only through a proxy, with the API key injected from outside the sandbox. In the reference design the proxy is a separate container on the agents' network, not a process on the agent's own localhost. For every request, the proxy:

  1. Refuses anything that is not a model call to the one provider endpoint the launch was configured for (other paths and hosts get a 403 and a log line);

  2. Removes any auth headers the container sent;

  3. Adds the real credential: the API key header, a bearer token (which the proxy refreshes itself when the token is short-lived), or an AWS SigV4 signature computed over the exact request;

  4. Forwards the request to the provider over HTTPS with certificate verification, and streams the response back.

For a compromised agent container this means there is no key, token, or credential file to read, copy, or send anywhere. The cloud metadata services are also unreachable from the container. What the container can still do is make model calls through the proxy, on your account, for as long as the run is up. For that reason, prefer a narrowly scoped credential for agent runs, and rotate it after any run whose transcript shows unexpected behavior. The proxy logs one line per request, with the client address, method, path, status, and size. It does not log headers or bodies. Save that log when the run ends (see the Retention section in Offline monitoring of agent transcripts).

Anything the agent sends to the model API leaves your network, in request bodies that the proxy does not inspect or log.

What each side holds, per auth route:

Route

What the credential proxy holds

What the agent container gets

API key

the key

proxy address + placeholder ANTHROPIC_API_KEY

OAuth token (claude setup-token)

the token

proxy address + placeholder CLAUDE_CODE_OAUTH_TOKEN

Workload Identity Federation (WIF)

the identity token, plus the access token the proxy exchanges it for and refreshes. With ANTHROPIC_IDENTITY_TOKEN_FILE, the token file's directory is mounted read-only into the proxy, so the proxy sees a rotated token. With an inline ANTHROPIC_IDENTITY_TOKEN, the proxy holds that value and there is no file to re-read

proxy address + placeholder CLAUDE_CODE_OAUTH_TOKEN

ant auth login profile

the profile directory, mounted read-only into the proxy

proxy address + placeholder CLAUDE_CODE_OAUTH_TOKEN

Bedrock

the bearer token, or the access-key set (the proxy signs each request)

proxy address, CLAUDE_CODE_SKIP_BEDROCK_AUTH=1, region; no AWS credential

Vertex

the service-account key, plus the access tokens the proxy mints from it and refreshes

proxy address, CLAUDE_CODE_SKIP_VERTEX_AUTH=1, region and project; no Google credential

The credential proxy container is part of the trusted side of the setup. It runs under plain runc rather than the sandbox runtime, as root with every capability dropped except the one it needs to read the mounted files, and with a read-only filesystem. It listens only on its address on the agents' internal network, and it connects to the provider directly rather than through the egress proxy. The orchestrator passes the credential to the running proxy over docker exec. For the WIF token-file and profile routes, the directory holding the token or profile is mounted read-only into the proxy, which reads the file from there. The credential is not in the proxy container's environment or on its command line, so that docker inspect on it shows the mounts and not a secret.

Start one proxy per launch, and remove it when the run ends or is interrupted. Check for proxies left behind by a run that was killed, and remove them before the next launch.

The credential proxy has two side effects. First, the agents' CLI is given a non-default API address, so a few CLI features that apply only to the default address are off inside agent containers. Its telemetry and update checks are switched off as well, because they would have no route out anyway. Second, on Vertex the proxy's path filter is the main control that stops an agent container from using the rest of the Vertex AI API (see Egress allowlist). A check of the key's IAM scope at launch is a second layer.

Egress allowlist

The egress proxy's allowlist is empty by default. Agent containers reach the model API only through the credential proxy (see The credential proxy). Each launch starts that proxy for the provider your credentials select, so nothing provider-specific is stored in the sandbox setup. Every other destination an agent requests through the egress proxy, including the model-API host itself, gets a 403 and a deny line in its log. The region values that select the Bedrock and Vertex endpoints (AWS_REGION, CLOUD_ML_REGION) are checked against a strict pattern before they are used, so that a malformed value cannot send the credential to a different host.

Vertex has one limit to be aware of. …aiplatform.googleapis.com serves the whole Vertex AI API, so a credential that is allowed to create custom jobs could run an arbitrary container with full internet egress in your project. Use two controls against this. Have the credential proxy forward only publisher-model calls on your configured project (…/publishers/anthropic/models/…:rawPredict, :streamRawPredict, and :countTokens) and refuse every other path. And check the key's IAM scope at launch and fail closed: refuse gcloud user ADC, check a service-account key with testIamPermissions, and refuse it if it holds any workload-creation permission. Grant the account a custom role that holds only aiplatform.endpoints.predict. The metadata server, Google's STS, and the IAM-credentials endpoints should not be reachable from agent containers.

If agents need to reach extra hosts, such as a package mirror or an in-scope target, add them to the egress proxy's allowlist as host:port entries. Add only what the engagement needs (see Scope and supervision).

Do not add the model-API host to this list. Agents do not need it, because their model calls go through the credential proxy. Adding it gives agent containers a direct route to the API, and the rest of the design, including what the permission classifier is told, assumes they have none.

Changing the allowlist means restarting the egress proxy, which drops any live agent connections. Do it between batches rather than during one, and confirm afterwards which list the proxy loaded.

Derive the provider endpoint from your credential configuration, and ignore a base-URL variable such as ANTHROPIC_BASE_URL that is set on the host: the credential proxy should always forward to the endpoint that the credentials select.

Building and operating a sandbox like this

This section collects what was learned from running agents under Kata with Firecracker in the reference implementation. It is not a set of installation steps.

Host requirements

The reference design needs a physical Linux machine, or a VM that can itself run VMs:

  • Linux x86_64 or aarch64 with KVM (/dev/kvm present and usable): bare metal, or a cloud VM with nested virtualization enabled.

    • GCE, Azure, and AWS offer nested virtualization on selected instance types (on AWS, for example, the C8i, M8i, and R8i families). AWS bare-metal (*.metal) instances also work.

    • Nested virtualization costs some performance. It is not expected to weaken the boundary.

  • Block-device storage for containers. Firecracker cannot share a host directory into a guest; the only storage it can attach to a VM is a block device. Container root filesystems therefore have to exist as block devices instead of the usual overlay filesystem. With Docker, containerd's devmapper snapshotter provides that. It needs rootful Docker (the reference design uses Engine 25 or newer) backed by the system containerd (2.0 or newer).

  • Outbound network access while you build the host and the images. Agents do not get this egress.

It will not work on:

  • a container or a Kubernetes pod, including Docker-in-Docker and container-based CI runners. The host's containerd, Docker daemon, and device-mapper cannot be configured from inside a container, and KVM is usually not available there either;

  • rootless Docker;

  • a VM without nested virtualization, which is most general-purpose cloud instance types unless you pick one that offers it and turn it on;

  • macOS or Windows, including Docker Desktop and WSL2.

The reference design does not support macOS or Windows hosts

Its sandbox needs KVM on a Linux host. Docker on a Mac runs inside a shared Linux VM that mounts the user's home directory and holds the Docker daemon, so it cannot provide per-agent hardware isolation. If you work on a Mac, run the agents on a Linux host with KVM (bare metal, or a cloud VM with nested virtualization) and drive them over SSH. If you build your own sandbox on macOS or Windows, see the hypervisors named under the General guidelines section above.

Switching Docker's image store has a side effect

Images and containers created on Docker's previous store become invisible to Docker after the switch to the devmapper snapshotter. They stay on disk. To see them again, put both storage settings in daemon.json (storage-driver and features.containerd-snapshotter) back to what they were and restart Docker.

Kata settings

Pin the Kata release, and verify its digest before you install it. The reference design starts from Kata's Firecracker profile and pins these settings in /etc/kata-containers/configuration.toml:

Setting

Value

Why

jailer_path

path to Kata's jailer binary

Starts Firecracker inside its jail: a chroot with only the VM's files, its own mount namespace, the container's network namespace, and Firecracker's seccomp filter. Left unset, Kata would run Firecracker without the jail.

enable_annotations

[]

A container cannot change any hypervisor setting through OCI annotations.

disable_guest_seccomp

false

Docker's seccomp profile for the container is also applied inside the guest, a second layer under the VM boundary.

sandbox_cgroup_only

true

Firecracker and its I/O threads are placed in the container's cgroup, so the VM is accounted to the container. Whether the container's --memory limit also caps the VM depends on the host (see Sizing).

static_sandbox_resource_mgmt

true

Firecracker cannot add CPUs or memory to a running VM, so the VM is sized once at boot from the container's limits.

default_vcpus / default_memory

4 / 4096 MiB

Per-VM baseline; the container's --memory is added on top. Raise for build-heavy targets (Firecracker allows at most 32 vCPUs per VM).

debug_console_enabled, enable_debug

false

No console into guests, and no guest console output in host logs.

[factory] enable_template

false

Off, because VM templating would share guest memory pages between VMs.

entropy_source

/dev/urandom

Non-blocking host entropy for guests.

The kernel, image, and kernel_params values that come with the Kata release are kept as they are.

Make the guest kernel and image immutable

Every VM boots from the same two files, Kata bind-mounts them into every jail, and the jailer is started as root, so ordinary file permissions would not stop a compromised Firecracker process from rewriting the image that every later VM boots from. Set the immutable flag on both (chattr +i). Clearing that flag takes a system call that Firecracker's seccomp filter does not allow, which is designed to stop even root inside the jail from doing it. Clear the flag yourself before you install a new Kata release. On a filesystem without chattr support, use a read-only bind mount instead.

Sizing

Memory

Each agent VM is sized once, at boot: default_memory plus the container's --memory. With a 4096 MiB baseline and a 4g container limit, an agent is an 8 GiB VM. The host allocates that memory to the VM as the guest uses it, not all at boot, but plan for the full amount per concurrent agent: ten agents in parallel can grow toward 80 GiB. Whether the container's --memory limit also caps the VM as a whole depends on how the host's cgroups are laid out. Check which of these applies on your host:

  • The container limit applies to the whole VM. An agent cannot use more host memory than --memory even though its VM is nominally larger, and a VM that exceeds the limit is killed from the host side.

  • Nothing on the host caps the VM below its boot size. An agent's memory use is bounded by the VM size (default_memory + --memory) and the guest kernel's own out-of-memory killer.

  • A parent cgroup caps it at a different value.

We have not confirmed this behavior on every kind of host. Either way, size the host by VM size.

CPUs

Each VM gets default_vcpus. Firecracker cannot add CPUs to a running VM, so for build-heavy targets raise the value before launch; the maximum is 32 per VM.

Disk

Every image and every container gets a virtual disk out of a device-mapper thin pool, which backs the image store. If the pool fills, writes inside containers fail with I/O or out-of-space errors. If the filesystem that holds the pool fills first, the pool becomes read-only and every container on it fails. Free space with docker rmi and docker system prune (freed blocks return to the pool). sudo dmsetup status <pool> shows usage. For a long-lived host, an LVM thin pool on a dedicated disk is the better layout.

Boot time

Every agent start boots a guest kernel. That takes well under a second on bare metal and a few seconds under nested virtualization, which is small compared with agent run times.

Hardening the host

  • Dedicate the host to this job.

    • Assume an agent could read anything on it, and keep no credentials or sensitive data there other than the model-API credential the agents need.

  • Keep every layer that the agent's traffic touches patched, not only the kernel: Docker, the proxy containers' base images, and any firewall or network appliance between the host and the model API. An out-of-date proxy or firewall on the sandbox boundary is itself attack surface.

  • Keep the kernel, KVM, and CPU microcode current, and leave the kernel's CPU-vulnerability mitigations at their defaults.

    • grep . /sys/devices/system/cpu/vulnerabilities/* should show no Vulnerable lines.

    • If the host is shared with workloads that must not observe each other, also follow Firecracker's production host setup guide on SMT and speculative-execution settings.

  • Disable swap (sudo swapoff -a, and remove it from /etc/fstab) so that guest memory is not written to swap.

  • /dev/kvm does not need to be world-writable, and Kata's install and runtime directories must stay owned by root and not writable by other users.

    • root:kvm with mode 0660 is enough for /dev/kvm, because Kata's runtime runs as root.

  • Never add --privileged, --device, --cap-add, or host networking to agent containers or to target services.

    • Under Kata those flags pass host devices and privileges into the VM.

  • Keep agents on the internal network behind the two proxies.

    • Firecracker does no packet filtering of its own; the host-side network and proxies are the egress control.

  • Keep enable_debug off outside troubleshooting; debug logging includes guest console output.

Validating a new host

After you set up a fresh machine, this sequence shows that the sandbox works end to end. Steps 3 and 4 make real model calls.

  1. Run the isolation checks under Verify isolation yourself. The two kernel versions must differ. Note whether the container memory limit caps the VM (see Sizing).

  2. Confirm that the agent CLI runs under the sandbox runtime, in the image you will use.

  3. Test the boundary: have Claude review the sandbox configuration, then run the supervised escape test, both as described in Test the sandbox before you rely on it.

  4. Run a small batch end to end against a target you know. Run this step with the credential mode you will use for real engagements, not a stand-in API key: the credential proxy handles each mode differently (token exchange for WIF, signing for Bedrock, token minting for Vertex), and this run confirms that yours works end to end. Check the credential proxy's log afterwards.

  5. Check that nothing was left behind. Once the batch has finished, no agent container, helper container, VM process, or credential proxy should remain. The egress proxy's log should show no deny lines other than the probes from steps 1 and 3, and the credential proxy's log should show only model calls.

Operating notes

Running agents under Kata with Firecracker differs from running plain containers in these ways.

Bind mounts are copies; live files have to be streamed

Firecracker has no host-filesystem sharing, so Kata copies a bind-mounted file or directory into the guest when the container starts, and later changes on the host are not seen inside. That is fine for inputs that do not change during a run, such as target source, and those can stay read-only mounts. No credential file is mounted or streamed, because agent containers hold none (see The credential proxy). A file that does change during a run has to be written into the container by the orchestrator, through docker exec, and kept current. The in-container copies are ordinary files that an agent could edit, so have the orchestrator read the originals on the host, and judge findings from a copy that the agent under test cannot reach.

Each agent needs a companion container for its network

A microVM's network interfaces must exist when the VM boots; Firecracker cannot add one later. Docker, however, connects a container's network only after the runtime has created the container. The reference design works around this. It first starts a small idle container that is attached to the right network and runs only sleep, and then starts the agent inside that container's network namespace (--network container:<name>). Kata finds the interfaces there at boot and attaches them to the VM. A companion serves exactly one VM, because Kata leaves the VM's host-side network devices behind in it, so create and remove the pair together and do not reuse a companion. One costs an idle process of roughly 12 MB.

Stopping a stuck agent

docker rm -f <agent-container> is enough. The orchestrator should detect the dead container, mark the run as failed, and remove the companion container when it tears the run down.

Everything is addressed by IP

Docker's embedded DNS runs in the host-side network namespace and is unreachable from inside a guest, so pass the proxies to agents as IP addresses, and give networked targets a static IP.

docker exec works, docker cp does not

docker exec into an agent container behaves as usual; Kata's agent inside the guest runs the command. docker cp into or out of an agent container does not work, because the container's files are on a virtual disk inside the guest. Use docker exec <container> cat <path> and the like.

Firecracker runs as root inside its jail

Kata starts the jailer with uid 0, so the Firecracker process is confined by its chroot, namespaces, seccomp filter, and cgroup, not by an unprivileged user id. The two files every VM shares, the guest kernel and image, are protected by the immutable flag instead (see Kata settings).

Kata has deprecated the runtime that Firecracker depends on

Kata runs Firecracker through its older Go runtime. Kata 4.0 made a newer Rust runtime the default for its other hypervisors and deprecated the Go runtime. Upstream says the Go runtime still receives critical bug fixes and CVE fixes, and that it may be removed no sooner than Kata 5.0. The Rust runtime does not list Firecracker among its hypervisors, and upstream tests its Docker integration mainly with QEMU. Plan for a migration. When you change the Kata version, repeat Validating a new host. Kata can also use Cloud Hypervisor instead of Firecracker; the reference implementation has not tested or hardened that configuration.

Logs

Kata's runtime, Firecracker, and guest-agent messages go to the journal: journalctl -t kata. containerd and Docker problems are in journalctl -u containerd and journalctl -u docker.

Did this answer your question?