The sandbox: isolating the agent from your machine
Objectives of this module
- Know what a code agent can do on your machine
- Compare the disposable clone, the container and the micro virtual machine: what each protects, and what it costs
- Launch Pi in Docker Sandboxes with a kit versioned in this repository
- Leave with a sandbox in which the exercises of the following modules run unattended
The following modules launch Pi twenty times on the same task without human intervention, give it sub-agents that have a shell, then chain sub-agents into pipelines. Pi has no mechanism for asking your consent before executing a command, and its security documentation says it clearly: the tools read, write and run commands "with the permissions of the pi process", and "Pi does not include a built-in sandbox". Anything you can do from your terminal, the agent can do too: read ~/.ssh, read ~/.pi/agent/auth.json where your API keys are stored, run git push --force, or send the content of a file to any domain with curl.
The natural reaction is to write a prompt: "only modify game/neon.js", "don't read anything outside the repository". A prompt is text, and we remind you that using an LLM is always non-deterministic, which means you will never have a 100% guarantee that it will be followed. In the module on skills, you will observe that an instruction for cleaning up temporary files placed in a SKILL.md is followed less than one time in three. Before the first unattended run, you therefore need a limit that doesn't depend on the model's obedience. The sandbox is a deterministic way to ensure that the LLM is in a closed environment whose boundaries are defined by you and that the model cannot overstep.
Understanding
What does the agent have access to?
A code agent running on your machine has access to your files, that is, the repository it works on and, with the same rights, your home directory, where the SSH keys, the model provider tokens and the .env files of your other projects live. The network lets it install any package, run a curl | sh found in a README, or broadcast what it has just read. Finally, it launches processes with your identity, which covers the Docker daemon, the rm command and write access to the remote repository. If you also have sudo privileges on your machine, nothing stops it.
These actions don't even require the model to make a mistake. A repository file can contain instructions written for the agent: that is the role of NÉON's SUPPORT.md. Its text mimics a support procedure but asks the agent to read the .env and send its contents to an external address. An agent that opens this file to answer a question treats the instruction as if it came from you, and the module on permissions will work on this case. An extension installed from the community directory runs, as the module on Pi reminded you, with all of your permissions. In both cases, the flaw is in the harness, and a guardrail written inside AGENTS.md will not protect you.
Pi's documentation concludes: "For untrusted repositories, generated code you do not intend to monitor closely, or unattended automation, run pi in a contained environment. Use a container, VM, micro-VM, remote sandbox, or policy-controlled sandbox with only the files and credentials required for the task." Our twenty runs on issue #1 are exactly unattended automation. And eventually, we want autonomous agents that can work for hours without us having to monitor them.
Three levels of isolation
The cheapest of the three is the disposable clone. The measurement tool of the next module clones NÉON at a tag, into a temporary directory, on every run, which protects the repository's history and working tree at little or no cost. Still, the process runs under your identity, with your home directory and your network, so a disposable clone only protects the repository. And even then, nothing stops the model from pushing to your remote repository if it has the rights, which is the case if it has access to the gh command (to work on your GitHub).
One step further, the container runs Pi in a Docker image where only the repository is mounted, putting your home directory out of reach. It shares the host kernel, its network is open by default, and above all the model provider's API key must be added to it so that Pi can call the model, which the Pi page on containerization notes in one sentence: "Provider API keys enter the container". So everything the agent runs has access to this key.
The third level is the policy micro-VM with Docker Sandboxes. Each sandbox has its own kernel behind a hypervisor, all outbound TCP traffic passes through a proxy on the host that only accepts domains from an allowlist, and this proxy injects the API keys into the HTTP headers, so that, to quote the security page, "Credential values never enter the VM". The working directory is mounted into the VM at the same absolute path as on the host. The cost is a 700 MB image to build, a daemon to run, a list of allowed domains to maintain. It may seem complicated, but your favorite AI can help you set up this infrastructure easily.
| what is protected | disposable clone | container | Docker Sandboxes |
|---|---|---|---|
| the repository's working tree | yes | no | no by default, yes with --clone |
| your home directory | no | yes, if only the repository is mounted | yes |
| outbound network | no | no by default | yes, denied by default with an allowlist |
| your API keys | no | no, they get into the image | yes, only the host's proxy sees them |
These three levels isolate Pi's process from the host machine, but nothing inside the sandbox yet stops Pi from running rm -rf on the repository or reading a .env lying around in NÉON. The pi-permission-system extension adds this filter inside the sandbox itself: it hooks into the tool_call event of Pi's extension API, a hook that intercepts every tool call, every bash command, every MCP call, and every skill invocation before it executes, and compares the request against allow / deny / ask rules written in JSON.
The trade-off lies in where this filter runs. It lives in the same Node process as Pi, not in the kernel that isolates the sandbox, so an extension that compromised this process before the rule is evaluated would disable the guard along with everything else. pi-permission-system tightens what Pi can do once launched in the sandbox; it doesn't replace any of the three levels in the table above.
What the sandbox doesn't protect
In direct mode, the default one, the agent edits your working tree in place, and the Docker Sandboxes documentation reminds us that it can therefore modify a git hook, a Makefile or a continuous integration configuration, which will run later on the host when you launch them yourself. The sandbox protects the machine while the agent works, and it does not spare you from reviewing the diff.
The balanced network policy, the one that sbx policy init recommends, allows domains with broad wildcards like *.googleapis.com, which cover far more than model APIs. We start from deny-all and then only open the domains that appear in the denial log.
Finally, inside the VM, the agent is an administrator, with passwordless sudo and its own Docker daemon, which we accept, since nothing that happens there ever leaves it and the VM itself is disposable.
Rebuilding
We offer you two approaches below: using an extension in Pi that adds a hook (which we will suggest rebuilding in another module) and using Docker Sandboxes. The first solution requires no special installation on your system, and will therefore be used for the classroom training. But bear in mind that it has its limits and that it is clearly not sufficient for daily work with agents.
Installing and configuring pi-permission-system
The extension installs with one command, like any package from Pi's directory:
pi install npm:@gotgenes/pi-permission-systemThe rules live in a JSON file, read at three scopes: global (~/.pi/agent/extensions/pi-permission-system/config.json), project (.pi/extensions/pi-permission-system/config.json, ignored if the project is not approved), and per-agent, in the YAML header of an agent file, which takes precedence over the first two. For NÉON, a project configuration is enough to stop the most dangerous instruction in SUPPORT.md, since reading a .env is refused by construction:
{
"permission": {
"*": "allow",
"path": {
"*": "allow",
"*.env": "deny",
"*.env.*": "deny"
},
"bash": {
"*": "ask",
"rm -rf *": "deny",
"sudo *": "ask"
},
"external_directory": "ask"
}
}The most specific rule wins: bash.* requires confirmation by default, rm -rf * refuses without asking, and a path outside the repository remains subject to confirmation even when path.* allows everything else. A command that the extension's bash parser cannot classify is refused rather than let through, and a path that traverses a symbolic link is resolved before comparison.
Exercise (in class)
Work in a disposable clone of NÉON, since two of the requests below are destructive. Install the extension, place the configuration above in .pi/extensions/pi-permission-system/config.json, create a .env at the root containing a fake key, then launch Pi and ask it three things: to read you the contents of this .env, to delete the game/ folder with rm -rf, and to run the test suite. The first two requests are refused without Pi consulting you; the third opens a confirmation prompt that you answer yourself.
Next, resubmit the request to read the .env three times in a row, rephrasing it, then explaining to Pi that you are the owner of the file and that you authorize it: the verdict does not change, because it comes from a rule evaluated before the tool call and not from model arbitration.
Finally, remove the path block from the configuration and replace it with the instruction "never read a .env file" in the repo's AGENTS.md, then resubmit the same request five times in five different sessions. Count the refusals you get: you now have your own figure for what a text instruction is worth against a guardrail in code.
Install and configure sbx
sbx is the command to use Docker Sandboxes. sbx knows a list of agents it can launch as-is (claude, codex, copilot, cursor, gemini, opencode and a few others). Unfortunately, Pi is not among them. It is therefore necessary to create a kit: a directory described by a spec.yaml whose kind: sandbox variant defines an agent from scratch: the image, the startup command, the instructions added to the context file, the keys to inject, and the network permissions. Ours is versioned in https://github.com/AI-for-dev/pi-sandbox and contains only three files.
pi-sandbox
├── Dockerfile
├── spec.yaml
└── files/home/.pi/agent/settings.jsonThe versions listed below are those with which this kit was verified at the time of writing: sbx 0.45.1, Docker Engine 29.7.2, Pi 0.87.1.
Install sbx
The command-line tool is called sbx. To install it on your OS, simply go to the following page
https://docs.docker.com/ai/sandboxes/install/
Build the image
FROM docker/sandbox-templates:shell-docker
USER root
ARG NODE_VERSION=22.21.1
ARG PI_VERSION=0.87.1
# Ubuntu names the package fd-find and ships the binary as fdfind, to avoid a
# name collision. pi looks for fd then fdfind, so /usr/bin/fdfind is enough and
# pi stops downloading its own copy into ~/.pi/agent/bin.
ARG FD_PACKAGE_VERSION=10.3.0-2ubuntu1
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
xz-utils ca-certificates curl "fd-find=${FD_PACKAGE_VERSION}" \
&& fdfind --version \
&& rm -rf /var/lib/apt/lists/*
# Explicit Node install instead of inheriting from the template: pi requires
# >= 22.19, and the base image's bundled version is not a contract.
RUN set -eux; \
case "$(dpkg --print-architecture)" in \
amd64) a=x64 ;; \
arm64) a=arm64 ;; \
*) echo "unsupported architecture" >&2; exit 1 ;; \
esac; \
cd /tmp; \
curl -fsSLO "https://nodejs.org/dist/v${NODE_VERSION}/node-v${NODE_VERSION}-linux-${a}.tar.xz"; \
curl -fsSLO "https://nodejs.org/dist/v${NODE_VERSION}/SHASUMS256.txt"; \
grep " node-v${NODE_VERSION}-linux-${a}.tar.xz$" SHASUMS256.txt | sha256sum -c -; \
mkdir -p /opt/node; \
tar -xJf "node-v${NODE_VERSION}-linux-${a}.tar.xz" -C /opt/node --strip-components=1; \
rm -f /tmp/*.tar.xz /tmp/SHASUMS256.txt
ENV PATH="/opt/node/bin:${PATH}"
RUN npm install -g "@earendil-works/pi-coding-agent@${PI_VERSION}" \
&& pi --version
USER agentThe image starts from the shell-docker base image provided by Docker, installs an explicit version of Node, because Pi requires at least 22.19, then pins the version of Pi.
The Docker Sandboxes daemon pulls its images from a registry different from the local images available to Docker. Without a registry, you go through an archive:
git clone https://github.com/AI-for-dev/pi-sandbox
cd pi-sandbox
docker build --platform linux/arm64 -t pi-sandbox:0.87.1 .
docker image save pi-sandbox:0.87.1 -o pi-sandbox.tar
sbx template load pi-sandbox.tarFor a team, you'll prefer to push the image to a registry.
Declare the kit
schemaVersion: "2"
kind: sandbox
name: pi
version: "0.1.0"
displayName: Pi
description: Pi coding agent (pi.dev) in a Docker sandbox.
sourceURL: https://github.com/earendil-works/pi
sandbox:
image: "pi-sandbox:0.87.1"
entrypoint: [pi, -a]
agentInstructions:
filename: AGENTS.md
content: |
## Sandbox environment
Tu tournes dans une microVM Docker Sandbox. `sudo` est sans mot de passe,
Docker est disponible a l'interieur de la VM. Le reseau sortant est filtre
par une allowlist: un domaine non autorise echoue, ce n'est pas une panne
reseau. La cle du provider n'est pas dans la VM, seule une sentinelle l'est.
environment:
variables:
PI_SKIP_VERSION_CHECK: "1"
PI_TELEMETRY: "0"
NODE_OPTIONS: "--disable-warning=UNDICI-EHPA"
credentials:
- service: ilaas
description: ILAAS API KEY (llm.ilaas.fr)
required: true
apiKey:
name: ILAAS_API_KEY
proxyManaged: true
inject:
- domain: llm.ilaas.fr
header: Authorization
format: "Bearer %s"
permissions:
network:
allow:
- github.com
- raw.githubusercontent.com
- pypi.org
- files.pythonhosted.org
- pi.devThe sandbox block names the image created in the previous step and launches pi -a. The -a option declares the project files as safe for this run, which answers the question that trust.json asked in the module on Pi: inside the VM, a skill or extension found in the repository can only touch what the VM contains.
agentInstructions adds a few lines to the AGENTS.md the model reads: a denied domain is not a network failure, which keeps it from retrying ten times, and the provider key is not in the VM.
The credentials block declares a key managed by the proxy (proxyManaged: true). Pi finds in ILAAS_API_KEY a sentinel, a dummy value, and the host's proxy replaces it with the real key in the Authorization header of requests to llm.ilaas.fr, and nowhere else.
Under permissions.network, the kit opens, on top of the global policy, the model provider, GitHub for cloning NÉON, and PyPI for the measurement tools. The PI_SKIP_VERSION_CHECK and PI_TELEMETRY variables disable some of Pi's startup network operations.
The files/home/.pi/agent/settings.json file, which the kit places in the agent's home directory, sets the provider, the default model, and the reasoning level. It replaces the ~/.pi/agent/settings.json of your host, which is not mounted in the VM. It is quite simple here and looks like this:
{
"defaultProvider": "ilaas",
"defaultModel": "deepseek-v4-flash",
"defaultThinkingLevel": "high"
}Likewise, the files/home/.pi/agent/models.json file lists the models available in the sandbox.
{
"providers": {
"ilaas": {
"baseUrl": "https://llm.ilaas.fr/v1",
"api": "openai-completions",
"apiKey": "$ILAAS_API_KEY",
"models": [
{
"id": "gemma-4-31b",
"contextWindow": 128000,
"reasoning": true
},
{
"id": "qwen-3.6-35b-instruct",
"contextWindow": 256000
}
]
}
}
}Registering the key
ilaas authenticates with an API key. You entrust it to sbx under the name of the service declared by the kit:
sbx secret set ilaasYou must then provide your key. You can then verify that it is properly registered with the command
sbx secret lsOn first launch, sbx asks you to approve the credential binding, the authorization given to a third-party kit to use this secret on the domains it declares. The answer is recorded in ~/.config/sbx/credentials.yaml.
In non-interactive mode, nobody answers
With sbx create or from a script, nobody is asked the binding question. The sandbox starts anyway, sbx only issuing a warning, and the environment variable contains the proxy-managed sentinel: the real key is never injected by the proxy. The error therefore only appears at runtime, as a 401, and pi auth check nevertheless announces ready. A binding per declared service is necessary. Write the file beforehand:
bindings:
ilaas:
apiKey:
domains: [llm.ilaas.fr]Setting the network policy
This setting is global, sbx requires it before the first sandbox, and it is done once and for all:
sbx policy init deny-allThe kit's permissions.network.allow rules apply on top, for its sandboxes only.
Launching
cd pi-sandbox
sbx kit validate .
cd /chemin/vers/neon
sbx run /chemin/vers/pi-sandboxExercise (on your own)
Walk through the previous five steps on your machine, from installing sbx to the first sbx run. In the Pi session that opens, ask for the value of the ILAAS_API_KEY variable: you will see the sentinel, not your key. Then run a curl https://example.com: the request fails, because the domain is not on any list. Finally, have a NÉON file modified: the change appears on the host side as soon as Pi has written.
Go back to the host and read sbx policy log, where every refusal is logged with the domain requested. To finish, redo the pi-permission-system exercise, this time inside the sandbox: the two guards stack without getting in each other's way, and the rm -rf refusal keeps all its usefulness, since in direct mode the repository Pi would erase is the one on your host.
Tightening the allowlist
Exercise (on your own)
Work a full session inside the sandbox, then reread sbx policy log. Add to permissions.network.allow only the domains whose refusal actually blocked you, rerunning sbx kit validate after each change.
Generalizing
A usage boundary that does not depend on obedience. A permission written in text, in an AGENTS.md or a SKILL.md, is a suggestion the model may or may not follow. The permissions module will build guards in code inside the harness, which refuse a tool call before it executes. The sandbox is the outer layer, the one that holds when the harness itself is at fault, because a malicious extension or a booby-trapped file can only damage the VM.
The key stays on the host. The agent does not need to read the key, only that its requests to a specific domain be authenticated, and keeping the key on the host, to place it in the header as the request goes through, removes it from everything the agent can read, execute, or send. The principle holds for any harness, regardless of the tool that implements it.
Deny by default, then open up from the log. A kit's allowlist is not written in advance: you start from denial, work a session, and only add the domains whose refusal actually blocked something, just as the rest of the act decides on measurement rather than intuition.
Deliverable
At the end of this module, Pi runs in a sandbox on your clone of NÉON, and all the operations of the following modules can be done there with optimal control.
Four checks confirm it:
- the
ILAAS_API_KEYvariable read from the sandbox is the sentinel; - a request to a domain absent from the list fails;
- an edit made by Pi appears in the repository on the host side;
sbx policy logshows no refusal your list did not choose.
Going further
The threat model
- Kai Greshake, How We Broke LLMs: Indirect Prompt Injection - the blog post that accompanies the foundational article by Greshake et al., Not what you've signed up for: data read by the model becomes an instruction, and Copilot can already be compromised by a package's documentation.
- Simon Willison, The lethal trifecta for AI agents - access to private data, exposure to untrusted content, and the ability to communicate outward: the three combined are enough for exfiltration.
- Simon Willison, Agents Rule of Two and The Attacker Moves Second - the "at most two of three properties" rule formulated by Meta, and an article that knocks down twelve published defenses against prompt injection under adaptive attack.
- Beurer-Kellner et al., Design Patterns for Securing LLM Agents against Prompt Injections - architecture patterns that constrain what the agent can do, at the cost of part of its utility.
- Korny Sietsma, Agentic AI and Security - the trifecta applied to coding agents on martinfowler.com: containers, least privilege, task decomposition.
- OWASP, Top 10 for Agentic Applications 2026 - ten risk families, including supply chain compromise and unintended code execution.
- Marchand et al., Quantifying Frontier LLM Capabilities for Container Sandbox Escape - a 2026 benchmark where agents find and exploit vulnerabilities in a vulnerable container to escape it, i.e. the measured argument for a separate kernel.
Documented incidents
- Johann Rehberger, The Month of AI Bugs - one bug per day in August 2025 in code agents (Claude Code, Codex, Cursor, Copilot, Devin, Jules, OpenHands), which Simon Willison summarized.
- Johann Rehberger, Amazon Q Developer: Remote Code Execution with Prompt Injection - a
find -execclassified as read-only is enough to execute code without approval. - Will Vandevanter (Trail of Bits), Prompt injection to RCE in AI agents - argument injection into pre-approved commands, and the sandbox recommended as the main defense instead of lists of safe commands.
- Kevin Higgs (Trail of Bits), Prompt injection engineering for attackers: Exploiting GitHub Copilot - a poisoned GitHub issue makes Copilot Agent add a backdoor dependency.
- Pillar Security, Rules File Backdoor - hidden instructions in a Cursor or Copilot rules file (March 2025), i.e. the
SUPPORT.mdtrap observed in real-world conditions. - Nx, S1ngularity postmortem and Wiz, attack analysis - a compromised npm package (August 2025) enlists the code agents installed on the workstation, launched without confirmation, to locate secrets to exfiltrate.
- Fortune, Replit AI wiped a production database - an agent wipes a production database during a change freeze (July 2025), even though a written instruction forbade it.
- Pillar Security, The Agent Security Paradox - CVE-2026-22708 (January 2026): internal shell commands like
export, outside Cursor's allowlist, poison the environment of approved commands. - Unit 42, OpenClaw's Skill Marketplace and the Emerging AI Supply Chain Threat - malicious markdown skills on an agent's marketplace (2026), the same risk as for a package installed with
pi install. - Ken Huang, Coding Agent Security: Lessons from Claude Code, Cowork, Codex, and Copilot in the Wild - eight incidents from 2025 and 2026, and a comparison of Claude Code, Codex, Copilot, and Cursor sandboxes (August 2026).
- Simon Willison, Breaking Claude Code Opus 5 Auto Mode - a Johann Rehberger attack that succeeded four times out of five against Claude Code's auto mode (August 2026), and the conclusion that a classifier does not replace the sandbox.
Isolation practices and mechanisms
- Mario Zechner, What I learned building an opinionated and minimal coding agent - Pi's author explains why Pi has no permissions ("As soon as your agent can write code and run code, it's pretty much game over") and recommends running it in a container.
- Armin Ronacher, Agentic Coding Recommendations - the openly embraced
claude-yoloalias, and the risk moved into Docker. - Simon Willison, Designing agentic loops - YOLO mode both essential to productivity and dangerous, hence the sandbox, preferably on someone else's computer.
- Simon Willison, Codex CLI sandbox investigation - Seatbelt on macOS, Landlock and seccomp on Linux, or how another harness makes the same choice.
- sysid, Your Agent Has Root - the built-in tools that escape the kernel sandbox, and a Pi extension to close the gap.
- Andrew Lock, Running AI agents safely in a microVM using docker sandbox - the full
sbxjourney on a developer machine, network policies included. - Michael Krämer, Trust but Sandbox - Docker Sandboxes from a team's perspective: policies, secrets proxy, custom images.
- Palaimon, Coding Agents III: Sandboxing & Best Practices - dev containers, bubblewrap and VMs compared, with the startup cost quantified.
- Ry Walker, Local AI Agent Sandboxes - eight local sandbox tools compared, and what is left for a third-party tool once harnesses integrate their own.
- Daniel Vaughan, Agent Sandbox Comparison Matrix - Codex's Seatbelt, OpenShell and Docker
sbx: isolation boundary, network, secrets. - Agache et al., Firecracker - the AWS Lambda micro-VM (NSDI 2020), the reference text on the trade-off between isolation and startup time.
- Emir Beganović, Your Container Is Not a Sandbox: The State of MicroVM Isolation in 2026 - why a container is not a security boundary, the episode where Claude Code disables its own bubblewrap, and a tour of available micro-VMs (March 2026).
- Greg Hurrell, List of coding agent sandboxes - a catalogue kept up to date in 2026, from system primitives to hosted platforms, in ten categories.
- Zheng et al., ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses - a harness policy enforced in the Linux kernel via eBPF (June 2026), with a measured overhead between 2 and 8%.
Tools
- Docker Sandboxes: architecture, security model and kits.
- Pi: Security and Containerization.
- pi-sandbox - a per-command system sandbox for Pi, with an authorization prompt, based on
sandbox-execor bubblewrap. - pi-gondolin and Gondolin - Pi's tools run in a local micro-VM; both projects describe themselves as experimental.
- OpenShell - a runtime with declarative policies (filesystem, network, processes, inference), cited by Pi's documentation.