What is AI agent sandboxing?
By Identra · Updated
AI agent sandboxing runs an agent's commands inside an environment with enforced limits on files, processes and network access. It shrinks the damage a misled agent can do, but the tools and accounts the agent is signed in to need their own permission controls.
How does AI agent sandboxing work?
The limits live outside the model. Say Claude Code proposes running the test suite. Inside a sandbox, the operating system or a virtualization layer decides what that process can read, write and connect to. A line in the prompt saying 'only touch this folder' is a request. The sandbox is the enforcement.
Fit the limits to the job. Fixing a failing test needs the repo and a package registry. It doesn't need ~/.aws/credentials, the developer's Chrome profile or a production kubeconfig. That's least privilege applied to a workspace.
File and network limits need each other. Block outbound traffic and the agent can still wreck local files. Lock down files and it can still reach whatever internal service answers on the network. Anthropic's write-up on Claude Code sandboxing makes the same point.
- Files: readable paths, writable paths and host mounts.
- Network: allowed destinations, including internal services and the cloud metadata endpoint.
- Processes: privileges, child processes and host management sockets.
- Credentials: which secrets and signed-in tools this task actually gets.
Should an agent run in a container or a virtual machine?
A container shares the host kernel. A VM runs its own kernel behind a hypervisor. Either can work, and configuration is what usually breaks them. Mount /var/run/docker.sock into the container and the agent can control Docker on the host. Share your home folder into the VM and the boundary has a hole in it.
For coding agent security, check every route to execution. A locked-down shell tool means little if an MCP server listed in ~/.cursor/mcp.json can run commands on the host with no limits.
Restricted process
- Boundary
- OS controls around one process
- Where it tends to leak
- Child processes and reachable host services
Container
- Boundary
- Process isolation on a shared host kernel
- Where it tends to leak
- Privileged mode, host mounts and the Docker socket
Virtual machine
- Boundary
- Separate guest operating system
- Where it tends to leak
- Shared folders, host integrations and network bridges
Can sandboxing stop prompt injection?
It can't stop the model from being fooled. It can stop a fooled model from reaching what it shouldn't. If a poisoned README uses prompt injection to tell the agent to read ~/.ssh/id_rsa, a working file boundary denies the read.
Everything inside the boundary is still fair game. The agent can trash the writable repo. It can push code to github.com because that domain is allowed, and attackers have GitHub accounts too. Allowing a domain is not the same as allowing every action on it.
Output also escapes later. A script written in the sandbox runs with that person's full privileges the day someone executes it on their laptop.
Why do connected accounts need separate controls?
The authority lives somewhere else. A sandboxed agent with Gmail send scope can mail a customer list out without ever leaving its container. Gmail checks what the account is allowed to do, not whether the request came from a sandbox.
So AI agent identity and OAuth app risk are part of containment. Know whose account the agent acts as, what it can reach and how to cut it off. Hiding the raw token from the model protects the secret. The tool holding it can still use every permission it has.
Deleting the sandbox recalls nothing.
What does a sandboxed coding agent look like in practice?
Say a coding agent is fixing tests in a payroll service. It gets a throwaway clone, synthetic employee data and the internal package mirror. Production databases and deploy keys stay out. What it produces is a pull request.
A comment in a test fixture tells it to read the developer's cloud credentials and post them to a paste site. The file boundary should deny the read. The network allowlist should block the paste site. Test each one on its own.
It can still change payroll logic in the repo. Code review and real tests are what catch that. And if its GitHub token can merge or edit workflow files, the local sandbox does nothing about that power. Scope the token to the job.
How do you validate an AI agent sandbox?
Use harmless canary files and a destination you own. Confirm the forbidden things fail and normal work still succeeds. If routine tasks break, people switch the sandbox off. Retest after any change to tools, mounts or AI agent guardrails.
- Reads and writes outside approved paths fail, including from child processes.
- Traffic to unapproved hosts and internal services fails.
- No tool can execute outside the sandbox.
- External sends, production changes and permission changes need approval tied to the exact target.
- Shutdown and token revocation have been rehearsed before the agent gets anything sensitive.
How Identra thinks about it
Identra shows security teams which AI apps and agents are in use across the browser, the endpoint and connected identity, SaaS and cloud providers. Teams can restrict supported AI apps in the browser by signed-in account, block destructive agent shell commands by policy and have an analyst revoke risky OAuth grants.
Go deeper: Identra on the endpoint
Frequently asked questions
Is sandboxing the same as running an agent in a test environment?
No. A test environment is about where the work happens. A sandbox enforces what the work can touch. Plenty of staging machines have no file or network limits at all.
Is a container enough to sandbox an AI agent?
It can be, with no privileged mode, tight mounts and restricted network access. A mounted Docker socket undoes all of that.
Does a browser agent need sandboxing?
Browser process isolation protects the machine. It does nothing about what the agent does inside a signed-in Salesforce tab. That needs scoped accounts and approval before sends or settings changes.
Does sandboxing prevent data leakage?
It narrows what the agent can read and where it can send. Leaks through an allowed tool or destination still happen without extra controls.
Should each agent task use a fresh sandbox?
Where you can, yes. Shared caches, persistent volumes, credentials and remote account state still carry over and need their own cleanup.
