Execution Environment and Sandbox (E)
This page corresponds to §3 of Agent Harness Engineering: A Survey. Execution is the physical substrate of an agent harness: the environment in which actions actually run. In LLM-agent systems, execution and sandboxing are tightly coupled because the agent may write files, run shell commands, install packages, browse the web, or operate a desktop.
Why Sandboxing Is First-Class
The paper argues that sandboxing serves three purposes in agent systems:
- Security: model-generated code is too large and too dynamic to rely on static review. Multi-step autonomy means harmful actions may happen without a human in the loop. Prompt injection can also redirect an otherwise benign agent into an attack carrier.
- Reproducibility: long-horizon tasks and evaluations need resettable execution state. SWE-bench, OSWorld, and similar benchmarks are practical because the environment can be restored to a known baseline.
- Liveness: without a sandbox, every risky action needs a permission prompt. At scale, users either abandon the workflow or reflexively approve everything. A sandbox creates a bounded area where the agent can act freely.
The paper’s useful phrase is that a sandbox is both a cage and a license. It restricts blast radius, but it also makes autonomous execution possible.
Sandbox Categories
The survey classifies sandboxes by workload rather than by isolation primitive:
| Category | Representative systems | Design signal |
|---|---|---|
| General-purpose managed sandboxes | Daytona, E2B, Modal, Northflank, OpenSandbox, Docker Sandboxes | API-exposed execution for arbitrary OCI images; increasingly moving toward dedicated-kernel microVMs. |
| Computer-use infrastructure | Anthropic Computer Use, CUA, OSWorld | Full GUI/OS interaction through mouse, keyboard, and screenshots; high fidelity, high attack surface. |
| Code and repository execution sandboxes | Judge0, OpenAI Code Interpreter, sandboxed.sh, langchain-sandbox, Repo2Run | Language/runtime preinstallation, high concurrency, often stateless and request-scoped. |
| Framework-integrated runtimes | OpenHands, GoEX, agent-infra sandbox, smolagents executors | Bundled into an orchestration loop; convenient but tightly coupled. |
| Browser evaluation environments | WebArena, VisualWebArena, BrowserGym, WorkArena | Both sandbox and benchmark substrate; natural surface for indirect prompt injection. |
| OS-level permission sandboxes | Anthropic sandbox-runtime, Claude Code sandboxing, IsolateGPT, AgentBound, transactional sandboxing | Narrow the host view rather than creating a full new OS image. |
| Sandbox abstraction layers | SWE-ReX, smolagents executor, K8s Agent Sandbox CRD, R2E-Gym, EnvScaler, SandMLE | Unify multiple execution backends so agent logic is decoupled from where it runs. |
Isolation technology is orthogonal: containers, gVisor-style user-space kernels, Firecracker/Kata microVMs, WebAssembly, bubblewrap, Seatbelt, seccomp, and full VMs can appear under multiple workload categories.
Threat Model
The agent setting amplifies traditional sandbox threats:
- prompt injection can cause the agent to run attacker-directed actions;
- goal misalignment can make sandbox escape instrumentally useful to the agent;
- compositional amplification lets one weak tool combine with others into a larger breach.
SandboxEscapeBench reports that frontier models can exploit Docker sandbox weaknesses under realistic configurations. This makes sandbox escape an agent-relevant threat today, not a purely theoretical concern.
The paper distinguishes infrastructure isolation from semantic/capability isolation. A container bounds damage after an action runs. A tool permission policy decides whether the action should have been allowed at all. Robust harnesses need both.
Deployment Modes
| Mode | Strength | Cost |
|---|---|---|
| Self-hosted | Lowest latency, tight local iteration, strong control over data | The developer owns security and operations. |
| Cloud/SaaS | Elastic scale and managed isolation | Network round trips and provider trust. |
| Hybrid/BYOC | Keeps data locality while using remote capacity | More complex identity, routing, and audit boundaries. |
Compliance, data residency, and auditability often push systems toward hybrid deployment even when latency and scale would otherwise favor a single mode.
Open Problem
Execution environments need to become both measurable and composable. The paper highlights three gaps:
- common security evaluations for prompt injection, goal misalignment, and compositional amplification;
- cheap reset and replay for large-scale training/evaluation, including learned surrogate environments whose fidelity still needs to be proven;
- portability layers that preserve semantics across Linux containers, macOS, Windows, browser environments, desktop VMs, cloud, self-hosted, and hybrid deployments.
The bundle-vs-compose choice between framework-integrated runtimes and sandbox abstraction layers should be treated as an empirical harness design question, not just a product preference.