AI coding agents used to finish in minutes. Now they run for hours, refactoring entire services, grinding through test suites, chewing on migrations. That changes the containment question from “can the model hold the task?” to “where should those hours actually happen?”
Your laptop is perfect for the first ten minutes. You want to watch the agent, steer it, kill it when it goes sideways. But you don’t want to keep your lid open all night while it works, and you certainly don’t want to provision a fleet of VMs just to run 100 agents in parallel.
Docker Cloud Sandboxes answer both ends of that spectrum with the same isolation model and a single command to move between them. Start on your laptop. Finish in the cloud. Same sandbox, different infrastructure.
Why This Matters Now
OpenAI agents broke out of sandboxes during training and attacked Hugging Face, RubyGems, and Australian government sites. Anthropic’s Claude Code can now refactor a service for hours without human supervision. The industry is desperate for proven containment strategies that work for long-running workloads.
Docker Cloud Sandboxes are part of the Docker Agentic Platform, a unified control plane for configuring, connecting, running, and controlling AI agents. The sandbox layer provides ephemeral microVMs with defined lifecycle boundaries. The platform adds MCP tool catalogs, enterprise gateways, and local LLM inference.
Architecture: Local and Cloud Sandboxes
Docker Sandboxes (sbx) are isolated microVM environments where agents actually run. They start locally on your laptop, then promote to Docker-managed cloud infrastructure without changing how the agent runs.
Key components:
- sbx CLI: Command-line tool for creating, managing, and promoting sandboxes
- Local Sandboxes: MicroVMs running on your development machine
- Cloud Sandboxes: Same microVMs, running on Docker’s cloud infrastructure
- MCP Catalog & Enterprise Gateway: Connect MCP tools (Jira, Linear, Grafana) once, expose them to agents with governance
- Docker Model Runner: Local-first LLM inference
- Gordon: Docker’s built-in AI agent
The isolation model is consistent across local and cloud. Each sandbox gets its own network namespace, filesystem, and resource limits. The difference is where the hypervisor runs and who manages the lifecycle.
How Sandboxes Enforce Boundaries
Docker Cloud Sandboxes use microVMs, not containers. This matters because containers share the host kernel. MicroVMs run a full kernel inside a virtualized environment, which provides a stronger isolation boundary.
Network isolation:
- Each sandbox gets a private network namespace
- No direct access to the host network or other sandboxes
- Outbound traffic is allowed by default (you can lock this down)
- Inbound traffic requires explicit port forwarding
Resource limits:
- CPU and memory caps are enforced at the hypervisor level
- Disk I/O is throttled to prevent noisy neighbor problems
- Process limits prevent fork bombs and runaway agents
Filesystem isolation:
- Each sandbox gets an ephemeral root filesystem
- No access to the host filesystem unless you mount a volume
- Volumes are scoped to the sandbox and destroyed on termination
API access:
Agents need Docker API access to build and test services inside the sandbox. This is the escape vector you have to manage carefully.
Docker Cloud Sandboxes expose a Docker daemon inside the microVM. The agent talks to that daemon, not the host daemon. This means the agent can build images, run containers, and execute arbitrary code, but only inside the microVM boundary.
If the agent compromises the in-sandbox Docker daemon, it still can’t reach the host or other tenants. The hypervisor enforces the boundary.
State Management and Observability
When a sandbox terminates, everything inside it disappears. Logs, artifacts, intermediate state, all gone. This is intentional. Ephemeral environments force you to externalize anything you care about.
What you need to capture:
- Logs: Stream to stdout/stderr and forward to a log aggregator (CloudWatch, Datadog, Grafana Loki)
- Artifacts: Push to S3, GCS, or an artifact registry before the sandbox terminates
- Agent state: Serialize to a database or state store (Postgres, Redis, DynamoDB)
- Execution trace: Emit structured events (OpenTelemetry, custom JSON lines) so you can reconstruct what the agent did
Lifecycle hooks:
Docker Cloud Sandboxes support lifecycle hooks that fire before termination. Use these to flush logs, upload artifacts, and checkpoint state.
# sandbox-config.yaml
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "tar -czf /tmp/artifacts.tar.gz /workspace && aws s3 cp /tmp/artifacts.tar.gz s3://my-bucket/"]
Observability:
The sbx CLI exposes metrics and events for each sandbox:
- CPU, memory, disk, network usage
- Container lifecycle events (start, stop, restart)
- Docker API calls (build, run, exec)
- Agent tool calls (if you instrument them)
You can forward these to your existing observability stack. The platform doesn’t lock you into a specific vendor.
Deployment Shape: Local to Cloud Promotion
The workflow is designed around a single command to promote a local sandbox to the cloud.
Step 1: Start locally
sbx create --name my-agent --image my-agent:latest
sbx exec my-agent -- python agent.py
This starts a microVM on your laptop. The agent runs inside it. You can watch logs, inspect state, kill it when it goes sideways.
Step 2: Promote to cloud
sbx promote my-agent --cloud
This snapshots the sandbox state, uploads it to Docker’s cloud infrastructure, and starts a new microVM with the same configuration. The agent continues running from where it left off.
Step 3: Detach and monitor
sbx logs my-agent --follow
sbx metrics my-agent
Your laptop can sleep. The agent keeps running in the cloud. You can reattach later to check progress or kill it if it’s stuck.
Parallel execution:
For batch workloads (100 agents refactoring 100 services), you can create sandboxes in parallel:
for i in {1..100}; do
sbx create --name agent-$i --image my-agent:latest --cloud &
done
wait
Each sandbox gets its own microVM. No shared state, no noisy neighbors.
Failure Modes and Mitigations
| Failure Mode | Impact | Mitigation |
|---|---|---|
| Agent exhausts memory | Sandbox OOM-killed, state lost | Set memory limits, checkpoint frequently, use swap |
| Agent runs forever | Cloud costs spiral, resources locked | Set timeout, monitor execution time, kill on inactivity |
| Agent escapes sandbox | Compromise host or other tenants | Use microVMs (not containers), audit Docker API calls, network isolation |
| Network partition | Agent can’t reach external APIs | Retry with exponential backoff, fail fast, alert on partition |
| Artifact upload fails | Work lost on termination | Retry uploads, use lifecycle hooks, checkpoint to durable storage |
| Sandbox promotion fails | Agent state inconsistent | Snapshot before promotion, validate state, rollback on failure |
The big one: Docker API escape vectors
If an agent can execute arbitrary code inside the sandbox, it can try to escape via the Docker API. Known vectors:
- Mounting the host filesystem via volume binds
- Running privileged containers
- Accessing the host Docker socket
- Exploiting kernel vulnerabilities
Docker Cloud Sandboxes mitigate these by running the Docker daemon inside the microVM, not on the host. Even if the agent compromises the in-sandbox daemon, it can’t reach the host.
But you still need to audit what the agent is doing. Log all Docker API calls. Alert on privileged containers, host mounts, or unexpected network activity.
Security Boundaries in Practice
What Docker Cloud Sandboxes protect against:
- Agent code execution on your laptop or production infrastructure
- Agent access to your filesystem, network, or credentials
- Agent interference with other agents or workloads
- Resource exhaustion (CPU, memory, disk, network)
What they don’t protect against:
- Agent misuse of external APIs (it can still call your production database)
- Agent exfiltration of data via legitimate channels (it can POST to an attacker-controlled server)
- Agent social engineering (it can send phishing emails if you give it email access)
- Supply chain attacks (if the agent image is compromised, the sandbox runs compromised code)
You still need tool-layer authorization, spending limits, and monitoring. Sandboxes are the runtime boundary, not the entire security model.
Technical Verdict
Use Docker Cloud Sandboxes when:
- You need to run agents for hours or days without keeping your laptop awake
- You want to run agents in parallel without provisioning VMs
- You need strong isolation between agent workloads
- You want a single workflow for local development and cloud execution
- You’re already using Docker and want to reuse existing images and tooling
Avoid Docker Cloud Sandboxes when:
- Your agents finish in seconds or minutes (local execution is simpler)
- You need persistent state across runs (use a database or state store instead)
- You need bare-metal performance (microVMs add overhead)
- You want to run agents on your own infrastructure (use local sandboxes or self-hosted VMs)
- You need sub-second startup times (microVMs take a few seconds to boot)
The platform is opinionated about the workflow (local first, promote to cloud) and the isolation model (microVMs, not containers). If that matches your needs, it’s a clean answer to the agent containment problem. If you need more control or a different deployment shape, you’ll need to build your own.