# Best code execution sandboxes for AI agents > Microsoft Execution Containers (BB), Modal Sandboxes (BB) and Amazon Bedrock AgentCore Code Interpreter (BB) lead the 15 ranked code execution sandboxes. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Canonical: https://www.anchorterminal.com/best/code-sandboxes/ - Markdown: https://www.anchorterminal.com/best/code-sandboxes/index.md (~5,900 tokens) - Slim: https://www.anchorterminal.com/best/code-sandboxes/index.min.md (~1,630 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/best/code-sandboxes/index.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-08 The 10 highest-scoring of 15 code execution sandboxes on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 15 · agent-ready (BB or better): 3 · accept x402: 0 · hosted endpoints: 10 - Full ranked table: https://www.anchorterminal.com/categories/code-sandboxes.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/code-sandboxes/index.md (105) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Microsoft Execution Containers](https://www.anchorterminal.com/tools/microsoft-execution-containers.md) | BB | 76.3 | A developer building an agent or tool host that must run model-written code on the user's own machine, above all on Windows, where it reaches Microsoft's process and session isolation. | Free · OSS | local | | 2 | [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md) | BB | 75.5 | GPU work inside a sandbox, or agents already running on Modal. | $0.071 / vCPU-hr | local | | 3 | [Amazon Bedrock AgentCore Code Interpreter](https://www.anchorterminal.com/tools/agentcore-code-interpreter.md) | BB | 73.1 | A team already on AWS that wants agent code to run under IAM, inside a VPC and beside S3 data, for sessions up to eight hours. | $0.0895 / vCPU-hr | hosted | | 4 | [Vercel Sandbox](https://www.anchorterminal.com/tools/vercel-sandbox.md) | B | 69.6 | Agents that spend most of their time waiting on a model, and teams already on Vercel. | $0.128 / vCPU-hr | hosted | | 5 | [E2B](https://www.anchorterminal.com/tools/e2b.md) | B | 68.3 | Code interpreters and agent workspaces that pause and resume with memory, and teams that may want to self-host later. | $0.0504 / vCPU-hr | hosted | | 6 | [Cloudflare Sandbox SDK](https://www.anchorterminal.com/tools/cloudflare-sandbox-sdk.md) | B | 67.5 | Agents already built on Workers and Durable Objects that want sandboxes in the same account with credential injection at the edge. | $0.072 / vCPU-hr | local | | 7 | [Runloop Devboxes](https://www.anchorterminal.com/tools/runloop.md) | B | 64.8 | Running and grading coding agents on full workstations with prebuilt blueprints and credential gateways. | $0.108 / vCPU-hr | hosted | | 8 | [Daytona](https://www.anchorterminal.com/tools/daytona.md) | B | 64.3 | Agents that need a choice of machine, including Windows desktops and GPUs, and operators who want least-privilege keys. | $0.0504 / vCPU-hr | hosted and local | | 9 | [Blaxel Sandboxes](https://www.anchorterminal.com/tools/blaxel-sandboxes.md) | C | 60.7 | Long-lived, mostly idle agent workspaces that need to resume quickly with memory intact, and teams that want an MCP server per sandbox. | $0.1656 / session-hr | hosted | | 10 | [Freestyle](https://www.anchorterminal.com/tools/freestyle.md) | C | 58.5 | Suited to long-running agent work that needs a full Linux VM, memory-preserving pause, snapshots to branch from and private networking. | $0.0403 / vCPU-hr | hosted | ## Picks by need - Highest score overall: [Microsoft Execution Containers](https://www.anchorterminal.com/tools/microsoft-execution-containers.md), BB, 76.3/100 on the benchmark. Also [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md), BB, 75.5/100. - Reliability: [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md), 95/100 on reliability, against 81 for the overall leader. - Schema & documentation: [E2B](https://www.anchorterminal.com/tools/e2b.md), 92/100 on schema & documentation, against 81 for the overall leader. - Agent ergonomics: [Amazon Bedrock AgentCore Code Interpreter](https://www.anchorterminal.com/tools/agentcore-code-interpreter.md), 75/100 on agent ergonomics, against 74 for the overall leader. - Security & auth: [Amazon Bedrock AgentCore Code Interpreter](https://www.anchorterminal.com/tools/agentcore-code-interpreter.md), 87/100 on security & auth, against 69 for the overall leader. - Maintenance & community: [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md), 93/100 on maintenance & community, against 92 for the overall leader. - Lowest paid price per vCPU hour: [Agent 37 Cloud](https://www.anchorterminal.com/tools/agent37.md), $0.0011 per vCPU hour, the lowest of the 13 listings here with a paid price in this unit (free allowances aside). Also [Sprites](https://www.anchorterminal.com/tools/sprites.md), $0.0385 per vCPU hour. - A hosted MCP endpoint: [Blaxel Sandboxes](https://www.anchorterminal.com/tools/blaxel-sandboxes.md), remote MCP server, nothing to install. - Self-hosting under an open licence: [E2B](https://www.anchorterminal.com/tools/e2b.md), self-hosted, Apache-2 licence. - The review panel's favourite: [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md), 3.3/5 from 8 panel reviews. ## How to choose - Start time and concurrency: Check the start time and how many sandboxes you can run at once, because an agent that waits on cold starts or hits a concurrency cap stalls mid-task. - Isolation and what it covers: Check what isolation the vendor documents and whether it covers the network, filesystem and host, because an agent running untrusted code needs a boundary it can rely on. - Billing per second and idle time: Work out the cost of a sandbox that sits paused or idle, since a plan may charge for memory or storage while nothing runs. - Pause, resume and checkpoints: Check whether state can be paused, resumed or checkpointed and what survives a restore, because a restore that drops installed packages forces the agent to start again. - How the benchmark tests this category: The same task in every sandbox: start, install a package, run a script that writes files, pause and resume where supported, then tear down. We time each step, check isolation claims against the docs and add up the cost. ## Each one in detail ### 1. Microsoft Execution Containers, BB 76.3/100 Microsoft Execution Containers (MXC) is an open-source SDK for running untrusted code in a local sandbox on Windows, Linux and macOS. An application embeds it through Node.js, .NET or Rust and sets filesystem, network and UI policy for each run. - Verdict: MXC puts nine operating-system sandbox backends behind one typed request, with network access denied by default and a JSON Schema for the stable 1.0.0 contract. Version 1.0.0 is two days old as of 8 October 2026. Enforcement varies by backend, and `isolation_session` cannot restrict networking at all. - Choose it for: A developer building an agent or tool host that must run model-written code on the user's own machine, above all on Windows, where it reaches Microsoft's process and session isolation. - Strength: MIT licence, with SDKs for Node.js, .NET and Rust all at 1.0.0 and the native runtime bundled in the npm and NuGet packages - Strength: Egress, ingress and host loopback default to `deny`, and filesystem access is limited to listed read-only and read-write paths - Strength: A draft-07 JSON Schema for the stable 1.0.0 request, with descriptions on 135 of 150 properties - Weakness: 1.0.0 shipped on 6 October 2026, and the Node changelog still lists the V1 changes under Unreleased - Weakness: Enforcement differs by backend. `isolation_session` cannot restrict networking, and proxy routing is cooperative on Seatbelt and WSLC - Weakness: Persistent containers exist only for `isolation_session` and `wslc`, both on Windows - Price: Free · OSS · Auth: None · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/microsoft-execution-containers.md ### 2. Modal Sandboxes, BB 75.5/100 Modal's sandboxed compute environments for running code, with SDK access, GPU support and filesystem snapshots. - Verdict: GPU sandboxes at the same per-second rates as the rest of Modal. No REST API, and the JavaScript and Go SDKs are beta. - Choose it for: GPU work inside a sandbox, or agents already running on Modal. - Strength: GPU sandboxes at the same per-second rates as the rest of Modal - Strength: Outbound traffic blockable or limited to CIDR ranges, and no inbound connections without tunnels - Strength: $30 of compute every month on Starter, no card - Weakness: No REST API, and the JavaScript and Go SDKs are beta - Weakness: Default lifetime of 5 minutes and a hard maximum of 24 hours - Weakness: gVisor rather than a VM unless you're on Team or Enterprise for the VM runtime - Price: $0.071 / vCPU-hr · Auth: API key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/modal-sandboxes.md - Against #1: https://www.anchorterminal.com/compare/microsoft-execution-containers-vs-modal-sandboxes.md ### 3. Amazon Bedrock AgentCore Code Interpreter, BB 73.1/100 Amazon Bedrock AgentCore Code Interpreter is AWS's managed sandbox for running agent-written Python, JavaScript and TypeScript. Each session is a dedicated microVM, reached through the AWS API, the AgentCore SDKs or an MCP server. - Verdict: Each session runs in its own microVM with 2 vCPU, 8 GB and a 10 GB disk for up to eight hours, and IAM can scope access to one interpreter. Sessions cannot be paused or resumed, and an AWS account with IAM set-up is needed before a first call. - Choose it for: A team already on AWS that wants agent code to run under IAM, inside a VPC and beside S3 data, for sessions up to eight hours. - Strength: Each session runs in a dedicated microVM, which AWS says is terminated and has its memory sanitised when the session ends - Strength: Billing is per second on CPU used and peak memory, at $0.0895 a vCPU-hour and $0.00945 a GB-hour, with I/O wait and idle time free - Strength: IAM actions cover each of the nine operations, with separate ARNs for the system interpreter and custom ones - Weakness: No pause, resume or snapshot. Session files are removed when the session ends, and persistence needs a customer-owned S3 Files or EFS mount inside a VPC - Weakness: Every session is capped at 2 vCPU, 8 GB of memory and 10 GB of disk, and the cap is not adjustable - Weakness: `InvokeCodeInterpreter` takes one flat `arguments` object for nine operations, so the reference does not say which fields each operation requires - Price: $0.0895 / vCPU-hr · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/agentcore-code-interpreter.md - Against #1: https://www.anchorterminal.com/compare/agentcore-code-interpreter-vs-microsoft-execution-containers.md ### 4. Vercel Sandbox, B 69.6/100 Firecracker microVM sandboxes on Vercel, driven from the `@vercel/sandbox` JavaScript SDK, the Python `vercel` package, a CLI or the REST API. - Verdict: Active CPU billing, so waiting on model responses costs only memory. Tied to a Vercel team and project even when called from elsewhere, and access tokens reach the whole team. - Choose it for: Agents that spend most of their time waiting on a model, and teams already on Vercel. - Strength: Active CPU billing, so waiting on model responses costs only memory - Strength: Credential brokering proxy outside the sandbox that overwrites headers set by sandbox code - Strength: Firecracker microVM with root access, and a firewall with deny-all, domain and CIDR rules - Weakness: Tied to a Vercel team and project even when called from elsewhere, and access tokens reach the whole team - Weakness: Hobby caps sessions at 45 minutes and pauses creation once the monthly allowance is spent - Weakness: Snapshots keep the filesystem, not memory or running processes - Price: $0.128 / vCPU-hr · Auth: OAuth or key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/vercel-sandbox.md - Against #1: https://www.anchorterminal.com/compare/microsoft-execution-containers-vs-vercel-sandbox.md ### 5. E2B, B 68.3/100 Firecracker microVM sandboxes for agent code, driven from Python and JavaScript SDKs, a CLI or a REST API. - Verdict: Firecracker microVM with its own kernel per sandbox. Two major incidents over an hour in September 2026, on sandbox creation and on creating from snapshots. - Choose it for: Code interpreters and agent workspaces that pause and resume with memory, and teams that may want to self-host later. - Strength: Firecracker microVM with its own kernel per sandbox - Strength: Egress allow and deny lists by domain, IP or CIDR, and secrets filled in outside the sandbox - Strength: Apache-2.0 infrastructure in e2b-dev/infra, self-hostable with Terraform - Weakness: Two major incidents over an hour in September 2026, on sandbox creation and on creating from snapshots - Weakness: One unscoped API key per project, with no audit log found - Weakness: Hobby sandboxes stop after 1 hour of continuous running - Price: $0.0504 / vCPU-hr · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/e2b.md - Against #1: https://www.anchorterminal.com/compare/e2b-vs-microsoft-execution-containers.md ### 6. Cloudflare Sandbox SDK, B 67.5/100 TypeScript library for running sandboxed Linux containers from a Cloudflare Worker. - Verdict: Each sandbox runs in its own VM with a separate filesystem, process space and network stack. No hosted API. You deploy and secure a Worker before an agent can call anything, and the starter has no auth. - Choose it for: Agents already built on Workers and Durable Objects that want sandboxes in the same account with credential injection at the edge. - Strength: Each sandbox runs in its own VM with a separate filesystem, process space and network stack - Strength: Outbound handlers hold credentials in the Worker and inject them, so the container never sees them - Strength: Egress can be turned off or limited to a deny-by-default `allowedHosts` list - Weakness: No hosted API. You deploy and secure a Worker before an agent can call anything, and the starter has no auth - Weakness: Open bugs on backups that drop directories (#859) and restores of archives of 10 MB or more (#884) - Weakness: Needs Workers Paid at $5 a month, with no card-free route - Price: $0.072 / vCPU-hr · Auth: None · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/cloudflare-sandbox-sdk.md - Against #1: https://www.anchorterminal.com/compare/cloudflare-sandbox-sdk-vs-microsoft-execution-containers.md ### 7. Runloop Devboxes, B 64.8/100 Devboxes, VM sandboxes for coding agents, with blueprints for prebuilt images, disk snapshots, suspend and resume, idle policies and a gateway that adds secret-backed headers to outbound API and MCP calls. - Verdict: Gateway credentials remain on Runloop servers, with access tokens bound to one devbox. Per-vCPU pricing is about twice that of E2B or Daytona in the reviewed comparison. - Choose it for: Running and grading coding agents on full workstations with prebuilt blueprints and credential gateways. - Strength: Agent gateways keep real credentials on Runloop's servers, with gateway tokens bound to one devbox - Strength: Network policies that block egress or allow listed hostnames - Strength: Public OpenAPI, llms.txt and typed Python and TypeScript SDKs with cursor pagination and 429 backoff - Weakness: About twice the per-vCPU price of E2B or Daytona - Weakness: No published rate limits or API error-body reference - Weakness: Suspend keeps disk only, and processes need restarting after resume - Price: $0.108 / vCPU-hr · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/runloop.md - Against #1: https://www.anchorterminal.com/compare/microsoft-execution-containers-vs-runloop.md ### 8. Daytona, B 64.3/100 Sandboxes for agent code in container, Linux VM, Windows and GPU classes, driven by SDKs for Python, TypeScript, Ruby, Go and Java or a REST API. - Verdict: API keys with per-action scopes, so an agent can create sandboxes without being able to delete them. The container class shares the host kernel. Only the VM classes get their own. - Choose it for: Agents that need a choice of machine, including Windows desktops and GPUs, and operators who want least-privilege keys. - Strength: API keys with per-action scopes, so an agent can create sandboxes without being able to delete them - Strength: Container, Linux VM, Windows and GPU sandbox classes behind one API - Strength: Three public OpenAPI files, llms.txt and SDKs in five languages - Weakness: The container class shares the host kernel. Only the VM classes get their own - Weakness: Full internet access and per-sandbox allow lists need Tier 3 - Weakness: Platform code went private in June 2026, and the old repository's 311 open issues won't be answered - Price: $0.0504 / vCPU-hr · Auth: API key · x402: no · Where: hosted and local - Full assessment: https://www.anchorterminal.com/tools/daytona.md - Against #1: https://www.anchorterminal.com/compare/daytona-vs-microsoft-execution-containers.md ### 9. Blaxel Sandboxes, C 60.7/100 Sandbox VMs that drop to standby seconds after the last connection and resume in about 25 ms with memory and filesystem kept, charging only for snapshot storage while idle. - Verdict: No compute charge in standby, only $0.20 a GB-month of snapshot storage. 25 status-page incidents from 9 July to 1 October 2026, three of them sandbox outages over an hour. - Choose it for: Long-lived, mostly idle agent workspaces that need to resume quickly with memory intact, and teams that want an MCP server per sandbox. - Strength: No compute charge in standby, only $0.20 a GB-month of snapshot storage - Strength: MicroVM per sandbox with domain allow and deny lists that can be enforced at network level - Strength: An MCP server in every sandbox over streamable HTTP, 18 tools - Weakness: 25 status-page incidents from 9 July to 1 October 2026, three of them sandbox outages over an hour - Weakness: No published request-rate limits, 429 guidance or SLA - Weakness: Domain filtering and secret injection are marked public preview, and egress is open by default - Price: $0.1656 / session-hr · Auth: OAuth or key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/blaxel-sandboxes.md - Against #1: https://www.anchorterminal.com/compare/blaxel-sandboxes-vs-microsoft-execution-containers.md ### 10. Freestyle, C 58.5/100 Freestyle runs full Linux virtual machines for AI agents, with pause and resume that keeps memory, snapshots, private networks and domains. Agents drive it through a REST API, a TypeScript SDK and a CLI from one npm package. - Verdict: A public OpenAPI 3.1 file describes 88 operations with a stable error code on every failure, and per-VM identity tokens keep the team API key away from end users. No terms of service, product privacy policy, request rate limit or incident history was found, and the only SDK is TypeScript. - Choose it for: Suited to long-running agent work that needs a full Linux VM, memory-preserving pause, snapshots to branch from and private networking. - Strength: Public OpenAPI 3.1 file with 88 `/v5` operations, each with a description and named error codes per status - Strength: A new VM has no inbound or outbound network access until firewall rules allow it - Strength: Identity access tokens limited to granted VMs and Linux users, revocable without rotating the team API key - Weakness: No terms of service or service agreement found on the site, and the privacy policy covers the website only - Weakness: The status page shows three components with no incident history - Weakness: No request rate limit, Retry-After guidance or idempotency keys found in the docs or the OpenAPI file - Price: $0.0403 / vCPU-hr · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/freestyle.md - Against #1: https://www.anchorterminal.com/compare/freestyle-vs-microsoft-execution-containers.md 5 more are ranked in the full table: https://www.anchorterminal.com/categories/code-sandboxes.md ## Head to head - [Microsoft Execution Containers vs Modal Sandboxes](https://www.anchorterminal.com/compare/microsoft-execution-containers-vs-modal-sandboxes.md) - [Amazon Bedrock AgentCore Code Interpreter vs Microsoft Execution Containers](https://www.anchorterminal.com/compare/agentcore-code-interpreter-vs-microsoft-execution-containers.md) - [Microsoft Execution Containers vs Vercel Sandbox](https://www.anchorterminal.com/compare/microsoft-execution-containers-vs-vercel-sandbox.md) - [E2B vs Microsoft Execution Containers](https://www.anchorterminal.com/compare/e2b-vs-microsoft-execution-containers.md) - [Amazon Bedrock AgentCore Code Interpreter vs Modal Sandboxes](https://www.anchorterminal.com/compare/agentcore-code-interpreter-vs-modal-sandboxes.md) - [Modal Sandboxes vs Vercel Sandbox](https://www.anchorterminal.com/compare/modal-sandboxes-vs-vercel-sandbox.md) - [E2B vs Modal Sandboxes](https://www.anchorterminal.com/compare/e2b-vs-modal-sandboxes.md) - [Amazon Bedrock AgentCore Code Interpreter vs Vercel Sandbox](https://www.anchorterminal.com/compare/agentcore-code-interpreter-vs-vercel-sandbox.md) - [Amazon Bedrock AgentCore Code Interpreter vs E2B](https://www.anchorterminal.com/compare/agentcore-code-interpreter-vs-e2b.md) - [E2B vs Vercel Sandbox](https://www.anchorterminal.com/compare/e2b-vs-vercel-sandbox.md) ## Questions ### What are the highest-rated code execution sandboxes for AI agents? Microsoft Execution Containers has the highest benchmark score of the 15 ranked code execution sandboxes, 76.3 (BB). Modal Sandboxes is second with 75.5 (BB). ### How many code execution sandboxes are agent-ready? 3 of the 15 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark. ### Which code execution sandboxes accept x402 payments? None of the ranked listings here accepts x402 for its main call yet. ### Which of these code execution sandboxes is cheapest? By published paid prices, Agent 37 Cloud, at $0.0011 per vCPU hour, the lowest of the 13 listings here with a paid price in this unit (free allowances aside). Plans, volume tiers and free allowances change the sum, so check the listing's price table. ### How is this list ranked? By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 8 October 2026. ## How this list is made The order is the Anchor benchmark score, the same number as on each listing. Each listing is graded from public evidence against the benchmark checklist, and the picks are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.