# Best code execution sandboxes for AI agents (slim) > Microsoft Execution Containers (BB), Modal Sandboxes (BB) and Amazon Bedrock AgentCore Code Interpreter (BB) lead the 15 ranked code execution sandboxes. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Full: https://www.anchorterminal.com/best/code-sandboxes/index.md (~5,900 tokens) · this version ~1,630 tokens · JSON https://www.anchorterminal.com/best/code-sandboxes/index.json · canonical https://www.anchorterminal.com/best/code-sandboxes/ - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-08 The 10 highest-scoring of 15 code execution sandboxes on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 15 · agent-ready (BB or better): 3 · accept x402: 0 · hosted endpoints: 10 - Full ranked table: https://www.anchorterminal.com/categories/code-sandboxes.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/code-sandboxes/index.md (105) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Microsoft Execution Containers](https://www.anchorterminal.com/tools/microsoft-execution-containers.md) | BB | 76.3 | A developer building an agent or tool host that must run model-written code on the user's own machine, above all on Windows, where it reaches Microsoft's process and session isolation. | Free · OSS | local | | 2 | [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md) | BB | 75.5 | GPU work inside a sandbox, or agents already running on Modal. | $0.071 / vCPU-hr | local | | 3 | [Amazon Bedrock AgentCore Code Interpreter](https://www.anchorterminal.com/tools/agentcore-code-interpreter.md) | BB | 73.1 | A team already on AWS that wants agent code to run under IAM, inside a VPC and beside S3 data, for sessions up to eight hours. | $0.0895 / vCPU-hr | hosted | | 4 | [Vercel Sandbox](https://www.anchorterminal.com/tools/vercel-sandbox.md) | B | 69.6 | Agents that spend most of their time waiting on a model, and teams already on Vercel. | $0.128 / vCPU-hr | hosted | | 5 | [E2B](https://www.anchorterminal.com/tools/e2b.md) | B | 68.3 | Code interpreters and agent workspaces that pause and resume with memory, and teams that may want to self-host later. | $0.0504 / vCPU-hr | hosted | | 6 | [Cloudflare Sandbox SDK](https://www.anchorterminal.com/tools/cloudflare-sandbox-sdk.md) | B | 67.5 | Agents already built on Workers and Durable Objects that want sandboxes in the same account with credential injection at the edge. | $0.072 / vCPU-hr | local | | 7 | [Runloop Devboxes](https://www.anchorterminal.com/tools/runloop.md) | B | 64.8 | Running and grading coding agents on full workstations with prebuilt blueprints and credential gateways. | $0.108 / vCPU-hr | hosted | | 8 | [Daytona](https://www.anchorterminal.com/tools/daytona.md) | B | 64.3 | Agents that need a choice of machine, including Windows desktops and GPUs, and operators who want least-privilege keys. | $0.0504 / vCPU-hr | hosted and local | | 9 | [Blaxel Sandboxes](https://www.anchorterminal.com/tools/blaxel-sandboxes.md) | C | 60.7 | Long-lived, mostly idle agent workspaces that need to resume quickly with memory intact, and teams that want an MCP server per sandbox. | $0.1656 / session-hr | hosted | | 10 | [Freestyle](https://www.anchorterminal.com/tools/freestyle.md) | C | 58.5 | Suited to long-running agent work that needs a full Linux VM, memory-preserving pause, snapshots to branch from and private networking. | $0.0403 / vCPU-hr | hosted | ## Picks by need - Highest score overall: [Microsoft Execution Containers](https://www.anchorterminal.com/tools/microsoft-execution-containers.md), BB, 76.3/100 on the benchmark. Also [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md), BB, 75.5/100. - Reliability: [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md), 95/100 on reliability, against 81 for the overall leader. - Schema & documentation: [E2B](https://www.anchorterminal.com/tools/e2b.md), 92/100 on schema & documentation, against 81 for the overall leader. - Agent ergonomics: [Amazon Bedrock AgentCore Code Interpreter](https://www.anchorterminal.com/tools/agentcore-code-interpreter.md), 75/100 on agent ergonomics, against 74 for the overall leader. - Security & auth: [Amazon Bedrock AgentCore Code Interpreter](https://www.anchorterminal.com/tools/agentcore-code-interpreter.md), 87/100 on security & auth, against 69 for the overall leader. - Maintenance & community: [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md), 93/100 on maintenance & community, against 92 for the overall leader. - Lowest paid price per vCPU hour: [Agent 37 Cloud](https://www.anchorterminal.com/tools/agent37.md), $0.0011 per vCPU hour, the lowest of the 13 listings here with a paid price in this unit (free allowances aside). Also [Sprites](https://www.anchorterminal.com/tools/sprites.md), $0.0385 per vCPU hour. - A hosted MCP endpoint: [Blaxel Sandboxes](https://www.anchorterminal.com/tools/blaxel-sandboxes.md), remote MCP server, nothing to install. - Self-hosting under an open licence: [E2B](https://www.anchorterminal.com/tools/e2b.md), self-hosted, Apache-2 licence. - The review panel's favourite: [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md), 3.3/5 from 8 panel reviews. ## How to choose - Start time and concurrency: Check the start time and how many sandboxes you can run at once, because an agent that waits on cold starts or hits a concurrency cap stalls mid-task. - Isolation and what it covers: Check what isolation the vendor documents and whether it covers the network, filesystem and host, because an agent running untrusted code needs a boundary it can rely on. - Billing per second and idle time: Work out the cost of a sandbox that sits paused or idle, since a plan may charge for memory or storage while nothing runs. - Pause, resume and checkpoints: Check whether state can be paused, resumed or checkpointed and what survives a restore, because a restore that drops installed packages forces the agent to start again. - How the benchmark tests this category: The same task in every sandbox: start, install a package, run a script that writes files, pause and resume where supported, then tear down. We time each step, check isolation claims against the docs and add up the cost. Each listing's verdict, strengths and weaknesses: https://www.anchorterminal.com/best/code-sandboxes/index.md