Best of · Agent runtime

Best code execution sandboxes for AI agents

The 10 highest-scoring of 15 code execution sandboxes on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.

  • 15 ranked
  • 3 agent-ready
  • 10 hosted endpoints
  • Updated 8 October 2026

Top three

Picks by need

Worked out from the scores, prices and facts, so they change when the research does.

Highest score overall

Microsoft Execution Containers BB

BB, 76.3/100 on the benchmark.

Also Modal Sandboxes, BB, 75.5/100.

Reliability

Modal Sandboxes BB

95/100 on reliability, against 81 for the overall leader.

Schema & documentation

E2B B

92/100 on schema & documentation, against 81 for the overall leader.

Agent ergonomics

Amazon Bedrock AgentCore Code Interpreter BB

75/100 on agent ergonomics, against 74 for the overall leader.

Security & auth

Amazon Bedrock AgentCore Code Interpreter BB

87/100 on security & auth, against 69 for the overall leader.

Maintenance & community

Modal Sandboxes BB

93/100 on maintenance & community, against 92 for the overall leader.

Lowest paid price per vCPU hour

Agent 37 Cloud D

$0.0011 per vCPU hour, the lowest of the 13 listings here with a paid price in this unit (free allowances aside).

Also Sprites, $0.0385 per vCPU hour.

A hosted MCP endpoint

Blaxel Sandboxes C

remote MCP server, nothing to install.

Self-hosting under an open licence

E2B B

self-hosted, Apache-2 licence.

The review panel's favourite

Modal Sandboxes BB

3.3/5 from 8 panel reviews.

The shortlist

#ToolGradeBest forPriceWhere
1 Microsoft Execution Containers
Microsoft
BB 76.3 A developer building an agent or tool host that must run model-written code on the user's own machine, above all on Windows, where it reaches Microsoft's process and session isolation. Free · OSS local
2 Modal Sandboxes
Modal
BB 75.5 GPU work inside a sandbox, or agents already running on Modal. $0.071 / vCPU-hr local
3 Amazon Bedrock AgentCore Code Interpreter
Amazon Web Services
BB 73.1 A team already on AWS that wants agent code to run under IAM, inside a VPC and beside S3 data, for sessions up to eight hours. $0.0895 / vCPU-hr hosted
4 Vercel Sandbox
Vercel
B 69.6 Agents that spend most of their time waiting on a model, and teams already on Vercel. $0.128 / vCPU-hr hosted
5 E2B
E2B
B 68.3 Code interpreters and agent workspaces that pause and resume with memory, and teams that may want to self-host later. $0.0504 / vCPU-hr hosted
6 Cloudflare Sandbox SDK
Cloudflare
B 67.5 Agents already built on Workers and Durable Objects that want sandboxes in the same account with credential injection at the edge. $0.072 / vCPU-hr local
7 Runloop Devboxes
Runloop
B 64.8 Running and grading coding agents on full workstations with prebuilt blueprints and credential gateways. $0.108 / vCPU-hr hosted
8 Daytona
Daytona
B 64.3 Agents that need a choice of machine, including Windows desktops and GPUs, and operators who want least-privilege keys. $0.0504 / vCPU-hr hosted and local
9 Blaxel Sandboxes
Blaxel
C 60.7 Long-lived, mostly idle agent workspaces that need to resume quickly with memory intact, and teams that want an MCP server per sandbox. $0.1656 / session-hr hosted
10 Freestyle
Freestyle Cloud Inc.
C 58.5 Suited to long-running agent work that needs a full Linux VM, memory-preserving pause, snapshots to branch from and private networking. $0.0403 / vCPU-hr hosted

5 more are ranked in the full table.

How to choose

  1. Start time and concurrencyCheck the start time and how many sandboxes you can run at once, because an agent that waits on cold starts or hits a concurrency cap stalls mid-task.
  2. Isolation and what it coversCheck what isolation the vendor documents and whether it covers the network, filesystem and host, because an agent running untrusted code needs a boundary it can rely on.
  3. Billing per second and idle timeWork out the cost of a sandbox that sits paused or idle, since a plan may charge for memory or storage while nothing runs.
  4. Pause, resume and checkpointsCheck whether state can be paused, resumed or checkpointed and what survives a restore, because a restore that drops installed packages forces the agent to start again.

How the benchmark tests this category. The same task in every sandbox: start, install a package, run a script that writes files, pause and resume where supported, then tear down. We time each step, check isolation claims against the docs and add up the cost.

Each one in detail

#1

Microsoft Execution Containers

BB 76.3/100

Microsoft Execution Containers (MXC) is an open-source SDK for running untrusted code in a local sandbox on Windows, Linux and macOS. An application embeds it through Node.js, .NET or Rust and sets filesystem, network and UI policy for each run.

Verdict MXC puts nine operating-system sandbox backends behind one typed request, with network access denied by default and a JSON Schema for the stable 1.0.0 contract. Version 1.0.0 is two days old as of 8 October 2026. Enforcement varies by backend, and isolation_session cannot restrict networking at all.

Choose it for A developer building an agent or tool host that must run model-written code on the user's own machine, above all on Windows, where it reaches Microsoft's process and session isolation.

Strengths

  • MIT licence, with SDKs for Node.js, .NET and Rust all at 1.0.0 and the native runtime bundled in the npm and NuGet packages
  • Egress, ingress and host loopback default to deny, and filesystem access is limited to listed read-only and read-write paths
  • A draft-07 JSON Schema for the stable 1.0.0 request, with descriptions on 135 of 150 properties

Weaknesses

  • 1.0.0 shipped on 6 October 2026, and the Node changelog still lists the V1 changes under Unreleased
  • Enforcement differs by backend. isolation_session cannot restrict networking, and proxy routing is cooperative on Seatbelt and WSLC
  • Persistent containers exist only for isolation_session and wslc, both on Windows

Price Free · OSSAuth Nonex402 nolocal

Full assessment

#3

Amazon Bedrock AgentCore Code Interpreter

BB 73.1/100

Amazon Bedrock AgentCore Code Interpreter is AWS's managed sandbox for running agent-written Python, JavaScript and TypeScript. Each session is a dedicated microVM, reached through the AWS API, the AgentCore SDKs or an MCP server.

Verdict Each session runs in its own microVM with 2 vCPU, 8 GB and a 10 GB disk for up to eight hours, and IAM can scope access to one interpreter. Sessions cannot be paused or resumed, and an AWS account with IAM set-up is needed before a first call.

Choose it for A team already on AWS that wants agent code to run under IAM, inside a VPC and beside S3 data, for sessions up to eight hours.

Strengths

  • Each session runs in a dedicated microVM, which AWS says is terminated and has its memory sanitised when the session ends
  • Billing is per second on CPU used and peak memory, at $0.0895 a vCPU-hour and $0.00945 a GB-hour, with I/O wait and idle time free
  • IAM actions cover each of the nine operations, with separate ARNs for the system interpreter and custom ones

Weaknesses

  • No pause, resume or snapshot. Session files are removed when the session ends, and persistence needs a customer-owned S3 Files or EFS mount inside a VPC
  • Every session is capped at 2 vCPU, 8 GB of memory and 10 GB of disk, and the cap is not adjustable
  • InvokeCodeInterpreter takes one flat arguments object for nine operations, so the reference does not say which fields each operation requires

Price $0.0895 / vCPU-hrAuth API keyx402 nohosted

Full assessment · Against #1, Microsoft Execution Containers

#4

Vercel Sandbox

B 69.6/100

Firecracker microVM sandboxes on Vercel, driven from the @vercel/sandbox JavaScript SDK, the Python vercel package, a CLI or the REST API.

Verdict Active CPU billing, so waiting on model responses costs only memory. Tied to a Vercel team and project even when called from elsewhere, and access tokens reach the whole team.

Choose it for Agents that spend most of their time waiting on a model, and teams already on Vercel.

Strengths

  • Active CPU billing, so waiting on model responses costs only memory
  • Credential brokering proxy outside the sandbox that overwrites headers set by sandbox code
  • Firecracker microVM with root access, and a firewall with deny-all, domain and CIDR rules

Weaknesses

  • Tied to a Vercel team and project even when called from elsewhere, and access tokens reach the whole team
  • Hobby caps sessions at 45 minutes and pauses creation once the monthly allowance is spent
  • Snapshots keep the filesystem, not memory or running processes

Price $0.128 / vCPU-hrAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Microsoft Execution Containers

#5

E2B

B 68.3/100

Firecracker microVM sandboxes for agent code, driven from Python and JavaScript SDKs, a CLI or a REST API.

Verdict Firecracker microVM with its own kernel per sandbox. Two major incidents over an hour in September 2026, on sandbox creation and on creating from snapshots.

Choose it for Code interpreters and agent workspaces that pause and resume with memory, and teams that may want to self-host later.

Strengths

  • Firecracker microVM with its own kernel per sandbox
  • Egress allow and deny lists by domain, IP or CIDR, and secrets filled in outside the sandbox
  • Apache-2.0 infrastructure in e2b-dev/infra, self-hostable with Terraform

Weaknesses

  • Two major incidents over an hour in September 2026, on sandbox creation and on creating from snapshots
  • One unscoped API key per project, with no audit log found
  • Hobby sandboxes stop after 1 hour of continuous running

Price $0.0504 / vCPU-hrAuth API keyx402 nohosted

Full assessment · Against #1, Microsoft Execution Containers

#6

Cloudflare Sandbox SDK

B 67.5/100

TypeScript library for running sandboxed Linux containers from a Cloudflare Worker.

Verdict Each sandbox runs in its own VM with a separate filesystem, process space and network stack. No hosted API. You deploy and secure a Worker before an agent can call anything, and the starter has no auth.

Choose it for Agents already built on Workers and Durable Objects that want sandboxes in the same account with credential injection at the edge.

Strengths

  • Each sandbox runs in its own VM with a separate filesystem, process space and network stack
  • Outbound handlers hold credentials in the Worker and inject them, so the container never sees them
  • Egress can be turned off or limited to a deny-by-default allowedHosts list

Weaknesses

  • No hosted API. You deploy and secure a Worker before an agent can call anything, and the starter has no auth
  • Open bugs on backups that drop directories (#859) and restores of archives of 10 MB or more (#884)
  • Needs Workers Paid at $5 a month, with no card-free route

Price $0.072 / vCPU-hrAuth Nonex402 nolocal

Full assessment · Against #1, Microsoft Execution Containers

#7

Runloop Devboxes

B 64.8/100

Devboxes, VM sandboxes for coding agents, with blueprints for prebuilt images, disk snapshots, suspend and resume, idle policies and a gateway that adds secret-backed headers to outbound API and MCP calls.

Verdict Gateway credentials remain on Runloop servers, with access tokens bound to one devbox. Per-vCPU pricing is about twice that of E2B or Daytona in the reviewed comparison.

Choose it for Running and grading coding agents on full workstations with prebuilt blueprints and credential gateways.

Strengths

  • Agent gateways keep real credentials on Runloop's servers, with gateway tokens bound to one devbox
  • Network policies that block egress or allow listed hostnames
  • Public OpenAPI, llms.txt and typed Python and TypeScript SDKs with cursor pagination and 429 backoff

Weaknesses

  • About twice the per-vCPU price of E2B or Daytona
  • No published rate limits or API error-body reference
  • Suspend keeps disk only, and processes need restarting after resume

Price $0.108 / vCPU-hrAuth API keyx402 nohosted

Full assessment · Against #1, Microsoft Execution Containers

#8

Daytona

B 64.3/100

Sandboxes for agent code in container, Linux VM, Windows and GPU classes, driven by SDKs for Python, TypeScript, Ruby, Go and Java or a REST API.

Verdict API keys with per-action scopes, so an agent can create sandboxes without being able to delete them. The container class shares the host kernel. Only the VM classes get their own.

Choose it for Agents that need a choice of machine, including Windows desktops and GPUs, and operators who want least-privilege keys.

Strengths

  • API keys with per-action scopes, so an agent can create sandboxes without being able to delete them
  • Container, Linux VM, Windows and GPU sandbox classes behind one API
  • Three public OpenAPI files, llms.txt and SDKs in five languages

Weaknesses

  • The container class shares the host kernel. Only the VM classes get their own
  • Full internet access and per-sandbox allow lists need Tier 3
  • Platform code went private in June 2026, and the old repository's 311 open issues won't be answered

Price $0.0504 / vCPU-hrAuth API keyx402 nohosted and local

Full assessment · Against #1, Microsoft Execution Containers

#9

Blaxel Sandboxes

C 60.7/100

Sandbox VMs that drop to standby seconds after the last connection and resume in about 25 ms with memory and filesystem kept, charging only for snapshot storage while idle.

Verdict No compute charge in standby, only $0.20 a GB-month of snapshot storage. 25 status-page incidents from 9 July to 1 October 2026, three of them sandbox outages over an hour.

Choose it for Long-lived, mostly idle agent workspaces that need to resume quickly with memory intact, and teams that want an MCP server per sandbox.

Strengths

  • No compute charge in standby, only $0.20 a GB-month of snapshot storage
  • MicroVM per sandbox with domain allow and deny lists that can be enforced at network level
  • An MCP server in every sandbox over streamable HTTP, 18 tools

Weaknesses

  • 25 status-page incidents from 9 July to 1 October 2026, three of them sandbox outages over an hour
  • No published request-rate limits, 429 guidance or SLA
  • Domain filtering and secret injection are marked public preview, and egress is open by default

Price $0.1656 / session-hrAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Microsoft Execution Containers

#10

Freestyle

C 58.5/100

Freestyle runs full Linux virtual machines for AI agents, with pause and resume that keeps memory, snapshots, private networks and domains. Agents drive it through a REST API, a TypeScript SDK and a CLI from one npm package.

Verdict A public OpenAPI 3.1 file describes 88 operations with a stable error code on every failure, and per-VM identity tokens keep the team API key away from end users. No terms of service, product privacy policy, request rate limit or incident history was found, and the only SDK is TypeScript.

Choose it for Suited to long-running agent work that needs a full Linux VM, memory-preserving pause, snapshots to branch from and private networking.

Strengths

  • Public OpenAPI 3.1 file with 88 /v5 operations, each with a description and named error codes per status
  • A new VM has no inbound or outbound network access until firewall rules allow it
  • Identity access tokens limited to granted VMs and Linux users, revocable without rotating the team API key

Weaknesses

  • No terms of service or service agreement found on the site, and the privacy policy covers the website only
  • The status page shows three components with no incident history
  • No request rate limit, Retry-After guidance or idempotency keys found in the docs or the OpenAPI file

Price $0.0403 / vCPU-hrAuth API keyx402 nohosted

Full assessment · Against #1, Microsoft Execution Containers

Head to head

All 105 comparisons in this category

Questions

What are the highest-rated code execution sandboxes for AI agents?

Microsoft Execution Containers has the highest benchmark score of the 15 ranked code execution sandboxes, 76.3 (BB). Modal Sandboxes is second with 75.5 (BB).

How many code execution sandboxes are agent-ready?

3 of the 15 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.

Which code execution sandboxes accept x402 payments?

None of the ranked listings here accepts x402 for its main call yet.

Which of these code execution sandboxes is cheapest?

By published paid prices, Agent 37 Cloud, at $0.0011 per vCPU hour, the lowest of the 13 listings here with a paid price in this unit (free allowances aside). Plans, volume tiers and free allowances change the sum, so check the listing's price table.

How is this list ranked?

By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 8 October 2026.

How this list is made

The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.

Full ranked table · 105 head-to-head comparisons · Best tools in every category

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.