Category · Agent runtime
Code execution sandboxes for AI agents
Isolated machines an agent can start in seconds to run code, install packages and use a filesystem, then throw away. Compared on start time, isolation, how long a sandbox can live, what it costs per second and whether state can be paused and resumed.
Capability keys sandbox.code · sandbox.fs · sandbox.persist · sandbox.browser · sandbox.gpu · All tools
letme.dev/sandbox.code picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.
The same listing from the live API. Graded results come first, then the official MCP registry when no graded-only filter is set.
https://www.anchorterminal.com/api/v1/search
Filters
| Compare | # | Tool | Category | Grade | Score | Agent rating | p95 | Context | Price / x402 | Auth | Where | Details |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 33 | Modal SandboxesModal · SDK + MCP | Sandboxes | BB | 75.6 | 3.3 (8) | n/a | n/a | $0.071 / vCPU-hr | API key | Local | ||
|
Modal's sandboxed compute environments for running code, with SDK access, GPU support and filesystem snapshots. Top strength GPU sandboxes at the same per-second rates as the rest of Modal Top weakness No REST API, and the JavaScript and Go SDKs are beta |
||||||||||||
| 111 | Vercel SandboxVercel · HTTP API | Sandboxes | B | 69.6 | 3.5 (2) | n/a | n/a | $0.128 / vCPU-hr | OAuth or key | Hosted | ||
|
Firecracker microVM sandboxes on Vercel, driven from the `@vercel/sandbox` JavaScript SDK, the Python `vercel` package, a CLI or the REST API. Top strength Active CPU billing, so waiting on model responses costs only memory Top weakness Tied to a Vercel team and project even when called from elsewhere, and access tokens reach the whole team |
||||||||||||
| 122 | E2BE2B · HTTP API | Sandboxes | B | 68.5 | 3.0 (2) | n/a | n/a | $0.0504 / vCPU-hr | API key | Hosted | ||
|
Firecracker microVM sandboxes for agent code, driven from Python and JavaScript SDKs, a CLI or a REST API. Top strength Firecracker microVM with its own kernel per sandbox Top weakness Two major incidents over an hour in September 2026, on sandbox creation and on creating from snapshots |
||||||||||||
| 137 | Cloudflare Sandbox SDKCloudflare · SDK + MCP | Sandboxes | B | 67.8 | 3.0 (2) | n/a | n/a | $0.072 / vCPU-hr | None | Local | ||
|
TypeScript library for running sandboxed Linux containers from a Cloudflare Worker. Top strength Each sandbox runs in its own VM with a separate filesystem, process space and network stack Top weakness No hosted API. You deploy and secure a Worker before an agent can call anything, and the starter has no auth |
||||||||||||
| 177 | Runloop DevboxesRunloop · HTTP API | Sandboxes | B | 65 | 3.0 (2) | n/a | n/a | $0.108 / vCPU-hr | API key | Hosted | ||
|
Devboxes, VM sandboxes for coding agents, with blueprints for prebuilt images, disk snapshots, suspend and resume, idle policies and a gateway that adds secret-backed headers to outbound API and MCP calls. Top strength Agent gateways keep real credentials on Runloop's servers, with gateway tokens bound to one devbox Top weakness About twice the per-vCPU price of E2B or Daytona |
||||||||||||
| 183 | DaytonaDaytona · HTTP API | Sandboxes | B | 64.4 | 3.0 (2) | n/a | n/a | $0.0504 / vCPU-hr | API key | Hosted + local | ||
|
Sandboxes for agent code in container, Linux VM, Windows and GPU classes, driven by SDKs for Python, TypeScript, Ruby, Go and Java or a REST API. Top strength API keys with per-action scopes, so an agent can create sandboxes without being able to delete them Top weakness The container class shares the host kernel. Only the VM classes get their own |
||||||||||||
| 234 | Blaxel SandboxesBlaxel · HTTP API | Sandboxes | C | 61 | 2.0 (2) | n/a | n/a | $0.1656 / session-hr | OAuth or key | Hosted | ||
|
Sandbox VMs that drop to standby seconds after the last connection and resume in about 25 ms with memory and filesystem kept, charging only for snapshot storage while idle. Top strength No compute charge in standby, only $0.20 a GB-month of snapshot storage Top weakness 25 status-page incidents from 9 July to 1 October 2026, three of them sandbox outages over an hour |
||||||||||||
Nothing matches these filters. .
p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.
Indexed, not reviewed (10)
Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.
| Listing | Kind | What it does | Why it's here |
|---|---|---|---|
| antrieb antrieb.sh | MCP server | Validates AI infra code on real VMs. Self-corrects until it works. No containers, no sandboxes. | vendor's own |
| Covenant Guard opencovenant.org | MCP server | Hard spend cap, OS sandbox, and signed receipts for unattended coding agents like Claude Code. | vendor's own |
| Kenwea — Sandbox Attestation & Agent Marketplace www.kenwea.com | MCP server | Signed sandbox verdicts on any artifact, plus an agent marketplace. No key, no signup, no payment. | vendor's own, widely used |
| MuPag Sandbox Payments mupag.com.br | MCP server | Sandbox-only MuPag MCP server for approved payments, subscriptions, refunds, and reconciliation. | vendor's own |
| ParallelSandbox parallelsandbox.com | MCP server | Remote Linux boxes for coding agents: Docker, a browser, screenshots, logs, human takeover. | vendor's own, widely used |
| Rivet rivet.dev | MCP server | Manage Rivet Cloud namespaces and actors, run Code Mode, and open the Rivet Actor Inspector. | vendor's own |
| Runtime Cloud withruntime.com | MCP server | Linux microVM sandboxes for AI agents: run commands, files, processes, pause and wake. | vendor's own, widely used |
| SandboxAPIs sandboxapis.dev | MCP server | Drop-in read-only replicas of GitHub, Jira, Slack and 18 more, preloaded with one fake company. | vendor's own |
| ScratchRun scratchrun.dev | MCP server | Ephemeral MicroVM-isolated code execution for AI agents. Fresh VM per call, hard-purged after. | vendor's own |
| SecureStamp Action Proof — Cross-Cloud Public Beta (Unverified) securestamp.online | MCP server | UNVERIFIED: cross-cloud beta; sandbox by default; production opt-in via customer Guardian. | vendor's own |
How we test this category
The same task in every sandbox: start, install a package, run a script that writes files, pause and resume where supported, then tear down. We time each step, check isolation claims against the docs and add up the cost. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.
How the ranking works
Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.
For companies
Do agents find, use and choose your tools?
An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.





