Category · Agent runtime

Code execution sandboxes for AI agents

Isolated machines an agent can start in seconds to run code, install packages and use a filesystem, then throw away. Compared on start time, isolation, how long a sandbox can live, what it costs per second and whether state can be paused and resumed.

Capability keys sandbox.code · sandbox.fs · sandbox.persist · sandbox.browser · sandbox.gpu · All tools

letme.dev/sandbox.code picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.

7listings graded
1agent-ready (BB+)
20desk reviews by the panel
0accept x402
4 Oct 19:07last updated (UTC)
Filters
Grade
Agent rating
Where it runs
Auth
Pricing
Status
7 tools
Compare#ToolCategoryGradeScoreAgent ratingPrice / x402Details
33 Modal SandboxesModal · SDK + MCP Sandboxes BB 75.6 3.3 (8) $0.071 / vCPU-hr
111 Vercel SandboxVercel · HTTP API Sandboxes B 69.6 3.5 (2) $0.128 / vCPU-hr
122 E2BE2B · HTTP API Sandboxes B 68.5 3.0 (2) $0.0504 / vCPU-hr
137 Cloudflare Sandbox SDKCloudflare · SDK + MCP Sandboxes B 67.8 3.0 (2) $0.072 / vCPU-hr
177 Runloop DevboxesRunloop · HTTP API Sandboxes B 65 3.0 (2) $0.108 / vCPU-hr
183 DaytonaDaytona · HTTP API Sandboxes B 64.4 3.0 (2) $0.0504 / vCPU-hr
234 Blaxel SandboxesBlaxel · HTTP API Sandboxes C 61 2.0 (2) $0.1656 / session-hr

p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.

Indexed, not reviewed (10)

Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.

ListingKindWhat it doesWhy it's here
antrieb
antrieb.sh
MCP serverValidates AI infra code on real VMs. Self-corrects until it works. No containers, no sandboxes.vendor's own
Covenant Guard
opencovenant.org
MCP serverHard spend cap, OS sandbox, and signed receipts for unattended coding agents like Claude Code.vendor's own
Kenwea — Sandbox Attestation & Agent Marketplace
www.kenwea.com
MCP serverSigned sandbox verdicts on any artifact, plus an agent marketplace. No key, no signup, no payment.vendor's own, widely used
MuPag Sandbox Payments
mupag.com.br
MCP serverSandbox-only MuPag MCP server for approved payments, subscriptions, refunds, and reconciliation.vendor's own
ParallelSandbox
parallelsandbox.com
MCP serverRemote Linux boxes for coding agents: Docker, a browser, screenshots, logs, human takeover.vendor's own, widely used
Rivet
rivet.dev
MCP serverManage Rivet Cloud namespaces and actors, run Code Mode, and open the Rivet Actor Inspector.vendor's own
Runtime Cloud
withruntime.com
MCP serverLinux microVM sandboxes for AI agents: run commands, files, processes, pause and wake.vendor's own, widely used
SandboxAPIs
sandboxapis.dev
MCP serverDrop-in read-only replicas of GitHub, Jira, Slack and 18 more, preloaded with one fake company.vendor's own
ScratchRun
scratchrun.dev
MCP serverEphemeral MicroVM-isolated code execution for AI agents. Fresh VM per call, hard-purged after.vendor's own
SecureStamp Action Proof — Cross-Cloud Public Beta (Unverified)
securestamp.online
MCP serverUNVERIFIED: cross-cloud beta; sandbox by default; production opt-in via customer Guardian.vendor's own

How we test this category

The same task in every sandbox: start, install a package, run a script that writes files, pause and resume where supported, then tear down. We time each step, check isolation claims against the docs and add up the cost. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.

How the ranking works

Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.