# Fix list: Together Code Sandbox From Anchor Terminal's listing at https://www.anchorterminal.com/tools/together-code-sandbox, the October 2026 research run, assessed 8 October 2026. Grade D, 53.6 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on Together Code Sandbox: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Reliability, 25 out of 100, up to 15 more on the total Why it scored 25: Read with the hosted lines. status.together.ai lists the website, the playground and inference models, and status.codesandbox.io lists the CodeSandbox API, website, editor, CI and two clusters. Neither names the sandbox management API at `api.bartender.codesandbox.io`, so half credit (10). No incident history for this service is readable on either page (5). No rate or concurrency limits found, and the migration guide says the concurrency count and limit are not exposed (0). Both SDKs retry 408, 429, 500, 502, 503 and 504 three times with exponential backoff from 0.5 seconds plus jitter, and the docs warn that `snapshots.create` is not idempotent. No `Retry-After` or idempotency keys found (10). No SLA found, and the terms make no uptime guarantee outside an order form (0). The SDK was announced on 1 October 2026 for organisations on an allowlist, which is not general availability (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 2. Payments & pricing, 20 out of 100, up to 10 more on the total Why it scored 20: No x402, MPP or L402 found (0). The pricing page lists Code Sandbox at $0.0446 per vCPU-hour and $0.0149 per GiB of RAM per hour without a login. It does not say whether the rate covers the new SDK as well as the legacy one (20). The billing docs say Together has no free trial and that platform access needs a $15 credit purchase (0). Access needs a browser signup, a payment and a request to Together to enable the organisation (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 3. Security & auth, 47 out of 100, up to 9.3 more on the total Why it scored 47: Project-scoped API keys, revocable, with optional expiry and no finer scopes. The docs say a key has full access to its project and can spend the credit balance (25). The in-VM document lists a `token` query parameter on the WebSocket exec path. We took 5 off, not the checklist's 10, because it is a per-sandbox agent token on one path and not the account key (-5). Each sandbox is a virtual machine. Every port is open to everyone unless an experimental inbound policy is set, no outbound rules exist, and project roles are admin and editor with no read-only role (8). The sandbox returns the output of code the agent ran, with no guidance on untrusted output found (8). Sandboxes can be listed with tags and `status_reason`. No audit log found (3). A signed security.txt valid to 30 September 2028 points to a private HackerOne programme. The trust centre is script-drawn and was unread (8). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 4. Agent ergonomics, 70 out of 100, up to 4.9 more on the total Why it scored 70: List calls take `limit` from 1 to 100, batch reads take IDs and exec output can be read from a `lastSequence`. No field selection (15). Cursor pagination on sandboxes and snapshots, with filters for status, source snapshot, tags and retired snapshots (20). Errors return `code`, `message` and per-parameter entries, the SDKs raise one `HttpError` with the status, and `status_reason` names causes such as `out_of_capacity` and `oom_killed`. No catalogue of codes found (14). Built-in retries, but no idempotency keys, sandbox IDs can't be chosen and `snapshots.create` can register duplicates (8). Python and TypeScript SDKs and a CLI, with defaults of 1 vCPU and 2 GiB. A snapshot must be built first because no default image exists (13). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 5. Transparency & trust, 60 out of 100, up to 3.5 more on the total Made of editorial 43, provenance 77. Why it scored 60: A closed service under Together's terms of service (19 May 2026), with the SDKs, CLI and OpenAPI documents under MIT (20). The privacy policy (17 December 2025) says data is not used for training without opt-in. Its table of data types covers inference, fine-tuning and GPU clusters with no sandbox row, and no retention period for sandbox filesystems or snapshots was found beyond an optional snapshot `ttl` (12). The legacy `@codesandbox/sdk` is marked deprecated with no end date, the deprecations page covers models only, and breaking SDK changes are marked in the changelog (8). No sub-processor list or data locations for sandboxes found. A docs example names a cluster `na-us-ce-01`, and the trust centre was unread (3). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Domain age: together.ai, registered 2017-12-16 (8 years) (11 of 15) - Endpoint on the vendor's domain: api.bartender.codesandbox.io is not on together.ai (0 of 15) - Terms of service: read, states 5 of the 7 things a reader expects, and has 1 clause that costs points (6.3 of 10) - Privacy policy: read, states 7 of the 8 things a reader expects (9.3 of 10) ## 6. Schema & documentation, 86 out of 100, up to 2.3 more on the total Why it scored 86: Two public OpenAPI documents in the repository and on GitHub Pages, the management API (3.1.0, 17 operations) and the in-VM API (3.0.3, 25 operations) (25). docs.together.ai has llms.txt and a Markdown twin of the page, and each package ships an `LLMS.md` with the references (10). The concept docs say when to snapshot memory rather than disk only and what ephemeral means, while 36 of 114 typed nodes in the management document carry a description (14). Enums for status, bounds on `cpu`, `memory_bytes`, `ttl` and `limit`, with `experimental` left loose (12). Examples in Python, TypeScript and the CLI, and an error object with `code`, `message` and `errors`. No list of error codes found, and 429 is not in the document (10). `/v1` paths, semver packages and a generated `CHANGELOG.md` (15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 7. Maintenance & community, 83 out of 100, up to 1.5 more on the total Why it scored 83: `together-sandbox` 4.1.1 on npm and PyPI on 7 October 2026 (30). Twelve tagged releases since 10 July 2026 (20). The repository has 3 stars and 6 open items, all pull requests from the team. The 30 most recent items are pull requests merged within days, so replies to outside reports can't be judged (12). Current official Python and TypeScript SDKs and a CLI, versioned together (15). Pull requests run the TypeScript unit tests, end-to-end tests run on pushes to main, and packages publish by OIDC trusted publishing. The Python tests are not in the pull request workflow, and we did not read the run results (6). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - How long an access request takes and what Together asks of an organisation before enabling the SDK. The docs only say to contact Together. - Whether the $0.0446 per vCPU-hour and $0.0149 per GiB-hour on the pricing page apply to the new SDK. The page lists them under Code Sandbox without naming either SDK, and a pull request adding billing usage to the SDKs was open on 8 October 2026. - The isolation technology. The repository docs call a sandbox a virtual machine and the legacy page says microVM, and neither names the hypervisor for the new service. - Rate limits, concurrency limits and maximum sandbox lifetime. None found. - unchecked: trust.together.ai is script-drawn and showed our reader only its title, so certifications such as SOC 2 are unconfirmed. - unchecked: GitHub Actions run results for the repository. We read the workflow files in the clone, not the runs. - unchecked: a DPA or sub-processor list. The two paths we tried on www.together.ai returned 404 and no link was found on the terms or privacy pages. - The lead described microVMs with forking. The new SDK has no live fork (the parent must terminate with a snapshot first), and the product page at together.ai/sandbox still shows the legacy `@codesandbox/sdk` and its 2 to 64 vCPU sizes, against 0.1 to 16 in the new SDK. ## Weaknesses - The SDK and CLI work only for organisations Together has enabled, and the docs say to contact Together for access - No rate limits, concurrency limits or SLA found, and neither status.together.ai nor status.codesandbox.io names the service - Every sandbox port is public unless an experimental inbound policy is set, and no outbound controls exist yet - Three major versions between 2 June and 28 July 2026, and memory snapshots were removed in 4.0.0 and restored in 4.0.4 - The management API answers at `api.bartender.codesandbox.io`, off Together's domain, and the privacy policy's data table has no sandbox row ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Confirm the organisation is on Together's allowlist before installing. A valid `TOGETHER_API_KEY` alone is not enough - Build a snapshot first with `snapshots.create`. No default image exists, and every sandbox needs `snapshot_id` or `snapshot_alias` - Set `ttl` at creation or call `terminate()`. Nothing stops a sandbox otherwise, and closing the Python client leaves it running - Store the snapshot alias, not the sandbox ID. A terminated sandbox can't restart, and its state lives at `sandbox:` - Exclude `snapshots.create` from retries with `should_retry`, and set `experimental.network_policy` before serving anything private on a port ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.