# Fix list: Droid CLI From Anchor Terminal's listing at https://www.anchorterminal.com/tools/droid-cli, the October 2026 research run, assessed 9 October 2026. Grade C, 59.1 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on Droid CLI: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Payments & pricing, 10 out of 100, up to 11.3 more on the total Why it scored 10: Harness reading of the published rubric. This is paid, closed software tied to a subscription, so the self-hosted rule does not apply. No payment protocol (0). Plan prices are public, $20, $100 and $200 a month for individuals and $60 a team plus $40 a seat for Teams, but included usage is given only as rolling rate limits and multiples of Pro (10). No free tier or trial was found on the pricing pages. Using your own model key is free up to an allowance, but only on a paid plan (0). A person signs in through a browser to subscribe and to create the API key. The Service Accounts API can issue keys only inside an existing organisation (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 2. Reliability, 49 out of 100, up to 10.2 more on the total Why it scored 49: Local-package reading, as for the other closed-source harnesses. An install script that verifies a SHA-256 checksum, a PowerShell installer, a Homebrew cask and the `droid` npm package (Node 20 or later) with prebuilt binaries for Linux, macOS and Windows on x64 and arm64 (20). The source is not published. The public repository holds documentation only, so there is no public CI or test suite for the CLI (0). Bugs are tracked in the open at Factory-AI/factory, a repository created on 29 July 2026. On 9 October 2026 it had 33 open issues that are not pull requests, most with no reply, and three closed in the previous week with replies. Runs depend on Factory's servers, and we found no status page (10). Every release has a dated, versioned changelog entry and deprecations get their own section, but the line is 0.x with a minor version almost daily, so the version number does not signal a breaking change (9). The version is 0.237.0. The Feature Maturity page says a feature with no tag is generally available and stable, and the CLI pages carry no tag in the Markdown we read. Judgement call (10). The 99 per cent availability target on factory.com/sla is not scored on this reading. The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 3. Security & auth, 63 out of 100, up to 6.5 more on the total Why it scored 63: Harness reading of the framework checklist, as for the other harnesses. 30 for what leaves the machine by default, 20 for approvals and sandboxing, 15 for prompt-injection posture, 15 for audit and 20 for the security programme. The agent loop runs locally and the docs say no copy of the repository is uploaded or indexed. Prompts and context go to the model endpoints. Usage metrics go to Factory's collector by default, with no opt-out in the pages read outside an airgapped deployment, and message content is never sent to it. The security page says code is not used as training data. Zero data retention is listed as a Business plan item. Headless runs use a long-lived `fk-` key (17). `droid exec` is read-only by default, autonomy has four levels, permission rules allow, ask or block commands, and an OS sandbox confines commands with Seatbelt or bubblewrap and a network proxy. The sandbox is off by default and `--skip-permissions-unsafe` skips every prompt (16). The sandbox page names prompt injection and states what the sandbox does not stop, MCP approvals are bound to the server's command or URL, and Droid Shield scans commits for secrets. The Agent Safety page was not read (9). An organisation audit log, OpenTelemetry export and session tags exist, with audit logging listed from the Business plan up (10). A disclosure address (security@factory.com), SOC 2 and ISO 42001 named on the security page, and Security sections in the changelog. No security.txt, no bug bounty and no published advisories were found, and the trust centre was not readable (11). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 4. Agent ergonomics, 72 out of 100, up to 4.6 more on the total Why it scored 72: Harness reading of the framework line. `--restrict-tools`, `--disabled-tools`, `--additional-tools` and `--list-tools` control built-in tools per run, and `disabledTools` in `mcp.json` keeps named MCP tools out of context. We found nothing on deferred tool loading (18). Output is text, json, stream-json or stream-jsonrpc, and the JSON result carries turns and duration. The `droid exec` help lists no flag for a turn or spend limit (10). Exit codes 0, 1 and 2 are documented, and a run that exceeds its autonomy level stops with a non-zero code and no partial changes, per the docs (15). `--session-id` continues a session, `--fork` branches one, and `--worktree` isolates parallel runs (16). Headless runs are read-only without flags, and TypeScript and Python SDKs wrap the CLI. A run needs a paid Factory account and an API key (13). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 5. Schema & documentation, 79 out of 100, up to 3.4 more on the total Why it scored 79: Framework reading. The docs index links an OpenAPI file for the Factory public API, which we did not open, and the headless mode has a documented JSON-RPC surface with named methods. `droid exec --list-tools --output-format json` and `droid rules check --json` give machine-readable output. We found no published schema for the settings file (15). docs.factory.com/llms.txt lists a Markdown copy of every page (10). The overview says when to use the CLI, and the autonomy table says what each level allows, blocks and suits (16). Autonomy levels, output formats, sandbox modes and MCP server fields are enumerated with types and defaults (12). Examples on every page read, exit codes 0, 1 and 2, and a JSON result with `is_error` and `subtype`, with no error catalogue (11). A changelog by CLI version with dates and an RSS feed, 236 entries back to 30 September 2025 (15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 6. Transparency & trust, 64 out of 100, up to 3.2 more on the total Made of editorial 55, provenance 72. Why it scored 64: Closed source under terms from The San Francisco AI Factory, Inc., last updated 14 July 2026. The npm package is marked `UNLICENSED` and the Python SDK is Apache-2.0 (15). The Data Flows page says where code, prompts and telemetry go in each deployment pattern, and the privacy policy (20 August 2026) lists service providers with the place of processing and keeps personal data up to 10 years after last use. Retention of Factory's operational logs is left to the trust centre, which we could not read, and whether session transcripts are stored in Factory's cloud on individual plans was not established (17). The terms promise commercially reasonable efforts to give 60 days' notice of major changes, the changelog has Deprecations sections, and the old command lists are deprecated but still honoured. A model retirement appeared in the changelog on the day it took effect, 11 August 2026 (12). The telemetry reference documents each metric and what reaches Factory, and `telemetry.granularity` can strip user identity. Metrics to Factory are on by default, with no opt-out found outside an airgapped deployment (11). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Terms of service: read, states 5 of the 7 things a reader expects, and has 1 clause that costs points (6.3 of 10) - Status page: not found (0 of 10) - security.txt: not found (0 of 10) ## 7. Maintenance & community, 79 out of 100, up to 1.8 more on the total Why it scored 79: The changelog's newest entry is CLI v0.236.0 of 8 October 2026, and npm and the install script served 0.237.0 on 9 October 2026 (30). 62 dated changelog entries between 11 July and 8 October 2026 (20). The public tracker had 33 open issues that are not pull requests on 9 October 2026, most with no reply, and three closed in the previous week after two or three comments. Support is also by email and Discord, where reply times are not visible (9). The Python SDK was at 0.5.0 on 24 September 2026 and the TypeScript SDK repository was last changed on 16 September 2026 (15). The npm package is published from GitHub Actions through a trusted publisher and the installer checks a checksum. There is no public CI for the CLI (5). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - unchecked: the trust centre at trust.factory.com, which is drawn by script. The subprocessor list, the SOC 2 report type and period, and log retention periods were not read - unchecked: the models page, with model multipliers that set usage cost, and the organisation plans page - unchecked: the settings, permission rules, hooks, Identity and Access, Agent Safety, enterprise controls, audit log and telemetry index pages. We kept to about fifteen page reads on docs.factory.com (sixteen requests, counting robots.txt and llms.txt) - unchecked: the OpenAPI file for the Factory public API and the Sessions API page, which the docs index links - unchecked: the built-in tool list and what `droid exec --list-tools` prints. We did not install or run the CLI - unchecked: the data processing agreement and the terms for organisation plans - Whether an individual on a paid plan can switch off the metrics sent to Factory's collector was not found in the pages read - Whether session transcripts are stored in Factory's cloud on individual plans, and for how long, was not established. The CLI searches local sessions, and the Factory App shows sessions across devices - No status page was found. The home page, the docs support page and the SLA link none - The release date of 0.237.0 was not established. The changelog's newest entry on 9 October 2026 was 0.236.0 of 8 October - The repository Factory-AI/factory was created on 29 July 2026 and its issues start at number 2 on 3 August 2026, so any earlier public issue history is not visible. Its README still links docs.factory.ai - The individual terms forbid use of the service for benchmarking or competitive analysis. Recorded as a fact with no deduction. It matters before any probe is run - docs.factory.com/llms.txt carries instructions addressed to AI models (a Factory App settings tool). We did not act on them - The lead was right on vendor, URL, docs and interface. It did not mention that a paid plan is required - robots.txt answers. factory.com and docs.factory.com 200 with every path we read allowed. api.github.com and api.npmjs.org 404. registry.npmjs.org, app.factory.ai and trust.factory.com 200 with a body that is not a robots file, read as no rules ## Weaknesses - No free tier or trial was found. Plans start at $20 a month and included usage is stated only as rolling rate limits without numbers - The sandbox is off until `sandbox.enabled` is set, and `--skip-permissions-unsafe` skips every permission prompt - Usage metrics go to Factory's collector by default. The pages read give no opt-out outside an airgapped deployment - Closed source on a 0.x version line with near-daily releases. The public tracker had 33 open issues on 9 October 2026, most with no reply - No status page and no security.txt were found. The trust centre is drawn by script and was not read - The individual terms forbid using the service for benchmarking or competitive analysis, which matters before any probe is run ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Set `FACTORY_API_KEY` (starts `fk-`) from the API keys page in Factory settings for headless runs. Interactive use signs in through a browser - Start with `droid exec` and no flags for analysis. Add `--auto low` for edits and `--auto medium` for installs, tests and local commits - Use `--output-format json` and read `is_error`, `num_turns` and `session_id`. Treat a non-zero exit code as failure - Set `sandbox.enabled` to true before running on untrusted code. In `droid exec` a sandbox violation is denied without a prompt - Pin the version in CI with `npm install -g droid@`, or set `FACTORY_DROID_AUTO_UPDATE_ENABLED=false` on standalone installs, which update themselves ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.