# Fix list: Statsig From Anchor Terminal's listing at https://www.anchorterminal.com/tools/statsig, the October 2026 research run, assessed 8 October 2026. Grade BB, 72.3 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on Statsig: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Payments & pricing, 40 out of 100, up to 7.5 more on the total Why it scored 40: No x402, MPP or L402 (0). Prices are public without a login. Pro is $150 a month with 5 million events, then $0.05 per 1,000 events (20). The Developer plan is free with no card and 2 million events a month (20). A person signs up in a browser, then creates a key or approves OAuth (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 2. Agent ergonomics, 71 out of 100, up to 4.7 more on the total Why it scored 71: The MCP V3 server shows 18 default tools, which is 15 on the checklist. Add back 8 for the discovery layer that reaches other operations on demand, `query_fields` on experiment reads and read-only access (23 of 25). `limit` and `page` on 49 list operations with pagination metadata, a cursor on the Logs Explorer query, and filters by tag, status, creator and date (17 of 20). Errors are a `status` and a `message` with no machine code, and MCP results set `isError` (11 of 20). No idempotency keys were found. Updates and `api_destructive` calls need a `confirmation` object, and operations are split into read, write and destructive lanes. We could not read the live tool annotations (10 of 20). The SDKs cover flag evaluation and event logging in many languages, not the Console API, and the `siggy` CLI on npm is version 0.0.4 from September 2024 (10 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 3. Security & auth, 74 out of 100, up to 4.6 more on the total Why it scored 74: OAuth for the MCP server with PKCE, dynamic client registration and a device code grant, issuing personal Console API keys that carry the user's role. Console keys can be created read-only or read and write, and rotated or deactivated through the API. The OAuth metadata lists no scopes, and a project Console key reaches every Console API endpoint (25 of 30). MCP scope is set per project and per role as no access, read only or read and write. Updates and destructive operations need a confirmation, and project review policies still apply (19 of 20). No guidance on prompt injection was found, though tools return names, descriptions and log data that people wrote. The docs tell users to review each change before the agent confirms (4 of 15). Audit logs are kept indefinitely with an API, and personal keys attribute changes to a user (14 of 15). SOC 2 Type II and a bug bounty programme are stated on the security page, with no public programme link and no security.txt (12 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 4. Reliability, 79 out of 100, up to 4.2 more on the total Why it scored 79: Graded on the hosted service with the hosted lines. Statuspage at status.statsig.com with 14 components, among them Console API, Config Delivery API and Log Event API (20). Three incidents between 10 July and 8 October 2026, all delays to Warehouse Native exports and results, marked minor or none (20 of 30). A major incident on 8 July 2026, when launched gates returned false through the V2 Config Delivery API for about 4 hours 44 minutes, falls two days outside the window. The docs give about 1,800 mutation attempts per 15 minutes per project, while the OpenAPI description gives about 100 per 10 seconds and 900 per 15 minutes, and no number was found for reads (12 of 15). The docs ask for exponential backoff with jitter on 429. No `Retry-After` header and no idempotency keys were found (10 of 15). The enterprise terms give 99.95 per cent monthly availability with credits, for customers with premium support, measured on the Console user interface (7 of 10). The Console API is versioned and the MCP V3 docs carry no beta label (10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 5. Schema & documentation, 82 out of 100, up to 2.9 more on the total Why it scored 82: Public OpenAPI 3.0 spec for Console API version 20240601 with 209 paths and 325 operations, and typed MCP tools per the tool reference (25). llms.txt, llms-full.txt, a Markdown copy of each page and a public docs MCP server (10). Every operation has a summary but only 29 of 325 have a description. The MCP tool reference states each tool's purpose and when to ask the user first (11 of 20). 1,164 enums across 188 schemas, with required fields marked (12 of 15). Examples on most responses. The spec documents 401 on 156 operations, 404 on 126 and 400 on 106, and no 429 (11 of 15). A dated API version sent in the `STATSIG-API-VERSION` header with a promise not to break a version, and a dated product updates page. No changelog for the API alone was found (13 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 6. Transparency & trust, 75 out of 100, up to 2.2 more on the total Made of editorial 64, provenance 85. Why it scored 75: Closed hosted service under a published self-service agreement and enterprise terms. The SDKs are open source under ISC (17 of 30). The terms, DPA and AI governance page mostly agree. The DPA names Amplitude, Inc. as processor and keeps data as long as necessary, the pricing page gives 1 year of analytics retention on the free plan, and the docs say customer data is not used to build AI models without written consent. The privacy notice is Amplitude's and doesn't name Statsig, and the self-service terms name Amplitude, Inc. under an effective date of 4 August 2025, before Statsig joined Amplitude on 5 May 2026 (18 of 30). Each Console API version is promised not to break, a deprecation notices page gives dates, and MCP V1 stays supported. No notice period is stated (12 of 20). Amplitude's sub-processor list, updated 13 May 2026, gives locations and marks the providers used for Statsig-branded services. EU hosting is an Enterprise option (17 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Terms of service: read, states 6 of the 7 things a reader expects, and has 2 clauses that cost points (5.1 of 10) - security.txt: not found (0 of 10) ## 7. Maintenance & community, 81 out of 100, up to 1.7 more on the total Why it scored 81: The newest product update is dated 12 September 2026 and the server SDKs released 0.23.1 on 28 September 2026 (30). Ten or more dated product updates since 10 July 2026 (20). Closed service with a public updates page and a Slack community. We did not read how quickly questions are answered (10 of 15). `com.statsig/statsig-mcp-server` is in the official MCP registry under the vendor's domain, though the entry dates from 19 September 2025 and lists the V1 URL (13 of 15). The server SDK repository has 25 CI workflow files and 124 commits since 10 July 2026, while the `siggy` CLI has not been published since September 2024 (8 of 10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - The lead named Statsig, Inc. as vendor. The self-service terms, the enterprise terms and the DPA name Amplitude, Inc., and a Statsig blog post says Statsig joined Amplitude on 5 May 2026 with the original team now at OpenAI - Amplitude has its own listing. Statsig is listed apart because it has its own domain, API, MCP server, terms, pricing and status page - provenance.privacy points at amplitude.com/privacy, where statsig.com/privacy redirects. It is the contracting entity's notice and the only one published, but it doesn't name Statsig - The self-service terms name Amplitude, Inc. under an effective date of 4 August 2025, which is before the May 2026 change of owner. We could not tell when the text changed - unchecked: the live `tools/list` response of the MCP server and its tool annotations, which need a key. Tool counts and inputs come from the docs tool reference - unchecked: whether the Console API and MCP server are open to the free Developer plan. The docs state no plan limit, and the pricing page lists API controls under Pro without saying what they are - unchecked: GitHub stars and issue response times. The GitHub API answered with a rate limit - unchecked: the bug bounty programme's scope and where to report. The security page names a programme without a link, and no security.txt exists - The docs and the OpenAPI description give different rate limits for the Console API (about 1,800 mutations per 15 minutes against about 900 requests per 15 minutes) - The 8 July 2026 incident, in which launched gates returned false through the V2 Config Delivery API for about 4 hours 44 minutes, is outside the 90-day window and is not one of the deduction types, so no deduction was made - Capabilities put experiments and flags first because the Console API and MCP tools centre on them. Analytics reads are limited to metric values, dashboard widget results and the Logs Explorer query ## Weaknesses - Console API errors carry only `status` and `message`, and the spec documents no 429 response - No idempotency keys were found, and most MCP update tools replace the whole resource - The project Console API key reads and writes everything in a project unless created read-only - The privacy notice is Amplitude's, last updated 31 August 2026, and doesn't name Statsig - The 99.95 per cent SLA needs Enterprise premium support and covers the Console user interface ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Connect to https://api.statsig.com/v3/mcp. The official MCP registry entry still lists the V1 URL, which has 93 tools - Call `get_context` first to learn the project, your permissions and the review settings - Read the resource before `gate_update`, `experiment_update` or `dynamic_config_update`. They replace the whole resource, so resend every field you want kept - Use `discover_tools` for partial edits, and fetch a fresh `operationHandle` when one expires - On a 429 slow down and back off with jitter. Rejected attempts count towards the quota of about 1,800 mutations per 15 minutes ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.