# Fix list: Shippo From Anchor Terminal's listing at https://www.anchorterminal.com/tools/shippo, the October 2026 research run, assessed 7 October 2026. Grade C, 61.9 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on Shippo: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Security & auth, 52 out of 100, up to 8.4 more on the total Why it scored 52: Self-serve API keys, live and test kept apart, shown once, deletable in the portal, with no scopes and no documented expiry. The OAuth flow for platforms has one scope, "*". JWTs for client-side use last 12 hours (20 of 30). A webhook can carry a self-made token in its URL query string. That is the receiver's secret, not the API credential, so we noted it without the deduction. No read-only key. Test keys can't spend money, and the MCP server separates reads from writes and its docs say writes ask for confirmation (9 of 20). Tracking events, metadata and address text come from carriers and other parties. The AI repository's security policy puts prompt injection in scope for reports, with no guidance in the API docs (5 of 15). The API portal is described as showing usage. No audit log was found in the docs (3 of 15). The trust centre lists SOC 2 Type II (November 2024 to November 2025), PCI DSS SAQ-A, an annual third-party penetration test and a bug bounty. security.txt returns 404, as does goshippo.com/security, which the repository names as the disclosure page (15 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 2. Reliability, 59 out of 100, up to 8.2 more on the total Why it scored 59: Read with the hosted lines and scored on the REST API at api.goshippo.com. Statuspage at status.goshippo.com with components for the REST API, the dashboard and each carrier (20). From 9 July to 7 October 2026 the history holds about 115 incidents outside maintenance, almost all on one carrier's rating, label or tracking service. None names the Shippo REST API. A third-party tracking vendor had elevated errors for 9 hours on 25 August, marked major, and USPS or Canada Post label creation failed for hours on several days, so we read this as one major rather than minor only (10 of 30). Per-minute limits per endpoint, live and test (15). The docs say a 429 is returned and nothing more. No Retry-After or backoff guidance was found, and label purchase takes no idempotency key, though the TypeScript SDK has a configurable backoff retry (4 of 15). No SLA found. API Premier lists paid 24/7 monitoring as an add-on (0). GA (10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 3. Payments & pricing, 40 out of 100, up to 7.5 more on the total Why it scored 40: No x402, MPP or L402 (0). Per-unit prices are public without a login. API Starter is 7 cents a label after 30 free a month, 1 cent a rate request, 2 cents a tracking request, 2 cents a US address validation and 8 cents elsewhere, plus the postage itself (20). Signup is free, and test keys create test labels and mock tracking at no charge. The docs say billing details are needed only for live label purchase. We didn't open an account to confirm (20). A person signs up in a browser and creates the key in the portal, or signs in through OAuth for the MCP server (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 4. Agent ergonomics, 64 out of 100, up to 5.9 more on the total Why it scored 64: Scored on the REST API, with the MCP server noted. List calls take `results` and `page`, and `expand` controls nested objects, but there's no field selection and a shipment returns every rate. The MCP server has four tools (list, describe, read-execute, write-execute) over 73 operations, which keeps definitions small (17 of 25). Pagination on every list. Date filters exist only on shipments, within a 90-day range (14 of 20). Status codes are documented in general terms and the error body is untyped. The goshippo/ai repository has an error reference with recovery steps for MCP use (12 of 20). No idempotency key on shipments, transactions or refunds. The Reporting API accepts `idempotency_key`. The MCP read and write split lets a client approve writes. Its annotations sit behind sign-in and weren't read (8 of 20). Official SDKs for Python, TypeScript, Java, C# and PHP. Shipments and transactions default to asynchronous, and objects can't be edited after creation (13 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 5. Transparency & trust, 61 out of 100, up to 3.4 more on the total Made of editorial 38, provenance 84. Why it scored 61: Closed service. The terms of use sit in a privacy centre that loads only with JavaScript, so we couldn't read them or confirm the contracting entity. SDKs and the goshippo/ai repository are MIT (12 of 30). The privacy notice has the same problem. The trust centre links a Data Processing Addendum and says data is encrypted in transit and at rest, and the release notes say shipments and rates are retrievable for 390 days. No retention periods were read (8 of 30). The versioning page says accounts are never forced onto a new API version, and notices carry dates, such as the 13 January 2026 notice to re-register FedEx accounts before 31 March. Carrier removals are announced on the day (12 of 20). The trust centre names AWS and says third parties are used. No subprocessor list was found. Webhook addresses are published for US and EU regions (6 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Terms of service: published, but our reader couldn't read it (7 of 10) - Privacy policy: published, but our reader couldn't read it (7 of 10) - security.txt: not found (0 of 10) ## 6. Schema & documentation, 83 out of 100, up to 2.8 more on the total Why it scored 83: OpenAPI 3.1 spec for the main API with 71 operations, and six further specs for addresses v2, estimates, reporting, invoices and surcharges (25). llms.txt at docs.goshippo.com and a Markdown copy of each page (10). Descriptions state purpose and some conditions, such as the 390-day limit on shipments and when a batch is asynchronous, but rarely when not to use an endpoint (13 of 20). 94 enums and required fields marked. The 400 response schema is an open object, and weights and dimensions are strings (11 of 15). 427 examples in the spec and samples in six languages. The spec documents 400 on 66 operations, 401 and 404 once each and never 429, and the status-code page is one line per code (9 of 15). Dated API versions set by the Shippo-API-Version header and dated release notes (15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 7. Maintenance & community, 77 out of 100, up to 2 more on the total Why it scored 77: The API release notes were last updated on 1 October 2026 (30). Six dated entries since 9 July, including the Reporting API on 15 September and a tracking fix on 23 September (20). Closed service with dated release notes and a support portal. The SDK repositories show 9 and 4 open issues, and the TypeScript SDK's last release is 2.18.0 on 4 February 2026 (9 of 15). The MCP server is in the official registry as com.shippo/shippo-mcp, verified by DNS. Official SDKs exist in five languages, but PyPI's latest is 3.9.0 from 26 November 2024 and the SDK page says the libraries are being updated (12 of 15). SDKs are generated and validated by CI workflows, and goshippo/ai has test and validation workflows, last tagged 23 July (6 of 10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - unchecked: the terms of use and privacy notice at privacy.goshippo.com load only with JavaScript, so the contracting legal entity, governing law, retention periods and any AI training statement weren't read - unchecked: the hosted MCP server's tool definitions and annotations, which need a Shippo sign-in - unchecked: whether signup asks for a card before a test key is issued. The docs say billing details are needed only for live labels - unchecked: the SOC 2 report, DPA text and bug bounty terms, which the trust centre gives on request - Whether the 390-day limit on retrieving shipments and rates (release note of 6 April 2026) had advance notice - Whether an SLA exists on API Premier. None is published - Whether the API returns Retry-After on a 429. The docs and the spec don't say ## Weaknesses - No idempotency key on label purchase or other writes. Only the Reporting API accepts one - API keys have no scopes or read-only mode, and the OAuth flow has one scope, "*" - The hosted MCP server runs on the live account only, with no test mode - Carrier-side incidents on the status page on most days from July to October 2026, many marked major or critical - Terms of use and privacy notice load only with JavaScript, and no SLA or subprocessor list was found - The Python SDK on PyPI is 3.9.0 from 26 November 2024 ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Build and test with a shippo_test_ key. Test and live object IDs are separate, and the hosted MCP server has no test mode - Send `Authorization: ShippoToken `, not Bearer, and pin `Shippo-API-Version: 2018-02-08` - Before retrying a failed POST /transactions, list recent transactions. There is no idempotency key, and a second purchase is a second charge - Pass `async: false` on shipments and transactions, or poll the object until its status leaves QUEUED - Set `results` on list calls (under 200). Defaults vary by resource, and list calls are limited to 50 a minute on live keys ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.