# The Agent Tool Benchmark (slim) > How Anchor Terminal scores model APIs, agent frameworks, MCP servers, data providers, scraping tools and payment protocols. Methodology v0.3 and the October 2026 research run, the checklist for every category with its points, the two categories still pending, renormalised weights, the computed provenance and data-quality scores, deductions for negative events and AA to F grade bands. - Full: https://www.anchorterminal.com/benchmark/index.md (~30,450 tokens) · this version ~2,030 tokens · JSON https://www.anchorterminal.com/benchmark/index.json · canonical https://www.anchorterminal.com/benchmark/ - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-05 Score 0 to 100 = Σ(category score × weight) ÷ 80 over the seven assessed categories, then negative events (0 to −15). Performance and Task success are pending in this run (no score, not in the total). Grades: AA 85+, A 78+, BB 70+, B 62+, C 54+, D 46+, E 38+, F below. Agent-ready = BB or better. Methodology v0.3, October 2026 research run (2026-10-01): scores from public evidence against the checklists below, with the reason, sources, confidence and open questions on every listing (`anchor.assessment`). Kinds: tools and APIs, model APIs, frameworks (ranked together), payment protocols (graded, not ranked). Transparency = (editorial + provenance) ÷ 2. Our own products (letme, LocalGhost, `own: true`) are graded twice independently, read strictly, with the lower award per checklist item kept by an auditor; no panel reviews and never letme's pick. | Category | Weight | This run | Measures | | --- | --- | --- | --- | | Reliability | 16% | 20% | Does the tool answer, and does it keep answering the same way? In this run it's assessed from public evidence (90 days of status history, documented rate limits and overload handling, SLAs, general availability). | | Performance | 10% | pending | How long an agent waits. | | Schema & documentation | 13% | 16.2% | Can a model understand what the tool does from the tool definition alone? We grade the text a model sees. | | Agent ergonomics | 13% | 16.2% | How much of an agent's context and how many round trips the tool consumes to get a job done. | | Security & auth | 14% | 17.5% | Can an operator give an agent this tool without giving it the keys to everything? | | Payments & pricing | 10% | 12.5% | Can an agent start using the tool, and pay for it, without a human in the loop? A machine payment protocol on the tool's own endpoints is the largest single signal. | | Task success | 10% | pending | Does an agent finish the job? A fixed suite of representative tasks per category, run monthly with a fixed reference model through the tool. | | Maintenance & community | 7% | 8.8% | Is anyone home? Release cadence and responsiveness predict how a tool behaves after the protocol moves under it. | | Transparency & trust | 7% | 8.8% | Can an operator find out who runs the tool, what it does with data, and what changed? Half of this category is the provenance score, computed from checked facts. | ## How this run was made Research agents on Anthropic's Claude models, 30 September to 2 October 2026, public evidence only (status history, docs, OpenAPI and tool definitions, pricing, terms, privacy, security pages, changelogs, repositories, registries), about twelve fetches per listing, no web search. Couldn't check is not absent: unchecked items score as absent, go in open questions and lower confidence. Anthropic's listings graded by the same checklist with a disclosure; the panel doesn't review them. ## Checklists (points) - Reliability, hosted: status page with history 20; last 90 days of incidents 0 to 30; rate limits documented 15; 429 handling and safe retries 15; SLA 10; GA 10. Local and SDKs: official package 20; CI and tests 25; open regressions 0 to 25; semver and changelog 15; 1.0 or stable 15. - Schema: machine-readable contract 25; llms.txt or Markdown docs 10; descriptions say when to use and not 0 to 20; typed inputs 0 to 15; examples and errors 0 to 15; versioning and changelog 15. - Ergonomics: context cost 0 to 25; pagination and size controls 20; recoverable errors 20; idempotency or MCP annotations 20; defaults and SDKs 15. - Security: credential model 0 to 30 (10 off for a secret in a URL); read-only and approvals 0 to 20; prompt-injection posture 0 to 15; audit logs 0 to 15; security programme 0 to 20. - Payments: machine payment protocol 40; per-call prices without login 20; free tier without a card 20; autonomous onboarding 20. - Maintenance: last release 0 to 30; three releases in 90 days 20; responsiveness 0 to 25; MCP registry or official SDKs 15; package health 10. - Transparency (editorial half): licence 0 to 30; data handling 0 to 30; deprecation policy 0 to 20; telemetry or subprocessors 0 to 20. ## Who's behind it {#provenance} Credibility shouldn't be a feeling. For every listing we record the legal entity named in its terms, its registrable domain and the date the registry says it was first registered, whether the hosted endpoint sits on a domain the vendor controls, and whether there are terms, a privacy policy, a status page, a changelog and a valid security.txt. Each check is worth fixed points and the total is the provenance score out of 100. It's half of Transparency and trust. | Check | Points | How it's scored | | --- | --- | --- | | Legal entity named | 20 | The terms, imprint or licence name the company or body that stands behind the service. | | Domain age | 15 | Years since the registrable domain was first registered, from the registry's own RDAP server. 10 years or more scores 15, 5 scores 11, 2 scores 7, 1 scores 3. Government domains score in full. Registration can predate the current owner, which is why this never stands alone. | | Endpoint on the vendor's domain | 15 | The hosted endpoint sits on a domain the vendor controls, not a shared host or a lookalike. Not scored for libraries, specifications and local servers. | | Terms of service | 10 | Published and reachable without a login. For software you run yourself with nothing hosted, an open-source licence counts. | | Privacy policy | 10 | Published and reachable without a login. Not scored for software you run yourself with nothing hosted. | | Status page | 10 | A public status page with incident history. | | Changelog | 10 | Dated release notes or a changelog an agent can read. | | security.txt | 10 | A valid RFC 9116 security.txt on the vendor's domain. An expired one scores half. | ## Data quality {#data-quality} For data providers, half of Task success will be a data-quality score. Task success is pending, so the score is published on the listing and isn't in the total yet. | Check | Points | How it's scored | | --- | --- | --- | | Coverage | 25 | Breadth and depth of what the data covers against the obvious alternatives, rated 1 to 5 by the panel and published with the evidence. | | Freshness | 15 | Real time 15, minutes 13, hourly 10, daily 7, static 3. | | History | 15 | 20 years or more 15, 10 years 12, 5 years 9, 1 year 5, less 2. | | Methodology published | 15 | How the data is collected, cleaned or calculated, in public. | | Sources disclosed | 10 | Where the data comes from, named. | | Licence | 10 | Open licence 10, clear proprietary terms 7, unclear 2. | | Official standing | 10 | Statutory register 10, regulated administrator 9, none 0. | ## When probes run Performance from p50 and p95 per call from three regions; measured availability and errors join Reliability; measured `tools/list` tokens join ergonomics; task suites score Task success (data quality half of it for data providers); weights return to the published ones; every change goes in each listing's history. ## Principles Score what a model sees. Every score has a weight, a reason, sources and a dated history. Rankings can't be bought. We don't host what we rank. Our own products are graded strictly and never picked. Agent-native payments weighted up on purpose. Provenance checks published. No reviewer reviews its own model's vendor. Methodology versioned. ## Data `/api/v1/rankings.json` (ranking with category scores) · `/api/v1/benchmark.json` (weights, effective weights, pending, bands) · `/api/v1/tools/{slug}.json` (includes assessment, provenance and data-quality checks) · CC BY 4.0, cite "Anchor Terminal Agent Tool Benchmark, methodology v0.3, October 2026 research run".