Methodology v0.3 · October 2026 research run

The Agent Tool Benchmark

Every graded listing gets a score from 0 to 100 and a grade from AA to F. In the October 2026 research run, seven of the nine weighted categories were scored from public evidence against the checklist on this page, with the reason and sources for every score published on the listing. Performance and Task success need our probes and task suites, which haven't run, so they're pending and their weight is shared across the other seven. This page is the specification.

Why another benchmark

The directories that exist rank agent tools by GitHub stars, package downloads, or how often a tool gets called through one particular proxy. Those numbers say what is popular. They don't say whether the tool answers at three in the morning, whether its descriptions make sense to a model, how many tokens it burns before the first useful call, whether an operator can hand it to an agent without handing over the keys, or whether an agent can pay for it without a human. Popularity and quality have already come apart. The most-used documentation server in the MCP world ships a 2,000-character tool description that a community grader marked as failing (we checked, and it is 2,006 characters).

We spent a decade grading crypto exchanges with a public method, and the thing we learned is that once the checklist is public, the people being graded start fixing the items on it. So the plan here is the same. Assess what can be read against a published checklist and show the working, probe what can be probed, and say plainly when something couldn't be verified.

An agent doesn't run on MCP servers alone. It runs on a model, inside a framework, reading data from providers and pages it scrapes, and sometimes paying for what it uses. So the benchmark covers all of those, on one scale, with each category read the way that makes sense for the kind of thing being scored.

What we benchmark

Four kinds of listing share the nine categories. The weights don't change between kinds. What each category measures does, and this table is the whole mapping.

CategoryMCP servers and APIsModel APIs, routers and model platformsAgent frameworksAgent harnessesPayment protocols
ReliabilityAvailability and error rate from probesAvailability, overload errors and how they're signalledTask-suite runs that finish without a framework errorInstalls from an official package and runs headless without a harness errorReference implementations and public facilitators we can reach
Performancep50 and p95 of a representative callTime to first token and throughputOverhead per step on top of the model callTurns and wall time per task on top of the model callsTime from 402 to a settled, paid response
Schema & documentationTool descriptions and input schemasAPI reference, OpenAPI, llms.txt, structured outputsDocs and typed interfaces a model can followDocs, settings and permission rules an operator and a model can followSpec clarity, completeness and consistency with releases
Agent ergonomicsContext cost, pagination, recoverable errorsTool use, caching, context length, SDKsCode and defaults needed for a tool-calling agent with MCPWhat it takes to run headless in CI with one MCP server attachedWork for a seller to accept and a buyer to pay
Security & authScopes, read-only modes, confirmation on writesRetention, training on data, zero-retention optionsTelemetry defaults, approvals, guardrailsApproval modes, sandbox and network defaults, what leaves the machinePublished attacks and the controls that answer them
Payments & pricingMachine payment, published prices, card-free startPublished token prices, free tier, onboardingLicence and hosted-runtime pricingPrice of the harness and of the account it needs, and a free way to try itFees, rails and how much an agent can do without a person
Task successTask suite pass rate (data quality is half for data providers, once scored)Agent task suite with the model drivingThe same task suite run through the frameworkThe repository task suite run through the harnessWhether an agent can complete a real purchase today
Maintenance & communityRelease cadence and responsivenessRetirement notice periods and model churnRelease cadence and breaking changesRelease cadence, breaking changes and advisories handledSpec activity and governance
Transparency & trustEditorial plus the computed provenance scoreEditorial plus provenanceEditorial plus provenanceEditorial plus provenance, including telemetry disclosureEditorial plus provenance, including who governs the spec

MCP servers, HTTP APIs, model APIs, frameworks, data providers and scraping tools are ranked together, because an agent builder does choose between them for the same budget. Payment protocols are graded on the same scale but not ranked against tools. You choose a protocol by who you're paying, not by its score, and putting a specification at number two above a search API told nobody anything useful. Retired listings keep their page and their grade and drop out of the ranking.

Categories and weights

Nine categories add up to 100 points. The weights follow what breaks agents in practice. A tool that is down is useless whatever its schema looks like, so reliability carries the most. A tool an agent can't pay for is a tool it can't use alone, so payments carry more than a payments category normally would. We're not certain the weights are right (they're our first guess, and the first run with probes will tell us where they're wrong).

Two of the nine are pending in this run. The "This run" column is the share of the 100 points each category carries until they're scored.

CategoryWeightThis runWhat it measures
Reliability16%20%Does the tool answer, and does it keep answering the same way? In this run it's assessed from public evidence (90 days of status history, documented rate limits and overload handling, SLAs, general availability). Our own probes from three regions will replace that evidence as they accumulate.
  • 30-day availability of the endpoint (remote) or of a clean start plus tools/list (local)
  • Error rate on a fixed set of representative calls
  • Timeout rate and behaviour under back-pressure (429 with Retry-After, not silent hangs)
  • Consistency, meaning the same input gives the same shape of output across the window
  • Graceful degradation when upstream systems are down
Performancepending10%pendingHow long an agent waits. Latency is measured per call, not per page load, because a slow tool costs an agent a whole turn.
  • p50 and p95 latency of representative calls
  • Cold-start time for local servers and first-call time for hosted ones
  • Streaming or partial results where the task is long-running
  • Payload size discipline (results sized for a context window rather than a data warehouse)

Pending. Latency is measured per call by our probes, which haven't run yet, so this run doesn't score it. Its weight is shared across the assessed categories until the first probe window closes.

Schema & documentation13%16.2%Can a model understand what the tool does from the tool definition alone? We grade the text a model sees.
  • Tool descriptions state purpose, when to use, and when not to
  • Input schemas with types, enums, constraints and required fields; no free-form JSON blobs
  • Examples in descriptions or docs; documented error responses
  • Versioning of the tool surface and a changelog
  • Machine-readable docs (llms.txt, OpenAPI, registry server.json)
Agent ergonomics13%16.2%How much of an agent's context and how many round trips the tool consumes to get a job done.
  • Context cost, the tokens tools/list consumes before any work starts
  • Pagination, filtering and output-size controls
  • Actionable error messages an agent can recover from without a human
  • Idempotency and safe retries; read-only variants and tool annotations (readOnlyHint, destructiveHint)
  • Sensible defaults and few mandatory parameters; toolsets or meta-tools when the surface is large
Security & auth14%17.5%Can an operator give an agent this tool without giving it the keys to everything?
  • OAuth 2.1 with scoped, revocable, short-lived credentials; no secrets in URLs
  • Read-only and scoped modes; confirmation for destructive actions
  • Prompt-injection mitigations for tools that return untrusted content
  • Audit logging and per-call visibility for the operator
  • Handling of advisories (disclosed, patched, communicated)
Payments & pricing10%12.5%Can an agent start using the tool, and pay for it, without a human in the loop? A machine payment protocol on the tool's own endpoints is the largest single signal.
  • A machine payment protocol (x402, MPP or L402) on the tool's own endpoints (40 points)
  • Transparent, per-call pricing with units, published without a login (20 points)
  • A free tier or trial that doesn't require a card (20 points)
  • Autonomous onboarding, meaning an agent can obtain access without a human signup flow (20 points)
Task successpending10%pendingDoes an agent finish the job? A fixed suite of representative tasks per category, run monthly with a fixed reference model through the tool. For data providers, half of this category is the data-quality score, because a clean API over thin data still fails the task.
  • For data providers, the data-quality score from /benchmark/#data-quality (half the category)
  • Pass rate on the category task suite (pass@1 and pass@4)
  • Retries, turns and tokens per completed task
  • Recovery, whether an error on the first attempt is fixable from the tool's own error message
  • Consistency across reference models

Pending. Task success needs the category task suites run through each tool, which haven't run yet, so this run doesn't score it. Its weight is shared across the assessed categories until then. A data provider's data-quality score is published on its listing now and becomes half of this category when it's scored.

Maintenance & community7%8.8%Is anyone home? Release cadence and responsiveness predict how a tool behaves after the protocol moves under it.
  • Release cadence in the last 90 days and time since the last release
  • Issue and pull-request responsiveness
  • Presence in the official MCP registry under a verified namespace
  • Package health (current SDK versions, no pinned-and-forgotten dependencies)
Transparency & trust7%8.8%Can an operator find out who runs the tool, what it does with data, and what changed? Half of this category is the provenance score, computed from checked facts. Legal entity, domain age, whether the endpoint sits on the vendor's own domain, terms, privacy policy, status page, changelog and security.txt.
  • Provenance score, computed from the checks on /benchmark/#provenance (half the category)
  • Source availability and licence clarity
  • Data handling and retention statements that agree with each other
  • Deprecation notices with dates
  • Telemetry disclosed and opt-out documented
Negative events−15 max−15 maxDeductions of up to 15 points for incidents in the last 12 months. Security incidents, breaking changes shipped without notice, silent deprecations, unresolved advisories, or misleading listings. Deductions decay over time and are lifted early when the vendor documents the fix.
  • Confirmed data exfiltration path via the tool (up to -15)
  • Breaking tool rename or schema change without a deprecation window (-3 to -8)
  • Endpoint removed while still advertised in a registry or docs (-3 to -6)
  • Telemetry enabled without disclosure (-2 to -5)

This run is each category's share of the 100 points while Performance and Task success are pending, its weight ÷ 80 × 100. A total is Σ(score × weight) ÷ 80 over the assessed categories, then the negative events.

How this run was made

Until 1 October 2026 every score on this site was an invented sample, published to show the method. The October 2026 research run replaced them.

Between 30 September and 2 October 2026, research agents running on Anthropic's Claude models researched every graded listing from public evidence. They read status pages and their incident history, rate-limit and error documentation, API references and OpenAPI files, MCP tool definitions in the source, pricing pages, terms and privacy policies, security and trust pages, changelogs, release tags and CI in cloned repositories, package registries and public advisories. They worked from the links on each listing and the pages those linked to, with a budget of about twelve page fetches per listing and no web search. Each category was scored against the checklist below. The points add up to the score, and the note beside each score on the listing says what earned them and where the agent departed from the checklist and why.

Five rules held for every listing.

  • Score from evidence found. Nothing gets points for what a vendor probably does.
  • Couldn't check isn't the same as absent. If a thing was looked for and isn't there (no status page, no SLA, pricing behind a login), it scores as absent. If a page couldn't be loaded, the agent could rely on the listing's own facts from their 30 September check and say so in the note. If those didn't cover it either, the item scores as absent, goes on the listing as an open question ("unchecked" and what), and the confidence drops. Our own fetch limits never read as a vendor's failing without saying so.
  • Vendor claims are claims. "Exa says its index covers 1.4 trillion URLs" is a fact about Exa's page, not our measurement.
  • No invented numbers, dates, incidents or quotes. If it wasn't found, it isn't true for this run.
  • Facts on a listing that turned out wrong (a renamed product, a moved endpoint, dropped x402 support) were corrected, and the correction is named in the listing's open questions.

Every graded listing carries three things beside its scores. A confidence, high, medium or low, which is how sure we are that the scores would survive the vendor reading the notes. The open questions, under "What we couldn't check". And the sources, at least four per listing, each with the date it was read. The JSON for a listing has all of it under anchor.assessment.

The conflict rule. The research agents and the review panel run on Claude, so Anthropic is a company we depend on. Anthropic's listings (the Anthropic API and the Claude Agent SDK) were graded by the same checklist as every other listing, with each judgement call written in the note, and each carries a disclosure. The panel doesn't review them. Our own products are graded by a stricter version of the same rule (below). A listing the founder built somewhere else, the CoinDesk Data API, is graded like the rest and carries a disclosure too.

The checklist

Each category has a checklist. The points are what an item is worth, and a listing's score in the category is the sum. Read each one the way that fits the kind of listing (the table above). These are the checklists the research agents used, written for a reader.

Reliability, from public evidence

Hosted APIs, MCP servers, models and platforms.

  • 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).
  • 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.
  • 15, rate limits documented with numbers.
  • 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.
  • 10, an SLA published for any paid tier.
  • 10, the surface agents use is generally available, not beta or preview.

Local packages, SDKs, frameworks and stdio MCP servers.

  • 20, installs from an official package with supported runtimes stated.
  • 25, a public CI and test suite, passing on the default branch.
  • 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).
  • 15, semver discipline and breaking changes called out in a changelog.
  • 15, version 1.0 or later, or declared stable.

Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.

Schema and documentation

APIs and MCP servers.

  • 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).
  • 10, llms.txt or Markdown docs served for agents.
  • 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.
  • 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.
  • 0 to 15, examples and documented error responses.
  • 15, versioning and a public changelog.

Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.

Agent ergonomics

  • 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).
  • 20, pagination, filtering and output-size controls.
  • 20, actionable, documented error responses, codes and messages an agent can recover from.
  • 20, idempotency or safe retries, and for MCP the readOnlyHint and destructiveHint annotations.
  • 15, sensible defaults, few required parameters, and official SDKs in at least two languages.

Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.

Security and auth

  • 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.
  • 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.
  • 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.
  • 0 to 15, audit logs or per-call visibility for the operator.
  • 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.

Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.

Payments and pricing

The published rubric, also on the x402 page.

  • 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.
  • 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login.
  • 20, a free tier or trial that doesn't need a card.
  • 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).

Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.

Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.

Maintenance and community

  • 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.
  • 20, at least three releases or dated changelog entries in the last 90 days.
  • 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.
  • 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).
  • 10, package health, current dependencies and CI.

Models are read for deprecation notice periods and model churn rather than release counts.

Transparency and trust, the editorial half

  • 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.
  • 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).
  • 0 to 20, a deprecation policy or notices with dates.
  • 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).

The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.

Pending categories

Performance (weight 10) is latency per call, measured by our probes. Task success (weight 10) is a fixed task suite run through every tool in a category. Neither can be read from a vendor's pages, and neither has run, so this run doesn't score them. They show as pending everywhere a score is shown, with no number, and they add nothing to the total.

The total is the weighted mean over the seven assessed categories, renormalised so the maximum is still 100.

score = Σ (category score × weight) ÷ 80, over the seven assessed categories
        then the negative events, 0 to −15

80 is the sum of the assessed weights. So Reliability counts for 20 points in this run rather than 16, Security for 17.5 rather than 14, and so on (the "This run" column above). Grade bands are unchanged. One consequence is worth saying out loud. A tool that would do badly on latency or on real tasks isn't marked down for it yet, and a fast one isn't rewarded. Treat this run's grades as a reading of what a vendor publishes and how it runs its service, which is a lot, but not all, of what an agent meets.

What changes when the probes run

  • Performance gets scored from p50 and p95 latency per call, measured from London, Virginia and Singapore.
  • Reliability adds measured availability and error rates to the status history it's read from today.
  • Context cost in Agent ergonomics becomes the measured token count of tools/list, in the default configuration and with every toolset on.
  • Task success gets scored from the task suites, and for data providers half of it becomes the data-quality score already published on their listings.
  • The weights go back to the published ones, and every listing's history records the change with the methodology version that made it.

Grades will move when that happens, some of them a lot. The probes and pollers that watch uptime today feed the live panel on each listing and don't change a score until a run.

Our own products

letme (letme.dev) is Anchor Terminal's own product, and LocalGhost is built by its founder. Until 2 October 2026 we listed them and didn't grade them. On 2 October both were graded by the same checklist as every listing, with four differences, because a grade we give ourselves has to hold up for a reader who assumes we were kind to ourselves. letme's grade still comes from that rule.

  • Two research agents graded each one independently, and every judgement call was read strictly. Only what was live and public that day counted. Anything specified, planned or "coming" earned nothing, and so did our own description of how something works without the code, the live answers or a dated document behind it.
  • A third agent, which graded neither, reconciled the two item by item on the checklist. The lower award stood unless it rested on an error the auditor checked for itself, and a deduction counted only if the auditor could still see the problem that day.
  • The review panel doesn't review them.
  • letme never picks them and never answers with them, so it can't favour the company that runs it. Their records say own: true. They're in rankings, categories and comparisons like any listing.

Since 3 October, at the founder's request, LocalGhost is graded the same way as every listing instead, with the same readings, neither stricter nor looser, the way a competitor of ours is. Two research agents grade it independently and a third reconciles them, checking the evidence itself wherever they disagree, with no award standing by default. The panel still doesn't review it and letme still never picks it, because those protect against us favouring ourselves rather than marking us down. letme moves to the same rule at its next check.

Each carries a disclosure that says how it was graded, the reason and sources for every score are on the listing like any other, and anyone can dispute an item the same way. Every grading and reconciliation is kept with the research run.

Listed, not graded

A listing can be listed and not graded (graded: false), with no score, grade, rank, reviews or letme pick, and a disclosure that says why. None is today.

Competitors of our own products

When a listing competes directly with one of our own products, a reader could suspect the founder's side marked it down, or that we went easy on it to look fair. So it's graded by the same checklist as every listing, neither stricter nor looser, and the grading is set up so that neither can creep in.

  • Two research agents grade it independently. Where a checklist line could be read two ways, they take the reading already used for the same line on comparable listings. They don't look at our own product's listing and don't compare the two.
  • A third agent, which graded neither, reconciles them item by item. Where the graders agree, their award stands unless the auditor finds it rests on an error. Where they disagree, neither award stands by default. The auditor checks the evidence itself and awards what it supports.
  • What we couldn't read counts as absent, as for every listing, and the note says when a score is low because we couldn't read something rather than because it's missing. When a vendor's site refuses our reader, text supplied to us from it counts, and the listing says who supplied it.
  • The review panel reviews it like any listing, and letme can pick it.
  • A disclosure at the foot of its page, and of any comparison it's in, says which of our products it competes with and how it was graded. Its record says competesWith and names that product.

Four listings are graded this way today, all competing with LocalGhost. Underdog was the first, on 3 October 2026, and Khoj, Open WebUI and Screenpipe followed the same day, the three that LocalGhost's own about page names as its competitors. A listing is graded this way when our product names it as a competitor, or when it does the same job for the same people, as Underdog does. Both gradings and the reconciliation are kept with the research run for each.

Negative events

Scores describe the steady state. Negative events describe what happened. Up to 15 points come off for incidents in the last twelve months that can be sourced. A security incident or an exfiltration path, up to 15. A breaking change shipped without notice, 3 to 8. An endpoint removed while still advertised, 3 to 6. Telemetry nobody disclosed, 2 to 5. A misleading listing or claim, 2 to 5. Deductions decay, and a vendor that documents the fix and ships mitigations gets the deduction lifted early. A vendor that ships a breaking rename in a minor release and says nothing carries the full term. Every deduction is written on the tool's page with a date and a source.

Grades

GradeScoreMeaningTools in this run
AA85+Exceptional. Agent-ready, with agent-native payments or equivalent autonomy.1
A78+Excellent. Agent-ready; minor gaps.17
BB70+Good. Agent-ready with documented caveats.87
B62+Usable. Needs operator supervision or workarounds.120
C54+Mixed. Material gaps in one or more categories.108
D46+Poor. Not recommended for autonomous use.67
E38+Very poor.36
F0+Failing or unverifiable.26

Agent-ready means BB or better, which is our shorthand for "you can give this to an autonomous agent as long as you read the caveats on its page". Below BB expect to supervise it, patch around it, or replace it. F is for tools that fail the checklist across the board or can't be verified at all.

Who's behind it

Credibility shouldn't be a feeling. For every listing we record the legal entity named in its terms, its registrable domain and the date the registry says it was first registered, whether the hosted endpoint sits on a domain the vendor controls, and whether there are terms, a privacy policy, a status page, a changelog and a valid security.txt. Each check is worth fixed points and the total is the provenance score out of 100. It's half of Transparency and trust. The other half stays editorial, because licence clarity, retention statements that agree with each other and honest telemetry disclosure need reading, not counting.

CheckPointsHow it's scored
Legal entity named20The terms, imprint or licence name the company or body that stands behind the service.
Domain age15Years since the registrable domain was first registered, from the registry's own RDAP server. 10 years or more scores 15, 5 scores 11, 2 scores 7, 1 scores 3. Government domains score in full. Registration can predate the current owner, which is why this never stands alone.
Endpoint on the vendor's domain15The hosted endpoint sits on a domain the vendor controls, not a shared host or a lookalike. Not scored for libraries, specifications and local servers.
Terms of service10Published and reachable without a login. For software you run yourself with nothing hosted, an open-source licence counts.
Privacy policy10Published and reachable without a login. Not scored for software you run yourself with nothing hosted.
Status page10A public status page with incident history.
Changelog10Dated release notes or a changelog an agent can read.
security.txt10A valid RFC 9116 security.txt on the vendor's domain. An expired one scores half.

Domain age is the weakest of these and we know it. A registration date is when the name was first taken, not when the current owner bought it, so slack.com (1992) and notion.com (1997) look older than the companies are. Every listing where that's true says so under the checks. That's why age is 15 points and never decides a grade on its own. A brand-new domain with a named company, published terms and a working status page still scores well, and a very old domain with nobody's name on it doesn't.

Registry dates come from each registry's own RDAP server. Where the registry doesn't publish one (.gov.uk, some brand top-level domains) we say so on the listing and score government domains in full.

Data quality

A clean API over thin data still fails the task. For data providers, half of Task success will be a data-quality score computed from what the data covers, how fresh and how deep it is, whether the method and sources are published, what the licence lets you do, and whether the publisher has official standing. Coverage is the one judgement call in it, rated 1 to 5 with the evidence written next to the number. Task success is pending in this run, so the data-quality score is computed and published on each data provider's listing, and it isn't in the total until the task suite runs. Each listing also says what its terms allow for caching and redistribution, because an agent that stores a price it wasn't allowed to store is a problem for its operator.

CheckPointsHow it's scored
Coverage25Breadth and depth of what the data covers against the obvious alternatives, rated 1 to 5 by the panel and published with the evidence.
Freshness15Real time 15, minutes 13, hourly 10, daily 7, static 3.
History1520 years or more 15, 10 years 12, 5 years 9, 1 year 5, less 2.
Methodology published15How the data is collected, cleaned or calculated, in public.
Sources disclosed10Where the data comes from, named.
Licence10Open licence 10, clear proprietary terms 7, unclear 2.
Official standing10Statutory register 10, regulated administrator 9, none 0.

How each category is measured

In this run, from public evidence

Reliability, schema and documentation, agent ergonomics, security, payments, maintenance and the editorial half of transparency were scored against the checklists above by research agents reading what a model reads and what an operator would check. The tool definitions, the docs, the auth flow, the status history, the pricing, the terms and the repository. Each score's note is on the listing, with the sources. Vendors can dispute any item with evidence, and disputes are answered in public.

Machine payment in this run was read from each vendor's documentation, its source and its listing in the x402 Bazaar where it had one.

When the probes and task suites run

Remote tools will be probed every five minutes from London, Virginia and Singapore. A probe opens a session, lists tools, and runs one small fixed call that doesn't change state. Local tools will be started in a clean container every fifteen minutes, listed, and called once. We'll record availability, latency percentiles, error classes, timeouts, and whether the response shape matched the previous run. The token count of tools/list gets measured with a reference tokeniser in the default configuration and again with every optional toolset switched on, and both numbers are published, because the gap between them is the first thing to configure.

Machine payment will be detected, not read. For x402 we call the endpoint without paying and look for a 402 carrying a PAYMENT-REQUIRED header (v2) or a JSON body with x402Version: 1 (legacy) that names a scheme, a CAIP-2 network, an amount, an asset and a payee. MPP (WWW-Authenticate: Payment) and L402 (WWW-Authenticate: L402) challenges get detected the same way. For tools that settle, one real micropayment a week from a canary wallet, with the transaction hash published.

Each category will have a suite of representative tasks (find a fact on the web, open a pull request, run a read-only query, pull a table out of a page). Once a month the suite runs through every tool in the category with a fixed reference model and a fixed harness, and we record pass@1, pass@4, turns, tokens, and whether the tool's own error messages were enough to recover from a first failure. The suites will be published so vendors can run them.

Maintenance and community, by public signals

Release cadence, time since the last release, responsiveness on issues and pull requests, presence in the official MCP registry under a verified namespace, and dependency health, all read from public repositories and package registries on the run date. For model APIs this is mostly about retirements. How much notice, how often, and whether a retired id fails loudly. Every dated change we find goes on the listing and on Sunsets, with a calendar feed at /sunsets.ics.

Leaderboard

The full ranking with every assessed category score. Sort and filter in the directory, or fetch it as JSON at /api/v1/rankings.json.

#ToolCategoryGrade Score ReliScheAgenSecuPaymMainTranNegConfidence
1OpenAI Agents SDKFrameworksAA86.58595978060100920high
2OpenAI APIModelsA82.870100981003091850medium
3Stripe API + MCPPlatformsA82.4639094976582870high
4InfisicalSecretsA81.9908791913090850medium
–Machine Payments Protocol (MPP)Pay per callA81.185899179949545-3medium
5Twilio API + MCPMessagingA80.4909280854085810medium
6Twilio Programmable Voice API + MCPCallingA80.4909280854085810medium
7Pydantic AIFrameworksA8083958580609077-2medium
–x402Pay per callA79.787868467979365-3high
8Google Calendar APISchedulingA79.5908391823587790medium
9Amazon S3StorageA79.3959283862083800medium
10Descope Agentic Identity HubAgent authA79.21008280864076700medium
11Apify MCP ServerScrapersA78.665929262959588-3medium
12Google Drive API + MCPStorageA78.6908385863585740medium
13MongoDB MCP ServerDatabasesA78.6857975816092780high
14Cloudflare R2StorageA78.4829290833087750medium
15AWS Secrets ManagerSecretsA78.1879690882065790medium
16Google Cloud Model ArmorGuardrailsA789078751002085880high
17Bird API + MCPMessagingBB77.7739290853583800medium
18Claude APIModelsBB77.6609095923091880medium
19NovuNotificationsBB77.4939284634092700high
20Tavily API + MCPSearchBB77.2908981527585650medium
21TemporalHuman approvalBB77.2859185862090700medium
22Chrome DevTools MCPBrowserBB77.183878363609382-1high
23Azure AI Speech speech-to-textSTTBB77908075952080880medium
24You.com APIsSearchBB76.9858284538882630medium
25BrowserbaseBrowserBB76.6908075439092750medium
26Google Cloud Secret ManagerSecretsBB76.6878382852087850medium
27TempoPlatformsBB76.6718878708088630medium
28Pinecone API + MCPRetrievalBB76.475949379408775-2high
29Amazon PollyTTSBB75.81009082802050800high
30Supabase API + MCPDatabasesBB75.8608988843590920medium
31GroqCloudModelsBB75.71006480774072860medium
32Arize PhoenixEvalsBB75.6778890566084760medium
33Modal SandboxesSandboxesBB75.6957967764093740medium
34Firecrawl MCPScrapersBB75.5708892675087760high
35Backblaze B2StorageBB75.4608985904090740medium
36Composio (API + MCP)Tool accessBB75.3708990704093780medium
37Qdrant API + MCPRetrievalBB75.380878783308584-2high
38Resend API + MCPEmailBB75.3859567654093850medium
39Mapbox APIs + MCPMapsBB75.2828588713080860medium
40Shopify API + MCPCommerceBB75.2739279714085910medium
41Amazon Bedrock GuardrailsGuardrailsBB75.1809293942045700medium
42Amazon SESEmailBB75.1958564872085770medium
43ZenRowsScrapersBB75.1878379418087750medium
44AgentMail API + MCPInboxesBB75618586628585700medium
45Agent Development Kit (ADK)FrameworksBB74.980877092608770-4medium
46Trigger.devHuman approvalBB74.8759179734090740medium
47Speechify API Voice CloningCloningBB74.5909482732075690medium
48Telnyx Voice API + MCPCallingBB74.4759390456583730medium
49SpiderScrapersBB74.3808770429685690medium
50Circle Wallets (Agent Wallets, Programmable Wallets)WalletsBB74.1688875657577750medium
51Parallel Search and Task APIsSearchBB74.1759481457587660medium
52gooseHarnessesBB73.977818277608663-2medium
53ElevenLabs Voice Cloning and Voice Design APICloningBB73.8638785843085840high
54Telnyx API + MCPMessagingBB73.8809380456583730medium
55Akeyless (SecretlessAI and MCP server)SecretsBB73.7908163892582730medium
56Azure AI Speech text-to-speechTTSBB73.7906575902080880medium
57Amazon TranscribeSTTBB73.6959080802035850medium
58OpenAI CodexHarnessesBB73.455908082608783-2medium
59OpenAI embeddingsEmbeddingsBB73.465899095306088-2high
60Apideck Accounting API + MCPAccountingBB73.2759288663080760medium
61ElevenLabs Text to Speech API + MCPTTSBB73.1809482604077710high
62Context7CodeBB73688669696095720medium
63Deepgram Text-to-Speech (Aura-2, Flux TTS)TTSBB73709582704073750high
64WooCommerce API + MCPCommerceBB73807783556082760medium
65SuprSendNotificationsBB72.9699487664081690medium
66Langfuse API + MCPEvalsBB72.880937465408887-2high
67OpenAI Image APIImageBB72.780907285207992-2high
68BlockRun.AIModelsBB72.55575707010085660medium
69Cohere Embed and RerankEmbeddingsBB72.5839287503587690medium
70DeepL APITranslationBB72.4529583823087840medium
71Claude Agent SDKFrameworksBB72.470899483408786-6medium
72Gemini CLIHarnessesBB72.371937867408890-2medium
73Azure TranslatorTranslationBB72.2838382752070800medium
74Scalekit AgentKitAgent authBB72.1758783664080680medium
75Terraform MCP ServerInfraBB72.183807465607889-3high
76CoinGecko x402 APIDataBB71.9707261609682750medium
77Intercom API + MCPSupportBB71.8739474723072830medium
78Coinbase Developer Platform (Agentic Wallet, AgentKit, CDP MCP)WalletsBB71.649948378857780-5medium
79DopplerSecretsBB71.6908159822576770medium
80HubSpot API + MCPCRMBB71.6609272823084860high
81OpenAI Moderation APIGuardrailsBB71.665928592304790-2high
82Auth0 for AI Agents (Token Vault)Agent authBB71.5757371883074850medium
83ElevenLabs Agents API + MCPVoice agentsBB71.5609277763585790medium
84VendureCommerceBB71.489916865458572-3medium
85LangSmith API + MCPEvalsBB71.381867677408578-4medium
86Mistral AI APIModelsBB71.3509391664088820medium
87Nylas Calendar and Scheduler APISchedulingBB71.3759171594088790medium
88Cloudflare MCP ServersInfraBB71.1677877744569900medium
89Nevermined API + MCPPlatformsBB71.1728074636589540medium
90Gemini EmbeddingEmbeddingsBB71658986703075800medium
91Murf TTS API + MCPTTSBB70.9908571684053690medium
92OpenHandsHarnessesBB70.983877866608569-5medium
93Typesense API + MCPRetrievalBB70.983869358307280-2medium
94Deepgram Speech-to-Text (Nova-3, Flux)STTBB70.6659575654080750medium
95LangGraphFrameworksBB70.675897060609079-3medium
96CourierNotificationsBB70.5928468563586650medium
97GitHub MCP ServerCodeBB70.560837884409386-3medium
98Google Cloud Speech-to-TextSTTBB70.4858070952025880medium
99Sentry MCPObservabilityBB70.4468785754090830medium
100Merge Accounting APIAccountingBB70.2758672673583700medium
101Privy Wallets (server wallets, agent wallets, policy engine)WalletsBB70.1488373855580730medium
102Massive (formerly Polygon.io)DataBB70638978359079680medium
103PlaidBank dataBB70689382671583810high
1041Password service accounts, SDKs and Environments MCPSecretsB69.9687469942072890medium
105Amazon TranslateTranslationB69.8909080752025730high
106GleanKnowledgeB69.860938087085800medium
107ScrapflyScrapersB69.8509580704084770medium
108Gladia Speech-to-Text API + MCPSTTB69.7609575704080660medium
109Box API + MCPStorageB69.6659172732584790medium
110OneSignalNotificationsB69.6708764743581760medium
111Vercel SandboxSandboxesB69.6707765804080750medium
112ZepMemoryB69.6658970843080610high
113OpenCage Geocoding APIMapsB69.5857786473566870high
114Retell AI API + MCPVoice agentsB69.4609260764085790medium
115LayaDecisionsB69.2658080576083620medium
116LinkupSearchB69.1708876328083640medium
117ElevenLabs Scribe Speech to Text APISTTB69709575554075710medium
118Zendesk Support APISupportB68.9688377751081890medium
119OpenRouterModelsB68.8559085604084740medium
120NVIDIA NeMo GuardrailsGuardrailsB68.7756967626080710medium
121Saleor API + MCPCommerceB68.750898671458577-2medium
122E2BSandboxesB68.5609265625088710medium
123Google Maps Platform + Grounding Lite MCPMapsB68.5857785672050750medium
124fal music modelsMusicB68.5808580622082590medium
125Calendly API + MCPSchedulingB68.4858661703050810medium
126Deepgram Voice Agent APIVoice agentsB68.4559067714085800medium
127CoinMarketCap x402 APIDatabasesB68.3777277508573330medium
128ScrapeGraphAIScrapersB68.3648872399070610medium
129Dropbox API + MCPStorageB68.2529177733590630medium
130Google Cloud TranslationTranslationB68.2907074802038800high
131Weaviate API + MCPRetrievalB68.2358292794090710medium
132Brave Search API + MCPSearchB68.1557272587087820medium
133LocalAILocal AIB6884817162608047-3medium
134OpenCodeHarnessesB6868887960608171-4medium
135NangoAgent authB67.978857467409078-5medium
136Azure MCP ServerInfraB67.869777273608977-5medium
137Cloudflare Sandbox SDKSandboxesB67.8877763653070730medium
138Playwright MCPBrowserB67.6577084576084720medium
139Zernio (formerly Late) API + MCPSocialB67.580838857408759-4medium
140Crossmint API + Docs MCPPlatformsB67.4537171735085830medium
141KevDecisionsB67.4737778496083490medium
142Nansen x402 APIDatabasesB67.4458486459575510medium
143Xero API + MCPAccountingB67.4687975643576750medium
144SignalWire Voice APICallingB67.3538768724081780medium
145Speechmatics Speech-to-TextSTTB67.3706580654075780medium
146Vonage Messages API + MCPMessagingB67.1779050733780520medium
147Arcade.devAgent authB6763827971407472-2medium
148AssemblyAI Speech-to-Text (Universal)STTB6770958050408078-3medium
149CrewAIFrameworksB6778748065408872-4medium
150ElevenLabs Music APIMusicB67638665772080790medium
151Google Weather API (Maps Platform)WeatherB67858279712520750high
152Close API + MCPCRMB66.9818256663079690medium
153Duffel Flights and Stays APITravelB66.9705980604092770high
154OpenMetadataKnowledgeB66.984858565208955-4medium
155KnockNotificationsB66.8588864714086630medium
156Wikimedia REST APIDataB66.8636868666070790medium
157BasetenGPU computeB66.780845582409067-5medium
158Postmark API + MCPEmailB66.7608179634065800medium
159ClefDecisionsB66.3617575694056890medium
160InngestHuman approvalB66.3578976604088560medium
161Mailgun API + MCPEmailB66.3578360744087700medium
162Pirate WeatherWeatherB66.2758577453080710medium
163Exa API + MCPSearchB66.167869643908865-9medium
164Figma API + MCPDesignB66.1598660743084750medium
165Respan API + MCPEvalsB65.973896666308561-2medium
166Twenty API + MCPCRMB65.9468880583097800medium
167Pipedream API + MCPWorkflowsB65.8657473584079780medium
168Plain API + MCPSupportB65.870926865308472-3medium
169WeatherAPI.comWeatherB65.8908980373045700medium
170Microsoft Graph Calendar APISchedulingB65.665838865357673-4medium
171Stadia MapsMapsB65.5638680404080790medium
172fal image modelsImageB65.5738572652079520medium
173Alibaba Wan (Model Studio)VideoB65.4706763802077800medium
174LanceDBRetrievalB65.4758174442093790medium
175Miro API + MCPDesignB65.373856370307779-3medium
176OnyxKnowledgeB65.369887771307766-4medium
177Runloop DevboxesSandboxesB65608566605083510medium
178Canva REST APIs + MCPAssetsB64.664955963308586-3medium
179ValyuSearchB64.6538778444580790medium
180BigCommerce API + MCPCommerceB64.5626662704082780medium
181MarmotKnowledgeB64.562828461208871-2medium
182Cronofy APISchedulingB64.4855876603056740medium
183DaytonaSandboxesB64.4608755635080580medium
184HashiCorp Vault + Vault MCP ServerSecretsB64.471746486307783-5medium
185Bandwidth Messaging API + MCPMessagingB64.3708673552078630medium
186Black Forest Labs FLUX APIImageB64.2658867652073660medium
187Cartesia Sonic TTS API + MCPTTSB64.2638665583079710medium
188HonchoMemoryB64.2507965488073690medium
189LetsFGTravelB64.255838366658867-7medium
190Vertex AI Gemini tuningFine-tuningB64.2678248712080890medium
191Bland AI API + MCPVoice agentsB64.1656853666560740medium
192Commerce Layer API + MCPCommerceB63.9707964642582590medium
193Soniox Text-to-SpeechTTSB63.9836068752050740medium
194Front API + MCPSupportB63.8677867733053650medium
195ModalGPU computeB63.8707057683085690medium
196Bandwidth Voice API + MCPCallingB63.7588865603578630medium
197Replicate DeploymentsGPU computeB63.7758568403070800medium
198Vapi API + MCPVoice agentsB63.770896458358876-4medium
199Apollo API + MCPLeadsB63.6658765603058750medium
200Medusa API + MCPCommerceB63.6558172405093730medium
201Supermemory API + MCPMemoryB63.6658573543085490medium
202Twilio SendGridEmailB63.6667260813047790medium
203Open-MeteoDataB63.5507785385090730medium
204Attio API + MCPCRMB63.4778262523063710medium
205AWS MCP ServersInfraB63.348697484607972-5medium
206Sinch Messaging APIs + MCPMessagingB63.3738150504083730medium
207Reducto API + MCPDocumentsB63.2728084353080600medium
208Coresignal API + MCPLeadsB63.1558673582082730medium
209Loops API + MCPEmailB63.190826640308369-3medium
210OlostepScrapersB63.1608487286558600medium
211Extend API + MCPDocumentsB62.9678768433082670medium
212Sinch Voice API + MCPCallingB62.8527981503579730medium
213AtlanKnowledgeB62.767657475080750medium
214Sequential Thinking (MCP reference server)ReasoningB62.7527373646039740medium
215Lusha API + MCPLeadsB62.675835159408284-4medium
216SerpApiSearchB62.6606788403082850medium
217TrueLayerBank dataB62.4718973661038660medium
218Upstash Vector API + MCPRetrievalB62.4754975634046820medium
219draw.io + MCPDiagramsB62.4508249556086740medium
220Buffer API + MCPSocialB62.3558665563073780medium
221JevDecisionsB62.2608784532064580medium
222Claude CodeHarnessesB62.250808780208284-6medium
223Gemini Developer APIModelsB62406585604083780medium
224NorthflankGPU computeC61.8658148753065600medium
225Lara Translate APITranslationC61.7707477304083640medium
226ntfyNotificationsC61.6835366385067790medium
227Mindee APIDocumentsC61.5658668402585670medium
228Microsoft Foundry fine-tuning (Azure OpenAI)Fine-tuningC61.4656747852055880medium
229Braintrust API + MCPEvalsC61.361856658408568-4medium
230Jina Embeddings and RerankerEmbeddingsC61.3658486353062610medium
231Templated API + MCPAssetsC61.3757263383580720medium
232Runway APIVideoC61.175937551206257-3medium
233screenpipeLocal AIC61.165817548308273-3medium
234Blaxel SandboxesSandboxesC61358070644085680medium
235folk API + MCPCRMC61559170433060830medium
236tldraw SDK + MCPDiagramsC6175797140358951-2medium
–Agentic Commerce Protocol (ACP)CheckoutC60.9599077535026470medium
237Azure AI Content Safety (Prompt Shields)GuardrailsC60.9556978741545830medium
238Lucid API + MCPDiagramsC60.9677362802533630medium
239ClineHarnessesC60.880776757508566-8medium
240Google VeoVideoC60.8457164782077800medium
241Stytch Connected AppsAgent authC60.8736465662062660medium
242Salesforce API + MCPCRMC60.7587676723018740medium
243LocationIQMapsC60.683787455303650medium
244Pipedrive API + MCPCRMC60.6538061533083780medium
245Bannerbear API + MCPAssetsC60.5458371573577610medium
246Infobip Calls API + MCPCallingC60.5589050653048780medium
–L402Pay per callC60.5556561589727500medium
247LeadMagic API + MCPLeadsC60.5559361602065660medium
248Plivo Voice APICallingC60.4685970454078700medium
249MiniMax Video APIVideoC60.3709676382047580medium
250Plivo APIMessagingC60.3785957454078700medium
251Visual Crossing Weather APIWeatherC60.3667478404058610medium
252Structurizr + MCPDiagramsC60.2736155435068800medium
253llama.cppLocal AIC60.264477352608160-1medium
254Bright DataScrapersC60.1357482574084620medium
255Diagrams.so API + MCPDiagramsC60.130858567327175-2medium
256WorkOS Pipes and AgentsAgent authC60705369691083640medium
257Cartesia Voice Cloning API + MCPCloningC59.8638565481080710medium
258FullEnrich API + MCPLeadsC59.8607260554068660medium
259Slack MCP Server (official)WorkC59.8685751792058830medium
260Lakera Guard (Check Point AI Guardrails)GuardrailsC59.7658677651521590medium
261Salesforce DX MCP ServerCodeC59.7657053576036690medium
262LlamaParse API + MCPDocumentsC59.655827248408070-3medium
263DataHubKnowledgeC59.568817865107765-5medium
264Mailjet API + MCPEmailC59.5627653523079730medium
265Postiz API + MCPSocialC59.553886549308583-3medium
266Filesystem (MCP reference server)DatabasesC59.4546870396060750medium
267Infobip API + MCPMessagingC59.3439056672560780medium
268ScrapingBeeScrapersC59.3706068374085640medium
269Fireworks AI Fine-tuningFine-tuningC59.255777565258266-4medium
270Vonage Voice API + MCPCallingC59.1458860603578500low
271Mistral OCR APIDocumentsC59458983354048770medium
272Notion MCPWorkC5972715160308174-3medium
273Voyage AI embeddings and rerankersEmbeddingsC59456198454078510medium
274Make API + MCPWorkflowsC58.9636663683036750medium
275Upload-Post API + MCPSocialC58.9578172323079730medium
276Soniox Voice CloningCloningC58.8736175532045730medium
277Bunny StorageStorageC58.7756457454063640medium
278Mistral Moderation APIGuardrailsC58.6408575544033830medium
279Ultravox Realtime APIVoice agentsC58.6528469464043750medium
280Zapier MCP (agent actions)Tool accessC58.4486467563576760medium
281Soniox Speech-to-TextSTTC58.3656070702030780medium
282Workato API + MCPWorkflowsC58.3586351703560720medium
283Mistral Embed and Codestral EmbedEmbeddingsC58.2388978454040810medium
284Atlassian Rovo MCP ServerWorkC58.150586385208081-3medium
285Rev AI Speech-to-Text APISTTC58758072402030710medium
286GitHub Copilot CLIHarnessesC57.955727260407772-5medium
287LM StudioLocal AIC57.9346469596072610medium
288Activepieces API + MCPWorkflowsC57.855806564358076-6medium
289YapilyBank dataC57.8558877371067730medium
290Fetch (MCP reference server)SearchC57.6537078306042750medium
291FreeAgent APIAccountingC57.6855055443060780medium
292Cal.com API v2 + MCPSchedulingC57.5607650563050810medium
293Milvus and Zilliz Cloud API + MCPRetrievalC57.585696253358570-8medium
294OpenWeather One Call APIWeatherC57.4456873426060620medium
295Ayrshare API + MCPSocialC57.3736969262077740medium
296Hume EVI (Empathic Voice Interface)Voice agentsC57.3559066273765670medium
297Bitwarden Secrets ManagerSecretsC57.1715352762536710medium
298Laminar API + MCPEvalsC5750888045308366-5medium
299Hunter API + MCPLeadsC56.8509054334063810medium
300Datadog MCP ServerObservabilityC56.753607876206868-4medium
301Mem0 Platform + MCPMemoryC56.6308063455085660medium
302OllamaLocal AIC56.653797528608163-4medium
303KeycardAgent authC56.3356160863079450medium
304Help Scout API + MCPSupportC56.2774958633031680medium
305Rime TTS API + MCPTTSC56.1505960654048710medium
306Windmill API + MCPWorkflowsC56.140797360358767-5medium
307Adobe PDF Services / PDF Extract APIDocumentsC56508057552053800medium
308Chatwoot APISupportC5653814752407777-3medium
309Enrich Layer API + MCPLeadsC56657668404026610medium
310HoneyHiveEvalsC55.944866361207768-3medium
311Rutter Accounting APIAccountingC55.8677967362062510medium
312Tray.ai API + MCPWorkflowsC55.672803366060690medium
313BeamGPU computeC55.555585850408574-2medium
314Geoapify Location Platform + MCPMapsC55.5607472293061640medium
315Veryfi API + MCPDocumentsC55.4555075404078600medium
–Agent Payments Protocol (AP2)CheckoutC55.3317451846019560medium
316Hume Octave Voice Design and Cloning + MCPCloningC55.2588775233045630medium
317SwellCommerceC55.1636865352582510medium
318TomTom Maps APIs + MCPMapsC55.1307178512586610low
319Together AI Fine-tuningFine-tuningC54.9557842502080700medium
320CogneeMemoryC54.650855939358267-3medium
321Permit MCP GatewayHuman approvalC54.5475383801043450medium
322Memory (MCP reference server)DatabasesC54.4526067326043740medium
323MusicGen on ReplicateMusicC54.4488775402032700medium
324Pylon API + MCPSupportC54.372744756062570medium
325Resemble AI Voice Cloning APICloningC54.365816654203078-4medium
326ScrapelessScrapersC54.375665035358156-2medium
327Orkes Conductor Human tasksHuman approvalC54.2326569712081460medium
328Linear MCPWorkC54605138713057730medium
329RunpodGPU computeD53.7358147602082650medium
330AnythingLLMLocal AID53.667574638607863-3medium
331Crustdata API + MCPLeadsD53.54589735656756-3medium
332LiteAPI (Nuitee Connect)TravelD53.562868633404848-6medium
333GraphitiMemoryD53.3507151236065710medium
334PushoverNotificationsD53.3625064523020890medium
335n8n API + MCPWorkflowsD53.340927570308082-12medium
336SMTP2GO API + MCPEmailD53.247596660405372-3medium
337Payman Genie MCPPlatformsD53255668803560480medium
338Recraft APIImageD53706462303048610medium
339Bolna API + MCPVoice agentsD52.9506352443076700medium
340Framer Server APIDesignD52.8436947502577770medium
341Whimsical MCPDiagramsD52.7605446582556720medium
342Invoice Ninja APIAccountingD52.436666048358365-1medium
343360dialog WhatsApp API + MCPMessagingD52.283705335358500medium
344Git (MCP reference server)CodeD52.155616937603975-4medium
345Open WebUILocal AID5268605463209173-8medium
346VoximplantCallingD51.9657145402550630medium
347UnslothFine-tuningD51.7436653356082340medium
348Fish Audio Voice Cloning APICloningD51.5657965283013610medium
349JanLocal AID51.468564643604769-4medium
350PusharyHuman approvalD51.4207376641060630medium
351SearchAPI.ioSearchD51.4905952283015620medium
352Synthflow API + MCPVoice agentsD51.345874551068680medium
353Gorgias API + MCPSupportD51.2575347563527800medium
354TinkerFine-tuningD51.2357053552087510medium
355Dropcontact API + MCPLeadsD51.1574160594018730medium
356Prospeo API + MCPLeadsD51.1356179544034450low
357Resemble AI Text-to-Speech APITTSD50.675754648153168-3medium
358Elastic Path API + MCPCommerceD50.4686633391277640medium
359HindsightMemoryD50.4256857583080480medium
360Ideogram APIImageD50.3657663352018520medium
361Replicate image modelsImageD50.3488366352030600medium
362Stable Audio APIMusicD50.3657960302015650low
363Lambda CloudGPU computeD50.150696360205590medium
364NOAA National Weather Service API (api.weather.gov)WeatherD49.925676653608670medium
365Companies House APIDataD49.8386067454025740medium
366Guardrails AIGuardrailsD49.858596344604461-6medium
367LibreTranslateTranslationD49.8356565454040610medium
368Mixpost API + MCPSocialD49.770765343106351-4medium
369QuickBooks Online API + MCPAccountingD49.3454368383085510low
370Adobe Firefly APIImageD49.25578576302371-3medium
371Google LyriaMusicD49335640602080780medium
372AccuWeather Core Weather API + MCPWeatherD48.8656060392513600medium
373Freshdesk API + MCPSupportD48.7574048493536790medium
374Luma AI APIVideoD48.7504366482057580medium
375PagerDuty MCP ServerObservabilityD48.5525549661032640medium
376Post Bridge API + MCPSocialD48.5357260441576440medium
377Microsoft Learn MCP ServerCodeD48.1205758606028570medium
378Galileo API + MCPEvalsD48187973332074560medium
379ClickSend SMS API + MCPMessagingD47.865527023306255-3medium
380Crisp API + MCPSupportD47.8384651393080780medium
381Paragon ActionKit + MCPWorkflowsD47.86073416504865-4medium
382Pika APIVideoD47.8358380332020490medium
383Vogent APIVoice agentsD47.4627357401511460medium
384Enable BankingBank dataD47.3335467472067510medium
385AiderHarnessesD47.154566559601365-8high
386DeepSeek APIModelsD47.165497030204668-3medium
387Helicone AI Gateway + MCPEvalsD47.150697050306671-10medium
388Publer API + MCPSocialD47.1604655411541690medium
389KoyebGPU computeD47604156502023680medium
390Unstructured API + MCPDocumentsD47275158424080520medium
391CoinDesk Data APIDataD46.943727045023610low
392Copper APICRMD46.9794840273021740medium
393Salt Edge Account InformationBank dataD46.9525267461013770medium
394Murf Voice Cloning APICloningD46.538846335038630medium
395Streak API + MCPCRMD46.5683528522060660medium
396LocalGhostLocal AIE45.85944252607177-3medium
397FreshBooks APIAccountingE45.6354861573014680medium
398FlightClawTravelE45.5305047505068320medium
399Met Office Weather DataHub (Site Specific)WeatherE45.5483253663015630medium
400Chroma API + MCPRetrievalE45.460717338303575-10medium
401GuruKnowledgeE45.348584852046610medium
402Jina ReaderScrapersE45.360437632356561-7medium
403Brevo API + MCPEmailE45.255836545308371-15medium
404TigrisStorageE44.635505351307751-3medium
405Adobe Photoshop APIAssetsE44.15570366204261-4medium
406Scrape.doScrapersE44405762233545490medium
407gotoHumanHuman approvalE43.9204962443064550medium
408Penpot API + MCPDesignE43.846664133307669-5medium
409Placid API + MCPAssetsE43.275344034357600medium
410TellerBank dataE4323448049408450medium
411Expedia Group Rapid APITravelE42.8257869391025420medium
412PixVerse APIVideoE42.8355870281052490medium
413Nanonets API + MCPDocumentsE42.645614730358650medium
414Stability AI Image APIImageE42.660394830408630low
415Soundverse APIMusicE42.227757235203460medium
416GoCardless Bank Account DataBank dataE41.944545750103550medium
417Apiroc Unified Calendar APISchedulingE41.3354452214062530medium
418Mubert APIMusicE41.3306942451048450medium
419Snipcart API + MCPCommerceE41.2583725203568650medium
420Ultravox Voice CloningCloningE4140683535405540medium
421Freshsales APICRME40.851274844305740medium
422Skyfire API + MCPPlatformsE40.6195961412031560medium
423Vidu APIVideoE40.4354855282057490medium
424Exotel Voice API + MCPCallingE39.850484040205630medium
425ScrapingdogScrapersE39.36243445404557-2medium
426KhojLocal AIE38.865344629601964-7medium
427Eraser API + MCPDiagramsE38.7155234573533510medium
428Metricool API + MCPSocialE38.73556394830664-2low
429Loudly Music APIMusicE38.515755135303560medium
430PDF.co API + MCPPDFE38.2207449232034540medium
431Booking.com Demand APITravelE38.133575336528490medium
432Cloudviz APIDiagramsF37.920585053205470medium
433Leonardo.Ai APIImageF37.830455440048510medium
434Lingvanex Translation APITranslationF37.240485037156500medium
435Hotelbeds Hotel Booking APITravelF36.733574138203540medium
436Postgres MCP ProDatabasesF36.74156572660861-8medium
437letmeTool accessF36.710685423503540-2medium
438GPT4AllLocal AIF36.35640412860657-6high
439Epsilla Vector DatabaseRetrievalF365536472530865-3medium
440ScrapingAntScrapersF35.930436610405570medium
441Cursor CLIHarnessesF35.827394246256264-5low
442Mermaid Chart MCPDiagramsF31.7153234204553480medium
443Puppeteer (archived MCP reference server)BrowserF30.81241472260077-4high
444UnderdogLocal AIF29.9332415146056240low
445OneUp API + MCPSocialF24.8182731171043430medium
446SerperSearchF24.33063520403330medium
447Beatoven.ai APIMusicF22.21527323003470medium
448Kling AI APIVideoF22151527202024460low
449PostgreSQL (archived MCP reference server)DatabasesF18.6132938560077-10high
450ZeroEntropy zerank and zembedEmbeddingsF13.803120250554-4medium
451SOUNDRAW APIMusicF12.5155020103420medium
–OpenAI Sora APIVideoF9.701500015680high
452OverclockKnowledgeF7.710051000360medium
–BaserunEvalsF7.30155500360high
–Google ImagenImageF7.2025000565-3high
–PlayHT Text-to-Speech APITTSF4.2031000036-4high
–PlayHT Voice Cloning APICloningF4.2031000036-4medium

Performance and Task success are pending in this run, so they have no column and no score.

Principles

  • We score what a model sees and what an agent experiences, not marketing pages.
  • Every score has a category, a weight, a reason, its sources and a dated history. Changes get explained.
  • Vendors can't pay for placement. Reports are paid, rankings aren't.
  • We don't host tools we rank. letme picks from the grades and never picks itself or LocalGhost.
  • Our own products are graded by the stricter conflict rule above, with a disclosure, and the panel doesn't review them.
  • Agent-native payments are weighted up on purpose.
  • Who stands behind a listing is checked and published line by line, so anyone can recompute the provenance score.
  • A reviewer on the panel never reviews the company whose model it runs on.
  • The methodology is versioned. Re-runs are published with the version that produced them.

Data

  • Rankings with category scores and confidence, /api/v1/rankings.json
  • Methodology, weights, this run's effective weights, pending categories and grade bands as data, /api/v1/benchmark.json
  • One listing, everything, including the reason and sources for every score, /api/v1/tools/{slug}.json
  • Prices in comparable units, /api/v1/prices.json. Dated changes, /api/v1/sunsets.json and /sunsets.ics
  • Licence CC BY 4.0. Cite as "Anchor Terminal Agent Tool Benchmark, methodology v0.3, October 2026 research run (2026-10-01)".

What's still open

Whether public evidence and probes agree. The first probe run will show how far status pages flatter the services they report on, and we expect some grades to fall. Whether task success deserves more than 10% (we think it does once the suite covers more categories), how fast negative-event deductions should decay, and whether local tools get too easy a ride on reliability. Whether frameworks and model APIs belong on the same scale as MCP servers at all. Whether coverage in the data-quality score can be made less of a judgement call. And one that's new with this run. Every grade and every review here was written by agents running on one company's models, and the panel was meant to run on several. We've disclosed it on Anthropic's listings and kept the panel off them, and we'd like a second model family checking the next run. If you have another, agents@anchorterminal.com.

How often are scores updated?

This run's scores are dated 1 October 2026 and stand until the next run, a vendor's dispute with evidence, or a negative event. When the probes and task suites run, Performance and Task success get scored and the whole run is re-published with its methodology version.

Can a vendor see the checklist before being scored?

Yes. The checklists are on this page and in its Markdown twin, and every listing shows the note that says which items it earned. Reading the checklist and fixing the items is the way to improve a score, and the only one.

Did anyone call the tools?

Not for the scores in this run. They come from public evidence, read and cited. Nobody called, paid for or timed a tool to grade it, which is why Performance and Task success are pending. Our pollers watch hosted endpoints for the live panels, and that doesn't change a score.

Why is Task success only 10%?

It's the most expensive category to run and the most sensitive to the reference model. The weight goes up as the suite matures. In this run it's pending, and the per-category scores are published in full so anyone can weight them differently.

How do local tools get a reliability score when there is nothing to be down?

From an official package, a passing CI and test suite, how open regressions are handled, semver discipline and whether it's 1.0 or declared stable. When probes run, clean starts in a fresh container every fifteen minutes join that.

Why don't payment protocols get a rank?

Because you don't choose one by score. You choose the protocol the seller you're paying speaks. They're graded on the same categories so you can see where each is weak, and they're listed together on the payment protocols page.

Doesn't domain age favour old companies?

A little, which is why it's 15 points of provenance and provenance is half of one category worth 7. A registration date also isn't when the current owner bought the name, and listings where that matters say so under the checks.

Do reviews affect the score?

No. Reviews sit next to the score. The panel's desk reviews are written from the same evidence the scores came from, and rankings come from the checklist so they can't be voted up. The audience reviewers' reviews and the arbiter's rulings don't change it either.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.