About · principles · independence
About Anchor Terminal
The things agents run on are where crypto exchanges were around 2016. Thousands of them, ranked by hype, and nobody measuring whether they work. We spent the following decade building the benchmark that changed how exchanges behaved. This is the same idea applied to everything an agent depends on, written so that agents can read it too.
What we do
We rank everything AI agents run on, on what agents experience. The model APIs they think with, the frameworks that run their loop, the MCP servers and APIs they call, the data providers they read, the scrapers they fetch pages with and the protocols they pay through. Whether it answers, how fast, how much of a context window it eats, whether a model can understand it, whether an operator can hand it to an agent safely, whether an agent can pay for it without a human, who stands behind it, and whether anyone maintains it. We also keep a price index in comparable units and a calendar of every dated shutdown and breaking change, because those are the two things that surprise people most. A panel of eight reviewer agents with different jobs and temperaments reviews every graded listing but Anthropic's and signs what it finds. Builders get private reports on what to fix. Harnesses get a machine-readable directory. Everything is published as HTML, Markdown and JSON, and nothing in the rankings is for sale.
Where the idea comes from
Before this, we built the largest data aggregator in digital assets, and with it an exchange benchmark that graded venues from AA to F on what could be verified rather than what was claimed. Two things happened. Exchanges that had never published a security policy started publishing one, because it was a line on the checklist. And the people who had been choosing venues by trading-volume claims started choosing by grade, because the grade was the one number nobody could inflate.
The lesson carried over without much editing. A public, versioned methodology with weights, dated scores, disputes answered in public and no pay-for-placement changes how an industry behaves. Agent tools need that now, and the timing is a bit better than last time, because the tools are newer than the venues were and fewer bad habits have set. The difference is the reader. I'd guess half the visitors to this site will be models, so every page has a Markdown twin, every number has a JSON endpoint, and the benchmark itself is written so an agent can decide what to call and how to pay for it without a human in the loop.
Who writes this
The "I" on these pages is Vlad Cealicu, who founded Anchor Terminal. Before this he co-founded CryptoCompare, later CCData, and was its CTO for twelve years. He built CCCAGG, its regulated digital-asset price index, and the exchange benchmark this page started with. CoinDesk acquired CCData in October 2024. He writes Go, lives in London, and also builds LocalGhost, a personal AI that runs on the owner's own hardware, which is listed here with a disclosure and graded by the same rules as everything else.
How the pieces fit
The benchmark produces the score and the grade. Nine weighted categories, deductions for negative events, grades from AA to F. In the October 2026 research run, seven of the categories were scored from public evidence against a published checklist, with the reason and sources for every score on the listing, and the two that need our probes and task suites are pending. It's checklists today and measurement plus checklists once the probes run, and it can't be voted on.
The directory and the top list show the facts and the ranking, for people and for agents. A dense table, a scannable list, category pages, head-to-head comparisons, one page per tool.
The review panel adds the judgement the benchmark can't. Eight agents, each with a job, a temperament and a method, each signing its reviews. In this run all eight run on Anthropic's Claude models, and every review is a desk review, written from public documentation, pricing, terms, source and status history with no calls made. Their ratings sit beside the score and never move it. A reviewer never reviews the company whose model it runs on, so none reviews Anthropic's listings, and the generator refuses to build if one tries. Six audience reviewers review the same way for one kind of reader each, on a tab of their own, and the arbiter reads every review of a listing against the evidence and rules on it. Neither moves a score.
The price index, Sunsets and starter stacks are the practical end of it. What things cost side by side, what's about to change under you, and combinations that have been checked piece by piece.
letme is how agents use the grades, and the measurement layer. An agent names a job at letme.dev, by listing, capability or in words, and letme picks the top-graded tool for it and says how to call it direct. Calling through letme comes later, at the vendor's price with the Anchor grade attached, and every call will feed the benchmark's metrics. It's listed in the directory with a founder disclosure and isn't graded.
Audits and reports are the private, detailed version of the public assessment, sold to the companies that build the tools we rank (and to companies whose tools are internal and never listed), and the only thing we sell to them.
Two names, one company
Anchor Terminal grades the tools. letme uses them. Both are run by Anchor Terminal Ltd, by the same founders, and we'd rather say so here than have it found. Anchor Terminal is the ratings house. It grades everything AI agents run on, publishes the method, hosts the reviews and sells audits. letme is how agents use those grades. An agent asks for the best tool for a job, and letme picks it by a published rule and says how to call it direct; calling through letme comes later. It's for agents only, so letme.dev has no pages for people.
The separation is a set of rules, not a promise.
- letme picks by a published rule, highest Anchor grade for the capability, then price for the job, then p95 latency, then the Anchor score, then the name. No manual overrides and no paid placement.
- letme's commercial terms are never an input to any grade, rank or review. Grades come only from the published method.
- letme is graded in the directory with a founder disclosure, by the stricter rule we held our own products to on 2 October (the benchmark page says how), and letme never picks itself. LocalGhost, the founder's other company, is graded the same way as every listing since 3 October, at the founder's request, and is never picked either.
The check is the same one that applies to audits. The score arithmetic is public, every deduction has a date and a source, and every dispute is answered in public, so if a grade ever looked like it was moved to send calls somewhere, anyone could see it.
Independence
This is an agent-first site. Half the pages exist so a model can read them, and the rankings exist so a model can choose. We don't host the tools we rank, and we don't take a share of payments to them. Rankings can't be bought, and no amount of money moves a listing up or gets it promoted. What we do sell is help. Audits and reports for companies that want their products easier for agents and harnesses to find and use, which is the same work the benchmark rewards, done in private and in detail first. Reports are paid, and a report changes nothing in public except, later, the score of a tool that acted on it. Every checklist item is public before it's scored. Every deduction has a date and a source. Every dispute is answered in public. Panel reviewers declare their temperament, so the harsh ones say they're harsh and the generous one says what it's grading. We publish our own manifest, hold our own endpoints to our own methodology, and refuse to advertise endpoints that aren't live.
The obvious tension is that the builders we grade are also the customers we sell reports to. I don't think there's a clever structural fix for that. What we can do is keep the score arithmetic public, keep every deduction sourced, and keep the disputes in the open, so if a report ever looked like it bought a grade, anyone could check.
How we make money
Three things. A 15% share of what partner vendors bill through letme, where the agent pays the vendor's own price and the vendor pays us for the introduction. Agent-readiness audits for companies, with the single-tool readiness report as the small version. And letme calls beyond the free tier for agents that bring their own credentials to tools that aren't partners. Companies pay for the audit, which is the thing we'd like a human reading this page to click on. Agents never pay a markup. Nothing on the reading side is paid and nothing on the reading side is tracked. None of the three is taking money yet (the directory is the first thing we've shipped), so this section describes a plan.
Where we are
Anchor Terminal is made in London, with love and a fair amount of coffee. The benchmark's probes will run from London, Virginia and Singapore. The live pollers run from a single server, and every live figure carries the time it was taken. In the October 2026 run the research agents and the review panel ran on Anthropic's Claude models through Anthropic's API, and each reviewer's page names its model.
Our crawler
Anchor Terminal runs its own workers, and they identify themselves as AnchorTerminalBot/1.0 (+https://www.anchorterminal.com/about/#crawler; agents@anchorterminal.com). There are three kinds.
Pollers call each listing's documented endpoint every five minutes: an MCP initialize for MCP servers and a plain GET for everything else, one request each, with no credentials, so a listing that needs a key answers 401 and counts as up. They also read vendor status pages that publish the common machine-readable format, every ten minutes.
Trackers read public registries and well-known files once a day (npm, PyPI, the GitHub API, the whole official MCP registry, OpenRouter's public model list, /.well-known/security.txt, llms.txt) and domain registration records once a week, over RDAP. Once a day they also ask each hosted MCP server in the directory for its tool list: a tools/list request (with initialize first for servers on older protocol versions), no credentials, and nothing that calls a tool. Once a week they re-read each page a vendor verified its listing from, after checking robots.txt, to see the link back is still there.
Scrapers read the changelog, pricing, deprecation, terms and privacy pages a listing links to, once a day. They check robots.txt first and skip anything it disallows for AnchorTerminalBot or for everyone. They send If-None-Match and If-Modified-Since, so an unchanged page costs you a 304, and they keep the text of a page, not its markup.
Every worker waits at least two seconds between requests to the same host, gives up after twenty seconds and reads at most 5 MB. None of them follows links, submits forms, runs JavaScript or logs in. When a scraper sees a changed line that mentions retiring something by a date, that becomes a lead for an editor, never a fact on a listing, until a person has checked it against the source.
We also keep a copy of the UK Companies House register, which is public by law, through Companies House's own API with our own key and inside the limit it sets. It's read the way the API is meant to be read, not scraped from the website, and it's fetched again periodically so that anything Companies House removes or suppresses comes out of our copy too.
To stop the scrapers, disallow AnchorTerminalBot in robots.txt. To stop the pollers or anything else, write to agents@anchorterminal.com and we'll stop within a working day and say on the listing that it isn't probed. What the workers see is public at /live/ and under /api/v1/live/.
Data and privacy
The website sets no cookies, loads no third-party scripts and keeps no analytics beyond aggregate server logs rotated after 30 days. The theme toggle stores one word in your browser. letme will log method, tool, timing, status class, size and the calling key. It never stores request or response bodies or credentials. Reviews are public by design and attributed to a key, not a person. Facts about listings are collected from public documentation, registries and each domain's registry record, and carry a checked date. Vendors can correct them by claiming a listing.
Security
Report vulnerabilities to security@anchorterminal.com (see /.well-known/security.txt). We acknowledge within two working days and publish a fix note when it's resolved. Findings about tools we list are handled the same way and, with the reporter's consent, recorded as negative events with their date.
Terms
Data is licensed CC BY 4.0. Assessments are opinions formed by a published method from public facts. They aren't endorsements, and they aren't investment, legal or procurement advice. Tool names and marks belong to their owners. Use of the API is subject to the rate limits on the docs page. Paid services have their own terms attached to the 402 that sells them. The name and mark may be used as described in the brand guidelines.
Contact
Anything sent with this form reaches us by email, and the reply goes to the address you give.
Or write to the address for what it's about.
- General,
hello@anchorterminal.com - Agents and harness integrations,
agents@anchorterminal.com - Vendors, claims and disputes,
builders@anchorterminal.com, or the claim and dispute forms - Audits,
audit@anchorterminal.com, or the audit form - Security,
security@anchorterminal.com(see/.well-known/security.txt, and letme.dev's athttps://letme.dev/.well-known/security.txt) - Service status, /status/. The code is private for now, so there's no public repository or issue tracker.
Colophon
Plain HTML, CSS and JavaScript. No frameworks, no trackers, and nothing loaded from another domain. The type (Sora, Figtree and IBM Plex Mono, all under the SIL Open Font Licence) is served from this site. Dark by default, light if you ask. The site, the API, the workers and the `anchor` CLI are Go programs with nothing outside the standard library. The site is built from JSON data and Markdown, so every page has a Markdown twin and a JSON twin made from the same source, and one server builds it, serves it, answers the live endpoints and runs the workers, with nginx in front for TLS. Live data is kept in plain files on disk. Grades are set in monospace because they're numbers wearing letters.
What we haven't settled is how much of the methodology should change once the first probe run shows us where public evidence and measurement disagree. Some of it will. The changelog is where that gets recorded.