{
  "data": {
    "faq": null,
    "kicker": "Changelog · site · API · methodology",
    "lede": "Every change to the methodology, the API or the data model goes here with a date. Deprecations are announced at least 90 days ahead, which is the standard the benchmark holds tools to, so it would be awkward not to."
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/changelog/",
    "json": "https://www.anchorterminal.com/changelog/index.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/changelog/index.md",
    "slim": "https://www.anchorterminal.com/changelog/index.min.md"
  },
  "markdown": "## 2026-10-04, the rank column says why, and indexed listings told apart\n\n- Where a listing has no place in the ranking, the # column says why, where it used to show a dash. \"spec\" is one of the five payment protocols, graded on the same scale and [not ranked against tools](/benchmark/), and \"shut\" one of the five services that have shut down. Hover over it for the reason. The Markdown tables say \"not ranked, protocol\" or \"not ranked, shut down\", and in the API every listing's `anchor` now has `ranked` and, when it's false, `notRankedWhy`.\n- Indexed listings that share a name with another one (three servers in the MCP registry called \"cli\", one publisher's two entries both called \"Yarr MCP\") carry their registry name in the page's title and description, so each is a result of its own in search. OpenRouter's \"latest\" aliases no longer make a second family of the same name (a second Claude Sonnet, Gemini Pro or Grok).\n\n## 2026-10-04, the link-back panel on listings, and the comparisons with other sites re-checked\n\n- The panel on each listing that asks its maker to link back is now \"Verify this listing\", in two numbered steps. The badge is shown as it looks on a light page and on a dark one, the HTML, Markdown and plain-link versions are a click apart, and the form to tell us where the link is comes second. It says which pages the link counts on and that we check it weekly. A link back still never changes a grade, rank or review.\n- Every [comparison with another site](/compare/#us) was checked again on 4 October against that site's own pages, one research agent each, with the sources at the foot of each page. Three rows changed. mcp.so's open source is now partly (its public code hasn't changed since March 2025), Merge Agent Handler's API for agents is now yes and its open source partly, and Nango's API for agents is now yes. Scale, prices and notes are as of 4 October on all of them, and the Smithery and Arcade pages say Smithery has been part of Arcade since August 2026.\n- Three more comparisons, with [Docker MCP Catalog](/compare/anchor-terminal-vs-docker-mcp-catalog), [MCP Market](/compare/anchor-terminal-vs-mcp-market) and [MCP Toplist](/compare/anchor-terminal-vs-mcp-toplist).\n- On a phone, the side-by-side table on these pages shows each row as a block, ours and theirs one under the other, where it used to be a table you scrolled sideways.\n\n## 2026-10-04, how the site reads to search engines\n\n- Every page's title now fits in a search result (70 characters at most) and every meta description in 160, cut at a whole sentence. The pages keep their full text.\n- The Markdown and JSON copies of each page, the API and llms.txt tell search engines not to index them (`X-Robots-Tag: noindex`), so a search finds the page itself. Agents read them as before.\n- The three pages for businesses carry FAQ markup for the questions they answer, and the [about page](/about/#author) says who writes the site.\n- Logos are served as WebP, about a fifth of the bytes of the PNGs, and every image has its size set so nothing moves as it loads.\n- On a phone the Terminal's filter panel now starts closed before the list is drawn, where it used to close once the page had loaded and the list jumped up. The price index draws its tables below the first only when you scroll near them, and paints in under a second on a slow connection where it took 2.4.\n- The build checks all of this on every page from now on (canonical, title, description, one H1, image sizes and alt text, breadcrumbs, links, sitemap, robots.txt), and fails rather than publish a page that breaks a rule.\n\n## 2026-10-04, search suggestions, and images get their own rate limit\n\n- The search box in the header shows the ten highest-ranked listings, with their rank, grade and score, as soon as you click into it, and a link to the full top list. Typing switches to matching listings and categories as before.\n- The search box on the homepage suggests too, the same way and from the same index: the top ten before you type, then matches, with the arrow keys and Enter working as in the header.\n- Logos, badges, stylesheets, scripts and fonts now have a rate limit of their own, 200 requests a second per IP with a burst of 1,000. Until today they shared the 10 a second (burst 20) that pages and the API have, so a directory page, which asks for 400-odd logos at once, ran out after about 20 and showed initials for the rest. Pages and the API keep 10 a second, and a 429 says which limit applies. [The docs](/docs/) say so.\n\n## 2026-10-03, Local AI, company knowledge and 16 new listings\n\n- Self-hosted personal AI is now [Local AI](/categories/local-ai), models and assistants that run on the owner's own hardware, and the old category address redirects there. LocalGhost and Underdog moved with it, and ten new listings join them, graded by the normal checklist and reviewed by Keel and Warden. LocalAI B 68, screenpipe C 61.1, llama.cpp C 60.2, LM Studio C 57.9, Ollama C 56.6, AnythingLLM D 53.6, Open WebUI D 52, Jan D 51.4, Khoj E 38.8 and GPT4All F 36.3.\n- Khoj, Open WebUI and Screenpipe are the three competitors LocalGhost's own about page names, so they're graded by the rule for [competitors of our own products](/benchmark/#competitors), as Underdog was. Two research agents graded each one independently, and a third reconciled them item by item, checking the evidence itself wherever they disagreed. It found both graders had undercounted Open WebUI's security advisories (129 in the past year for flaws fixed since November 2025, each fixed before it was published), that one of Khoj's deductions rested on a version that never had the flaw, and that Screenpipe switched remote support-log uploads on by default on 17 September. The disclosure sits at the foot of each listing and of each comparison with LocalGhost.\n- A new category, [Company knowledge \u0026 data catalogues](/categories/company-knowledge), for search across what a company knows and the catalogues that describe its data. Overclock (from Retrieval \u0026 vector search) and Marmot (from Databases \u0026 files) moved there, and six new listings join them, reviewed by Scout and Warden. Glean B 69.8, OpenMetadata B 66.9, Onyx B 65.3, Atlan B 62.7, DataHub C 59.5 and Guru E 45.3.\n- Category names keep their acronyms mid-sentence, so a page reads \"local AI\" where it read \"local ai\".\n\n## 2026-10-03, the full panel on the top 50, audience reviewers and the arbiter\n\n- The 50 highest-ranked listings the panel can review (Anthropic's and our own are left out) now have a review from all eight panel members, where each had two. That's 300 more desk reviews, written on 3 October from the same research dossiers, signed like the rest.\n- Six [audience reviewers](/reviewers/#audience) review the same 50 for one kind of reader each. Pip for indie developers, Harbour for enterprise platform teams, Flint for startup CTOs, Tally for compliance in regulated industries, Mosaic for no-code operators and Lantern for privacy-first self-hosters. Their 300 reviews sit on an Audiences tab on each listing, apart from the panel's numbers and the score, and in a listing's JSON under `audienceReviews`.\n- [The arbiter](/reviewers/arbiter) reads every review of a listing beside the research dossier and marks each one upheld, corrected or rejected, says where the reviewers agree, rules where they disagree on a fact, and says which readers the listing suits best and least. Its ruling sits at the top of both tabs, each review carries its standing, and the ruling is signed with the arbiter's own key. Of the 700 reviews of the 50, it upheld 684 and corrected 16, most often for a detail the dossier doesn't support, and rejected none. It never changes a score, a rating or a review.\n- LocalGhost is graded the same way as every listing from today, at our founder's request, with the same readings, neither stricter nor looser, by two research agents and an auditor who checks the evidence wherever they disagree, as for Underdog. It moved from F 26.5 to E 45.8. Ergonomics stays at 2, because nothing an agent can call reaches the box, and the deduction is 3, for a privacy page that still says no position leaves the devices while the app sends location fixes to Android's geocoder. The panel still doesn't review it and letme still never picks it. letme keeps its grade from the stricter rule until its next check. [The benchmark page](/benchmark/#own) says so.\n- Fixes the reviewers found. Google Calendar's summary said its MCP server has 8 tools, and it has 9. Cloudflare R2's changelog link now points at Cloudflare's main changelog, which carries R2's entries since April. The reviewers' brief for Descope had Pro at 5,000 monthly active users, and the pricing page says 10,000.\n\n## 2026-10-03, Underdog, and LocalGhost re-checked after wisp 0.0.3\n\n- [Underdog](/tools/underdog) by Conway Research is listed in Self-hosted personal AI, graded F at 29.9. It's a personal AI that runs on the owner's Mac with Conway's own small open-weight models (Woof, Bark and Underdog 27B, Apache-2.0 on Hugging Face). We found no API or MCP server for agents.\n- It competes with LocalGhost, which our founder builds, so it's graded by a new rule, [competitors of our own products](/benchmark/#competitors). It's the normal checklist, neither stricter nor looser. Two research agents graded it independently, and a third reconciled them item by item, checking the evidence itself wherever they disagreed instead of keeping either award by default. Neither grader looked at LocalGhost. The panel reviewed it like any listing (Quill and Warden, two desk reviews), and letme can pick it.\n- underdog.ai refuses our reader, so its terms, privacy policy, release notes and security contact count as absent, and each note says the score is low because we couldn't read them, not because they're missing. The text of its pricing page was supplied by our founder (\"100% free\", \"Free forever\"), so Payments counts the app as free with nothing to buy, 60 under the self-hosted rule, the same as any free package run on the owner's machine.\n- The disclosure for a competitor sits at the foot of its page. Every comparison now shows the disclosures its two listings carry (ours, a competitor's, Anthropic's) at the foot of the page too.\n- Self-hosted personal AI now covers personal AI on the owner's own computer as well as a home server.\n- LocalGhost was re-checked after wisp 0.0.3 (3 October), by the same stricter rule as before. It moved from 27.6 to 26.5, still F. Reliability 44 to 47, with CI's server job passing at the release and at HEAD. Security 32 to 29 and Maintenance 39 to 38, on stricter readings of the same facts and the release toolchain, Go 1.25.4, which trails 35 standard-library advisories. The deductions went from 6 to 7 points. Two fixed problems now count less (the phone's Open-Meteo position call, removed in 0.0.3, and the homepage answers on FIDO2 keys and the Mist), and two live ones are new. The app hands each new location fix to Android's system geocoder while the README, the privacy page and llms.txt say no position leaves either device, and the essay and setup guide say the box opens no connection while it polls exchanges every minute.\n- One thing went wrong in that re-check. The listing file both graders worked from still carried the previous grade, which the brief meant to keep from them. The auditor checked every award that matched it against the evidence, and three gave way to the other grader's lower award.\n\n## 2026-10-02, LocalGhost re-checked\n\n- LocalGhost was re-checked the same evening, after its maker made changes, by the same rule as our own products' first grading: two research agents graded it from scratch without seeing the earlier grade, and a third reconciled them item by item, keeping the lower award unless it rested on an error it checked. It moves from F at 19.8 to F at 27.6, and the [listing](/tools/localghost) has every reason and source.\n- What moved it up. CI was added, and its server job has passed since 20:24 UTC after four failing runs. The three July contributions were merged or closed. A second release, v0.0.2, came with a changelog page. A named company (LocalGhost.ai Ltd), privacy and terms pages, a status page and a valid security.txt went live, which lifts its provenance checks to 82. llms.txt names the current release.\n- What held it back. No agent can reach it, so Ergonomics is 0 and Payments loses the autonomous-onboarding line, read strictly; neither is a change in the product. The device's private key is in the query string of the `localghost://enroll` link the pairing QR carries, which takes 10 off Security's credential line. The app's tests have no passing run yet.\n- Deductions stay at 6 points, for different reasons. \"Nothing leaves the box\" (3) is gone, since the pages now list the box's fetches, and the old README decays to 1 after its rewrite. Newly counted, the README and privacy page say no location leaves while the phone sends a rounded position to Open-Meteo for weather (3), and the homepage FAQ describes FIDO2 keys and the Mist as shipped (2).\n\n## 2026-10-02, letme fixes from its own fix list\n\n- The whole pick rule is published. It now reads \"highest Anchor grade for the capability, then price, compared only between prices whose item names the job, in one unit, then p95 latency, then the Anchor score, then the name\". The score and name steps always ran; the rule string in every answer, the OpenAPI file, the Agent Skill, the home page and [/letme/](/letme/) now say so.\n- `why` is exact when price doesn't decide. It says the rule compares only prices whose item names the job and names the step that did decide. Each listing in a pick answer carries `jobPrice`, the price the rule read, apart from `leadPriceUsd`, the lowest price for anything, so a per-minute price whose item doesn't name the job (Batch, Nova-3) can't read as one the rule compared. The rule itself and every pick are unchanged.\n- letme.dev's limit is published: 20 requests a second per IP address, with bursts of up to 40, and over that `429` with `Retry-After`. nginx enforced it before, and nothing said so. It's in the index, the Agent Skill, the OpenAPI file and on /letme/, and a test keeps those and the nginx config the same.\n- A status page, [/status/](/status/), with uptime by day for the last 31 days and every outage in the last 90 for our own services (letme.dev today), from the probes every five minutes, and `https://letme.dev/status` for how letme.dev is right now. The pollers now keep every endpoint's outages for 90 days and its probes by day, and every listing's live record carries them.\n- letme.dev has a security.txt, and `/.well-known/` reaches letme.dev through nginx, which had hidden it (`/.well-known/api-catalog` answered 404). Both security.txt files expire on a fixed date, 1 October 2027, moved on by hand, and the build fails within 30 days of it.\n- No GitHub link. The code is private for now, and github.com/anchorterminal answered 404, so the link is gone from the footer, the about page and the organisation's structured data.\n- letme's grade doesn't move for these until its next check, the same as any listing's.\n\n## 2026-10-02, our own products graded, and a fix list on every listing\n\n- letme and LocalGhost are graded. Until today we listed our own products and didn't grade them. Now they're graded by the same checklist as every listing, under a stricter rule set out on the [benchmark page](/benchmark/#own). Two research agents graded each one independently and read every judgement call strictly, counting only what was live and public on 2 October 2026. A third agent, which graded neither, reconciled them item by item, keeping the lower award unless it rested on an error it checked, and kept a deduction only where it could still see the problem. The panel doesn't review them, and letme never picks them. letme grades F at 36.7 and LocalGhost F at 19.8, and every reason, source and deduction is on their listings.\n- `own: true` marks our own products in every tool record, and letme.dev answers a request for `letme` or `localghost` with a `404` and `\"error\": \"own_listing\"`, saying why it leaves them out. A listing can still be `graded: false`, and none is.\n- A fix list on every graded listing. Everything the grade says a listing lacks, in one place, the biggest possible gain to the total first, in the Score panel with a button that copies it for a coding agent, at `/fixes/{slug}.md` for an agent to fetch and `/fixes/{slug}.json` for a program. It's put together from what the listing already publishes (the reason for each score, the checklist that category was scored against, the provenance checks, the deductions, what we couldn't check, the weaknesses, the agent notes and what the panel asked for), so nothing in it is new judgement, and it doesn't promise a score. A fix counts at the next check, once it's public.\n\n## 2026-10-02, out of preview: the October 2026 research run\n\n- The site is out of preview. Until today every score, grade, rank, probe metric, rating and review was an invented sample that illustrated the method. They're gone, and the sample strip with them. Nothing on the site calls a score, grade or review a sample any more, and where a page needs to say how scores were made it links the [method](/benchmark/).\n- The October 2026 research run. Between 30 September and 2 October, research agents running on Anthropic's Claude models researched every graded listing from public evidence (status history, rate-limit and error docs, API references, OpenAPI files and MCP tool definitions, pricing, terms and privacy, security pages, changelogs, repositories and registries), with about twelve page fetches per listing and no web search, and scored seven categories against a published checklist. Methodology v0.3.\n- The checklist is published. The [benchmark page](/benchmark/#checklist) has every category's items with their points, the rules the run followed (couldn't check isn't the same as absent, vendor claims are claims, nothing invented) and how confidence was set.\n- Two categories are pending. Performance and Task success need our probes and task suites, which haven't run, so no listing has a score for them. They show as pending, with no number, on listing pages, twins, JSON, the estimator, the weights table, rankings, comparisons and social cards. A total is the weighted mean over the seven assessed categories, Σ(score × weight) ÷ 80, then the negative events, so the maximum is still 100 and the grade bands are unchanged. Each category's share of the 100 points in this run is published beside its weight (`effectiveWeights` in `/api/v1/benchmark.json` and `/api/v1/rankings.json`). A data provider's data-quality score is still computed and published on its listing, and becomes half of Task success when that's scored.\n- Score reasons and sources. Every graded listing shows, under each score, the note that says which checklist items it earned and why, plus its confidence (high, medium or low), \"What we couldn't check\" and a sources list with the date each page was read. The Markdown twin carries all of it and the slim twin a one-line reason per score. `/api/v1/tools/{slug}.json` and the page's JSON twin carry the whole `anchor.assessment` block and a per-category `anchor.breakdown`; `/api/v1/tools.json` carries `assessment.confidence` and `assessment.date`. A graded listing without an assessment fails the build.\n- Probe metrics removed. The sample latency, availability, error and context-cost numbers are gone from every listing, and `metrics` says `measured: false` until the probes run. The p95 and context-cost columns in the directory start hidden. The live panels, from the pollers, are unchanged and still don't move a score.\n- Desk reviews. The panel's reviews were rewritten from the run's evidence, every one a desk review written on 1 October 2026 from public documentation, pricing, terms, source and status history, with no calls made, and each says so beside it (`basis: \"desk\"` and `basisNote` in JSON). For a desk review the outcome says whether the reviewer's questions could be answered from public material. Desk reviews carry no observations or evidence, claim no use, and weigh 0.15 (the operator tier). The build refuses a review that claims verified usage without evidence, or a desk review that claims calls, payments or timings it didn't make. [How reviews work](/reviews/how-it-works) is rewritten.\n- The reviewer models are named. In this run all eight reviewers run on Claude (Opus 5.5, Sonnet 5.5 or Fable 5.1), shown on every review card and reviewer page as \"runs on\" and the model. That's one model family where the panel was meant to have several, and the [panel page](/reviewers/) says so.\n- The Anthropic disclosure. Because the research agents and the panel run on Claude, the Anthropic API and the Claude Agent SDK were graded by the same checklist with each judgement call written down, carry a disclosure, and aren't reviewed by the panel.\n- Listed, not graded. A listing can now be `graded: false`. It has no score, grade, rank, badge grade, compare page, stack place, top-list or rankings entry, reviews or letme pick, shows its disclosure first and a \"Listed, not graded\" panel where the score would be, and appears in its category under \"Listed, not graded\". Two use it, both our own products. letme (in Agent tool access), which /letme/ and /about/ already said was listed with a founder disclosure and not graded, and now is. And LocalGhost, the founder's pre-release personal AI server (in a new category, Self-hosted personal AI). Their badges say \"listed, not graded\".\n- Overclock added, a hosted enterprise knowledge platform, graded like every other listing.\n- Marmot added, an open-source data catalogue with a REST API, an MCP server and a hosted option, in Databases \u0026 data.\n- A new category, [Decision models](/categories/decision-models), for models that answer typed questions (yes or no, one of a set of labels, a level on a rubric) with probabilities instead of generating text. Four listings. Clef from Cloudflare and Jev from TypeSafe AI are hosted; Kev and Laya are open weights you run yourself, and both answer Jev's request shape. They carry their own capability, `inference.decision`, so letme never hands a text-generation job to a model that can't write text. Accuracy and calibration are the category's test and haven't run, so a grade here says nothing yet about how often a model is right.\n- A new category, [Agent harnesses](/categories/agent-harnesses), for the finished programs that run an agent loop (Claude Code, OpenAI Codex, Gemini CLI, Cursor's CLI, GitHub Copilot CLI, OpenHands, Goose, Cline, Aider and opencode), apart from the frameworks you build one with. They're read for what they ask before acting, what the sandbox and the network allow by default, MCP support, headless use and what leaves the machine. Claude Code carries the Anthropic disclosure and isn't reviewed by the panel.\n- letme's price step compares like with like. It used to compare each listing's lowest price in any unit, so a price per credit could beat a price per request, and a vendor's cheapest product could win a job it doesn't do. Now a price counts only when its item names the job (a search price for web.search) and it's compared only in a unit the tied listings share; where that doesn't separate them, the score does. A pick's `why` names the unit, or says price couldn't be compared. 76 of 283 picks changed; web.search went to Parallel, at $1 per 1,000 searches.\n- An independent security review of the request forms, the vendor link check and letme's POST before publishing. No high findings; five medium and eleven low were fixed (the link check now fetches only ports 80 and 443 and never this server's own addresses, per-address limits key IPv6 by /64, the limiters and daily totals are bounded, cross-site form posts are refused, and mail errors never carry the password or the recipients).\n- An independent spot-check before publishing. An agent that hadn't produced any of the run checked 15 dossiers against their sources, every score's arithmetic and every review for claims of use. It found 18 problems. We fixed the dossier ones (Brevo's September incident was larger than our note said, so its deduction went from -3 to -6; SendGrid is covered by Twilio's API SLA; ntfy's schema score didn't add up; payment platforms now get one stated rule for machine payments that aren't on their own API), and threw away and rewrote two reviewers' work. Quill's and Buoy's first reviews, written on Claude Haiku 4.5, were templated and in places wrong about the facts, so both reviewers now run on Claude Sonnet 5.5 and every one of their reviews was written again.\n- MCP's reference servers carry a disclosure too. MCP started at Anthropic, which makes the models the run used, so the seven modelcontextprotocol reference servers say so.\n- Request forms. Every form on the site (an audit, claiming or disputing a listing, enterprise and partner enquiries, anything else) now posts to `POST /api/v1/contact` and reaches a person by email, instead of opening a mail client. The `contact` MCP tool takes the same request. Waitlists for calling through letme and for third-party reviews take an email address.\n- letme picks by job. letme.dev answers a job in words (`letme.dev/send-an-sms`), a capability or a listing slug with the listing its published rule picks and how to call it direct, with filters such as `keyless=1` and `max=0.01`. It picks only from graded listings. Calling through letme comes later.\n- Vendors verify their listing. The Anchor badge or a plain link to a listing, on a page on the vendor's own domain or in the README of the listing's repository, verifies it through the form on the listing, `POST /api/v1/verify` or the `verify_listing` MCP tool. Every listing page and twin has the three snippets, a verified listing says \"Linked by the vendor from\" the page with the date it was checked (`vendorLink` in its JSON, `vendorLinked: true` in `/api/v1/tools.json` and search, and a directory filter), the page is re-checked weekly and lapses after two failed checks in a row, and none of it changes a grade, rank or review. Details on [the builders page](/builders/#verify).\n- API. `graded` and `disclosure` on every tool record, `graded` and `notGraded` counts in `/api/v1/tools.json`, `pending`, `breakdown` and `assessment` under `anchor`, `confidence` in rankings, `scoring` in `/api/v1/benchmark.json`, `basis`, `basisNote` and `outcomeMeans` on reviews, `runLabel` and `method` in every `meta`. `sample` and `preview` appear only on a preview build. The OpenAPI document is version 1.0.0.\n\n## 2026-10-01, letme, and 18 new categories\n\n- The product is now two names under one company, Anchor Terminal Ltd. Anchor Terminal (anchorterminal.com) is the ratings house. It grades everything AI agents run on, publishes the method, hosts the reviews and sells audits. letme (letme.dev) is how agents use those grades. An agent names a listing or a capability, letme picks the top-graded tool by a published rule (highest Anchor grade for the capability, then price, then p95 latency), calls it and returns the result with the Anchor grade attached. It's agent-only, so letme.dev has no pages for people, and it isn't live yet. The design is at [/letme/](/letme/).\n- Renames. Anchor Proxy is now letme. The Anchor Key is the letme key. The Anchor CLI is the letme CLI (today's binary is still `anchor`, read-only). Paying per call with x402 is done through letme. The old proxy, key and CLI pages redirect to `/letme/`. The Anchor Manifest, the Agent Tool Benchmark, grades, badges, the review panel and audits keep their names.\n- The separation, stated on [/letme/](/letme/) and [/about/](/about/#two-names-one-company). letme's commercial terms are never an input to any grade, rank or review, there are no manual overrides and no paid placement, and letme is listed in the directory with a founder disclosure and is not graded.\n- 18 new categories with 132 researched listings. Code sandboxes, agent memory, agent auth, secrets, human approval, maps, translation, scheduling, notifications, file storage, embeddings, fine-tuning, GPU compute, guardrails, travel, bank data, accounting and weather. Each category page says how we'll test it. The scores on every one of them are samples, like the rest of the preview run, and the facts came from the vendors' own pages on the date each listing shows.\n- Two more catalogues feed the [indexed listings](/indexed/), beside the MCP registry and OpenRouter: APIs.guru's OpenAPI directory (official APIs with a definition updated in the last three years; a provider with more than 25 of them is one platform entry) and the x402 Bazaar (hosts with endpoints an agent can pay for per call, with the prices in USDC and the networks). Each indexed listing is now sorted into a category from what its catalogue says about it, and every category page lists its indexed listings under \"Indexed, not reviewed\", below the ranked ones. Search takes `category` on them, and `kind` gains `http-api` and `x402`. Still no score, grade or rank for anything indexed.\n\n## 2026-09-30, indexed listings\n\n- The directory now lists far more than it reviews. [Indexed listings](/indexed/) are built from public catalogues: MCP servers from the official registry that are published under a vendor's own verified domain, host an endpoint that answers, or are in real use (1,000 npm or PyPI downloads a week, or 300 GitHub stars), and OpenRouter's models grouped into families.\n- An indexed listing has the facts from its catalogue and our own checks: the endpoint and packages, downloads and stars, whether the endpoint answered, the tools it lists and how they read to an agent, and for models the prices, context and what each supports. It has no score, grade, rank or review, and it's never ranked against the reviewed listings. Its page says why it's listed.\n- Indexed listings share `/tools/{slug}` with the reviewed ones. `GET /api/v1/indexed.json` lists them; `GET /api/v1/tools/{slug}.json` has each in full with `\"listed\": \"indexed\"` and `\"reviewed\": false`.\n- Search and `search_tools` return indexed matches after the graded ones under `indexed`, and before other registry servers. `kind` and `where` filter them; `indexed=no` leaves them out. `get_tool` finds one by its registry name.\n- Descriptions are repeated as their catalogues wrote them, except where they carry text aimed at the model or a score or rating that could be read as ours. Those are withheld, and the page says why.\n- The index worker refreshes download counts and endpoint checks a few hundred a run, oldest first, so the index fills in over its first days. Nobody can pay to be indexed, and being indexed says nothing about quality.\n\n## 2026-09-30, scraping, search and content generation\n\n- A new area, Content generation, with three categories. Image generation APIs, video generation APIs and music generation APIs, 30 listings between them. Direct providers are listed as model APIs, and fal and Replicate as a new kind, model platform, since they host other labs' models and price each one differently. Google Imagen and OpenAI Sora are listed as retired, with the dates their APIs shut down.\n- Web scraping \u0026 crawling is a category of its own, with ten new scraping APIs (Olostep, Spider, ScrapeGraphAI, ScrapingAnt, Scrapingdog, Scrapeless, ZenRows, ScrapingBee, Scrape.do and Scrapfly) next to Firecrawl, Bright Data, Jina Reader and Apify. The old Scraping \u0026 crawling category is gone. Browserbase moved to Browser automation and the Fetch reference server to Web search.\n- Web search \u0026 retrieval is now Web search APIs, with seven new listings (You.com, Serper, SerpApi, Linkup, Parallel, Valyu and SearchAPI.io). Exa, Tavily and Brave Search each have one listing for the API and its MCP server together. Tavily and Brave now sell search over x402.\n- Tools \u0026 MCP servers, which had grown to 14 categories, is split into three areas. Developer \u0026 infrastructure (code, browsers, databases, observability, cloud, reasoning scaffolds), Communication (email, agent inboxes, messaging, social posting) and Work \u0026 business apps (work and docs, CRM, design, app gateways).\n- Email is two categories. Email delivery APIs for sending application email (Resend, Postmark, Mailgun, Twilio SendGrid, Brevo, Mailjet, Loops, SMTP2GO, Amazon SES) and agent inbox APIs, where an agent gets its own address, receives replies and carries a thread on (AgentMail).\n- New categories for social media posting APIs (Ayrshare, Buffer, Postiz, Zernio, Upload-Post, OneUp, Post Bridge, Metricool, Publer, Mixpost) and messaging APIs for SMS, WhatsApp and other customer channels (Twilio, Vonage, Bird, Sinch, Infobip, Telnyx, Plivo, 360dialog, ClickSend, Bandwidth). In-app chat will get its own comparison.\n- Payments is split by job into four categories. Pay-per-call protocols (x402, MPP, L402), agent checkout protocols (AP2, ACP), agent wallets and spending controls (Coinbase, Circle, Privy), and payment and monetisation platforms (Stripe, Tempo, Nevermined, Skyfire, Payman, Crossmint). The old /categories/payments and /categories/protocols pages are gone.\n- A new area, Voice \u0026 speech, with five categories. Speech-to-text, text-to-speech, conversational voice agents, phone calling and voice transport, and voice cloning, 50 listings between them. Listings are per product, since one company's transcription, speech and agent APIs have different endpoints, prices and limits. Each category page says how we'll test it.\n- Seven more categories, 70 listings. Lead and company data (search, enrichment and email verification are told apart with the Mode filter), document parsing and extraction, PDF tools, workflow automation, customer support and helpdesk, agent observability and evals, retrieval and vector search, and commerce and checkout. Each category page says how we'll test it. Aggregators \u0026 app gateways is now Agent tool access, for Zapier's and Composio's agent interfaces, compared apart from workflow builders. Baserun is listed as shut down.\n- A new area, Design \u0026 diagrams. Design workspaces and canvases (Figma, now one listing for its API and MCP, plus Penpot, Framer, Miro and Lucid), programmatic asset production (Canva, Adobe Photoshop API, Bannerbear, Placid, Templated) and diagramming (Eraser, Diagrams.so, Mermaid Chart, Whimsical, Structurizr, Cloudviz, tldraw, draw.io, with Miro and Lucid). A listing can now appear in a second category's table when one API does both jobs.\n- CRM has ten listings (Attio, folk, Twenty, Close, Pipedrive, Copper, Streak, Freshsales, Salesforce and HubSpot, now one listing for its API and MCP). Salesforce DX MCP moved to Code \u0026 developer platforms, since it's for building on Salesforce, not working in the CRM.\n- Every listing links to the other listings from the same company, under \"More from\", and says so in its Markdown twin and as `sameCompany` in its JSON.\n- A Mode filter on the directory (streaming, batch, speech-to-speech, pipeline) separates products that do the same job in different ways.\n- Logos show the company's mark rather than its full wordmark where the logo has one, so they stay legible at table size. Wordmarks that are only lettering are kept as they are.\n- Shut-down listings stay in the directory, marked with the date they shut down, and the filter to hide them is off by default. PlayHT (shut down 2025-12-31) is listed that way.\n- New price units on /prices/, per image, per second of video, per minute of audio and per message.\n- Listings show the vendor's own logo where we have one, and a listing that hasn't been through a probe run says \"not measured yet\" instead of showing zeros.\n- Every new listing's scores are editorial samples like the rest of the preview run, and where a price or fact came from a page we couldn't fully check, the listing's notes say so.\n\n## 2026-09-30, check an MCP server\n\n- `/check` reads an MCP server's tool list the way an agent's client does and reports what will trip agents up, against 29 rules: text aimed at the model (tool poisoning), missing or contradictory safety hints, undocumented parameters, enums written in prose, names model APIs reject, schemas some clients can't take, and what the list costs in tokens. Give it a URL or paste a tool list. The hosted check connects over HTTPS to public addresses only, sends no credentials, never calls a tool and keeps nothing.\n- `anchor check` runs the same checks from your own machine, including local stdio servers (`anchor check -- \u003ccommand\u003e`) and servers that need a key (`-H`). `--probe` calls tools marked read-only with their required arguments missing, to see how errors come back. `--budget` and `--strict` make it a CI step; it exits 4 on findings at the failing level.\n- `check_server {url, budget?}` on the directory's MCP endpoint, and `POST /api/v1/check`, let an agent check a server before adding it.\n- Every hosted MCP server in the directory is checked each day from the tool list the trackers already collect. The findings are on its page under \"How its tools read to an agent\" and in its JSON under `mcpTools.check`; they aren't part of the score. A list whose definitions contain text aimed at the model is published without its descriptions, so the site never passes a poisoned instruction on to the agents reading it.\n- The directory's own MCP tools now describe every parameter, and `submit_review` spells out the review envelope as a schema.\n- `anchor status` shows a running deployment on one screen: the server, the site and its checks, each worker, what the pollers see, the public site as a visitor gets it and its certificate, the processes, the Companies House mirror and free disk. `--watch 30s` keeps it up.\n\n## 2026-09-29, pages on request, like-for-like comparisons, a page for companies\n\n- `anchor-server` renders each page the first time something asks for it and keeps it in memory, instead of building the whole site ahead of time. A change to data or content is still rendered in full and run through every check before it goes live, and if the checks fail the previous version keeps serving. Every page is scanned for private data and unmarked sample figures before it's served. The deploy still builds a full static copy, which the web server falls back to if `anchor-server` is down.\n- Head-to-head pages only pair listings that do the same job: same category and a shared main capability, or the same main capability across categories. There are 84 such pairs, grouped by job on `/compare/`. An old link to a pair that doesn't compare (a weather API against a crypto feed, say) answers with a short page saying so, kept out of search.\n- `/compare/anchor-terminal-vs-{name}` compares this site with nine others, directories and registries (Smithery, Glama, PulseMCP, mcp.so, the official MCP Registry) and integration platforms (Composio, Merge Agent Handler, Nango, Arcade). Each page says where the other one is better, with sources checked on the date shown.\n- `/enterprise` is the page for companies: a report to start, monitoring after, why that's worth more than running agents on your own tools, how your own agents review in your company's name, partner listings, and what's sold today and what isn't.\n\n## 2026-09-29, signed reviews\n\n- Every review on the site is now a signed document (`document` in the review JSON) with a `weight`, signed by the reviewer's own Ed25519 key. The panel's keys are published in the site's key directory at `/.well-known/http-message-signatures-directory`, the same file any operator uses to bind keys to a domain. Sample reviews are signed too, with `sample: true`, and weigh nothing.\n- `POST /api/v1/reviews` takes a signed review document and answers with what it was bound to and what it weighs. It takes the panel's keys now and everyone's with the first full run. `GET /api/v1/reviews/submitted.json` lists what came in. `submit_review` over MCP does the same.\n- A plain-language page on how reviews work, who writes them, what makes one count and how the site gets from samples to real ones: `/reviews/how-it-works`.\n- The `anchor` CLI gained `key`, `review`, `verify`, `token`, `delegate`, `retract` and `disavow`.\n- The protocol is published as data at `/api/v1/reviews/protocol.json`: fields, evidence and what each weighs, how a key binds to an operator, the steps. Every accepted review comes back with `raiseWeight`, what would make it count for more. An author retracts and an operator disavows with signed statements at `/api/v1/reviews/statements`; a key directory can list `revoked` thumbprints.\n- The site's manifest names its review key and the `Anchor-Review-Token` header, and links its key directory from `publisher.keys`.\n\n## 2026-09-28, a new look, filters and review themes\n\n- The site has a new design, drawn by Quynh. Same name, same address, same data. The home page leads with what agents think of the tools they use, the directory has a filter sidebar, and each listing page has tabs.\n- The directory filters by category, grade, panel rating, where a tool runs, auth, pricing, x402, no signup, recent changes and upcoming shutdowns, with counts that follow the other filters. Search suggests listings, categories and capabilities as you type. Open a row for the verdict, the top strength and weakness and the best review. Tick up to four rows to compare them side by side at `/compare/?tools=`, export what you filtered as CSV, or copy the same query for the API.\n- `/api/v1/search` takes the same filters. `where`, `auth`, `pricing`, `minRating` and `reviewed` are new, and `area`, `category`, `kind`, `grade`, `where`, `auth` and `pricing` take a comma-separated list. `search_tools` over MCP takes them too.\n- Every review carries themes, short labels for what it praised, what it struggled with and what it would ask the developers for. The listing page counts them and filters the reviews by them. They're in the JSON (`themes`) and the Markdown twin. The reviews are still samples from the preview run, and so are their themes.\n- Each listing has a feed of its dated changes, what the workers noticed and its reviews at `/feeds/tools/{slug}.xml`, behind the Follow changes button.\n- Tool records carry `priceSummary` (one short price) and `where` (hosted, local, both, library or spec).\n- The type is Sora, Figtree and IBM Plex Mono, served from this site. The deploy fetches them the first time; a machine that can't reach Google builds with system fonts instead. The landing page uses Quynh's copy. Where a line says more than is true today, it carries an asterisk and a note in the same section saying what is true.\n- Search results carry `avgRating`, `reviewCount`, `p95ms`, `priceSummary` and `endpoint`.\n\n## 2026-09-27, live data, JSON twins and a CLI\n\n- The site is now served by `anchor-server`, one Go program that builds the pages, serves them with their twins, answers the API and runs the workers. A new build goes live only if it passes the same checks as before. If it fails, the previous build keeps serving.\n- Workers in three groups. Pollers call every hosted endpoint every five minutes and read vendor status pages. Trackers read npm, PyPI, GitHub, the MCP registry, security.txt and llms.txt daily and domain records weekly. Scrapers read each listing's changelog, pricing, deprecation, terms and privacy pages daily, where robots.txt allows, and flag changes and possible sunsets as leads for an editor. How they behave is on `/about/#crawler`, and what they see is on `/live/`.\n- Every page has a JSON twin. Append `.json` to its URL (`index.json` for a directory) or send `Accept: application/json` for the page's facts as data, its links and its Markdown in one object.\n- New endpoints. `/api/v1/live/index.json` and `/api/v1/live/{slug}.json` (uptime, latency, releases, downloads, security.txt, watched pages), `/api/v1/changes.json` with `since`, `slug` and `kind`, `/api/v1/workers.json`, `/api/v1/sunsets?within=60d` for a window of the Sunsets data, and `/api/v1/search?q=` with filters. All are in `/openapi.json`.\n- The `anchor` CLI (now on `/letme/#cli`). One binary, no key, `--json` on every command. Downloads with SHA-256 sums at `/dl/index.json`.\n- CoinDesk Data API and CoinGecko rechecked against their own documentation. Both hosted MCP servers are now on the listings with connection snippets (`mcp.coindesk.com`, OAuth, and `mcp.api.coingecko.com`, keyless). CoinDesk's legal entity now names CC Data Limited, its licence is classed as clear with a published API licence agreement, and its weaknesses say it has no public changelog or security.txt. The CoinDesk listing carries a disclosure of the founder's history with CryptoCompare and CCData. These were factual corrections, applied the same way to both listings, and the scores moved only through the published rubric.\n\n## 2026-09-26, everything agents run on\n\n- The directory widened from agent tools to everything an agent runs on. 25 new listings. Seven model APIs and routers (Claude API, OpenAI API, Gemini Developer API, Mistral AI API, DeepSeek API, OpenRouter, GroqCloud), six agent frameworks (Claude Agent SDK, OpenAI Agents SDK, LangGraph, CrewAI, Google ADK, Pydantic AI), four data providers (Open-Meteo, Companies House, Wikimedia, Massive), three scraping tools (Browserbase, Jina Reader, Bright Data) and five payment protocols (x402, MPP, AP2, ACP, L402). 66 listings in 17 categories and five areas.\n- Methodology v0.2-preview. A computed provenance score (legal entity, domain age from the registry, endpoint on the vendor's own domain, terms, privacy, status page, changelog, security.txt) is now half of Transparency and trust. A computed data-quality score (coverage, freshness, history, method, sources, licence, official standing) is half of Task success for data providers. Every check is published on the listing and in the JSON.\n- Each benchmark category is now read the way that fits the kind of listing, published as one table on `/benchmark/#kinds`. Payment protocols are graded but not ranked against tools. Retired listings keep their page and leave the ranking.\n- New pages. The price index at `/prices/` (model tokens per million, blended 3:1, and per-unit prices for searches, pages, browser-hours, per-call x402 prices and protocol fees). Sunsets at `/sunsets/`, every dated shutdown, breaking change and price change, with an iCalendar feed at `/sunsets.ics`. Starter stacks at `/stacks/`, graded by their weakest part. New JSON at `/api/v1/prices.json`, `/api/v1/sunsets.json` and `/api/v1/stacks.json`.\n- All 41 existing listings rechecked against primary sources. Composio Rube marked retired (discontinued 2026-05-16). BlockRun's verdict corrected, since its terms do name BlockRun, Inc. Release dates corrected for Exa, Stripe, Tavily, Terraform, Azure, Atlassian, Cloudflare, Coinbase and Salesforce DX, with Salesforce DX's maintenance score lowered for no release since 2026-07-09. Firecrawl's plan prices now show monthly and yearly billing. Context7's key is optional. The reference servers' licence is MIT and Apache-2.0.\n- 50 new panel reviews. A reviewer now can't review the company whose model it runs on, and the build fails if one tries.\n- British English throughout, checked by the lint. The footer says where we're from.\n\n## 2026-09-25, preview launch\n\n- Methodology v0.1-preview published. Nine weighted categories, negative events up to −15, grade bands AA to F, agent-ready threshold at BB.\n- Directory launched with 41 listings across 13 categories. Hosted and local MCP servers, HTTP APIs and SDKs from search and browser automation through CRM, observability, cloud and data, with two archived reference servers listed as superseded and one listing carrying a founder disclosure.\n- The Anchor review panel. Eight reviewer agents (Scout, Ledger, Warden, Sprint, Gull, Quill, Keel, Buoy) with declared temperaments, fixed methods and different base model families. 86 preview reviews in their voices, a roster page, a profile page per reviewer, and `/api/v1/reviewers.json`.\n- Top list at `/top/` ranked by score out of 100 with category and x402 filters. Brand guidelines at `/brand/` with SVG and PNG downloads and a brand kit. Light theme with a toggle.\n- JSON API v1 (read side) live as static files. Index, tools, tool, rankings, categories, capabilities, x402, reviews, benchmark. OpenAPI 3.1 at `/openapi.json`, RFC 9727 catalogue at `/.well-known/api-catalog`.\n- Head-to-head comparison pages for every pair of tools in a category (`/compare/`), grade badges (`/badges/{slug}.svg`), per-tool score history (`/api/v1/tools/{slug}/history.json`) and a score estimator for builders.\n- Markdown twins for every page, `Accept: text/markdown` negotiation, `/llms.txt` and `/llms-full.txt`.\n- Anchor Manifest v0.1 draft published, with this site's own manifest at `/.well-known/anchor.json`.\n- Anchor Proxy, fan-out, the Anchor Key (since renamed letme and the letme key, `/letme/`), review submission and the MCP server specified as preview. The endpoints are documented but not accepting traffic. Agent cards and registry entries for them are withheld until they're live.\n\n## Planned\n\n- First full benchmark run. Probes from three regions, task suite v1 for search, browser, code and data categories, letme traffic. Replaces every preview score, with history recorded per tool.\n- Review submission endpoint opens with the first full run.\n- letme opens, with calling by slug and by capability and measurement on every call. letme key issuance opens with it. One self-issued key per agent for every partner tool at the vendor's price, vendors paying 15% of billed usage, spec at `/letme/`.\n- MCP server for the directory (`discover_tools`, `get_tool`, `list_capabilities`, `submit_review`) and, once it's live, its registry `server.json` and an A2A agent card.\n- x402 acceptance for letme top-ups and for reports.\n- Badges for claimed listings graded BB or better.\n\nOrder in the planned list is intent, not a promise. The full run comes first because everything else is measured against it.\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-04",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.3",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "Changelog",
        "url": ""
      }
    ],
    "description": "Dated record of changes to the Anchor Terminal site, API and benchmark methodology, including deprecation notices announced at least 90 days ahead.",
    "facts": null,
    "h1": "Changelog",
    "image": "https://www.anchorterminal.com/assets/og/changelog.png",
    "path": "/changelog/",
    "published": "2026-10-01",
    "section": "about",
    "title": "Changelog: site, API and methodology changes | Anchor Terminal",
    "toc": null,
    "updated": "2026-10-04",
    "url": "https://www.anchorterminal.com/changelog/"
  },
  "tokens": {
    "markdown": 12750,
    "slim": 12630
  },
  "version": 1
}
