{
  "data": {
    "faq": null,
    "kicker": "Sample · fictional tool · full structure",
    "lede": "What a readiness report looks like end to end, for a fictional web-search MCP server called Harbourline Search. Every number here is invented to show the format, and the grade in it is fictional too. Real reports will come from the published checklist, our probes and task suite run against your tools, and a line-by-line read of the tool definitions. We haven't delivered one for a customer yet."
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/builders/sample-report",
    "json": "https://www.anchorterminal.com/builders/sample-report.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/builders/sample-report.md",
    "slim": "https://www.anchorterminal.com/builders/sample-report.min.md"
  },
  "markdown": "## Summary\n\n|  |  |\n| --- | --- |\n| Tool | Harbourline Search MCP (fictional) |\n| Assessed | 2026-10-01, methodology v0.3 |\n| Current grade | B (63.4), rank 5 of 8 in web search |\n| Projected after top five fixes | A (79–81) |\n| Biggest single lever | Context cost. 9,400 tokens for `tools/list`, and four tools would do the work of eleven |\n| Biggest risk | `search_raw` returns unfiltered page HTML including any instructions found on the page |\n\n## 1. Where agents fail (task suite)\n\nTwelve web-search tasks, four attempts each, with the reference model and harness named in the run notes.\n\n| Task | pass@1 | pass@4 | Median turns | Where it went wrong |\n| --- | --- | --- | --- | --- |\n| Find the publication date of a named report | 4/4 | 4/4 | 3 | (none) |\n| Find three primary sources for a claim | 2/4 | 4/4 | 7 | Model called `search_news` for an academic query; description of `search_news` doesn't say it's limited to the last 30 days |\n| Extract a table from a results page | 1/4 | 2/4 | 9 | `fetch_page` returns HTML; model ran out of context twice at 40k tokens |\n| Compare prices across five vendor pages | 0/4 | 1/4 | 11 | Rate limit at request 12 returned `500 Internal Server Error` with body `\"limit\"`; model retried immediately, four times |\n| Answer a question requiring a date filter | 3/4 | 4/4 | 4 | `date_from` accepts `YYYY-MM-DD` but the description says \"ISO date\"; one attempt sent a full timestamp and got an empty result, not an error |\n\nTranscripts for every failed attempt are in the attached JSON with the turn index and the model's visible context at the point of failure.\n\n## 2. What your schema costs\n\n`tools/list` in the default configuration comes to **9,412 tokens** (reference tokeniser). The full surface with `--enable-experimental` is 14,880.\n\n| Tool | Description chars | Schema tokens | Note |\n| --- | --- | --- | --- |\n| `search` | 1,842 | 1,310 | Lists every parameter twice, once in prose and once in the schema |\n| `search_news` | 1,790 | 1,265 | 90% identical to `search` |\n| `search_images` | 1,650 | 1,180 | Rarely called; could be a parameter of `search` |\n| `fetch_page` | 920 | 740 | Does not mention output size or that HTML is returned |\n| `fetch_page_markdown` | 940 | 750 | The tool agents should use; described last |\n| (six more) |  | 4,167 |  |\n\nThe rewrite we'd propose is three tools. `search` (with `type` and `date_from`), `fetch` (markdown by default, `raw: true` for HTML) and `crawl`. Projected 2,100 tokens, no loss of capability. Draft descriptions are in the appendix.\n\n## 3. Descriptions that mislead\n\n- `search` says \"Powerful semantic search over the entire web with advanced ranking\", which says nothing about when to use it instead of `search_news`. Models pick by name, because the name is the only signal they have.\n- `search_news` doesn't state the 30-day window. Two of the twelve task failures trace to that omission.\n- `fetch_page` doesn't say it returns HTML, or how large the return can be. Models call it, receive 30k tokens, and lose the thread.\n- `date_from` says \"ISO date\", which models read as a timestamp. The server accepts only `YYYY-MM-DD` and returns an empty result on anything else. Say the format, or better, accept both.\n\n## 4. Errors agents can't recover from\n\n| Error observed | Frequency | Recovered without a human? | Should say |\n| --- | --- | --- | --- |\n| `500 {\"error\":\"limit\"}` on rate limit | 6.1% of calls in bursts | No: models retry immediately and make it worse | `429`, `Retry-After: 2`, body `{\"error\":\"rate_limited\",\"retryAfterMs\":2000,\"limit\":\"10/s\"}` |\n| Empty result on bad date format | 1.4% | Rarely | `400 {\"error\":\"invalid_date_from\",\"expected\":\"YYYY-MM-DD\",\"got\":\"2026-09-01T00:00:00Z\"}` |\n| `504` from upstream index | 0.3% | Yes, after a retry | Fine; add `Retry-After` |\n\n## 5. Onboarding friction\n\nTime to first successful call for an agent with no prior account is **not achievable without a human**. Three steps need a person. Create an account (email confirmation), add a card (required even for the free tier), copy a key. With x402 on `/search` and `/fetch`, time to first call for a funded agent is one request. With programmatic key issuance behind a card-free tier, it's two.\n\n## 6. Reliability and latency\n\n30 days, three regions, five-minute probes, plus 41,000 proxied calls.\n\n| Region | Availability | p50 | p95 | Notes |\n| --- | --- | --- | --- | --- |\n| London | 99.91% | 610 ms | 2.4 s | Two 12-minute outages, both 03:00–04:00 UTC |\n| Virginia | 99.97% | 380 ms | 1.6 s |  |\n| Singapore | 99.62% | 1,150 ms | 4.9 s | No regional endpoint; every call crosses the Pacific |\n\nYour status page showed 100% for the month. The London outages line up with a nightly index rebuild. A status entry, or a `503` with `Retry-After`, would have turned a negative event into a documented maintenance window.\n\n## 7. Security posture (external)\n\n- API key accepted in the query string (`?key=`) as well as the header. Query-string keys end up in logs and referrers. Deprecate with a date.\n- No read-only mode is needed (all tools are read-only), but the tools carry no `readOnlyHint` annotation, so hosts can't tell.\n- `fetch_page` returns page content verbatim. Two of our probe pages contained instruction-shaped text and both came back unmarked. Wrap fetched content with a boundary marker and a warning, and document it.\n- No `/.well-known/security.txt`.\n\n## 8. Prioritised fixes\n\n| # | Fix | Effort | Expected movement |\n| --- | --- | --- | --- |\n| 1 | Return `429` + `Retry-After` for rate limits; structured `400` for bad parameters | Small | Ergonomics +18, Task success +12 → +4.2 total |\n| 2 | Collapse eleven tools into three; rewrite descriptions (drafts attached) | Medium | Schema +22, Ergonomics +10 → +4.2 total |\n| 3 | Markdown by default from `fetch`; document size and add `max_length` | Small | Ergonomics +8, Task success +8 → +1.8 total |\n| 4 | x402 on `/search` and `/fetch` at your published prices | Small (Go: ~40 lines) | Payments +40 → +4.0 total |\n| 5 | Card-free tier or programmatic key issuance | Medium | Payments +20 → +2.0 total |\n| 6 | Deprecate query-string keys; add `readOnlyHint`; wrap fetched content | Small | Security +14 → +2.0 total |\n| 7 | Status page entries or `503`s for the nightly rebuild; Singapore endpoint or CDN | Medium | Reliability +6, Performance +10 → +2.0 total |\n\nProjected after fixes 1 to 5, 79 to 81 (A). After all seven, 82 to 84 (A).\n\n## Appendix\n\n- `report.json`, every table above as data, plus transcripts and probe logs.\n- `descriptions.md`, the proposed tool definitions.\n- A re-run, meaning a reduced-price re-assessment after the changes ship, with a before/after diff.\n\n*Harbourline Search is fictional. Any resemblance to a real search server with eleven tools and a 500 on rate limits is a coincidence the real server should look into.*\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-04",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.3",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "For builders",
        "url": "https://www.anchorterminal.com/builders/"
      },
      {
        "name": "Sample report",
        "url": ""
      }
    ],
    "description": "A complete sample agent-readiness report for a fictional MCP server (Harbourline Search), showing task-suite failures, schema token costs, misleading descriptions, unrecoverable errors, onboarding friction, reliability by region, security posture and a prioritised fix list with expected score movement.",
    "facts": [
      "fictional tool",
      "8 sections",
      "prioritised fixes"
    ],
    "h1": "Sample readiness report",
    "image": "https://www.anchorterminal.com/assets/og/sample-report.png",
    "path": "/builders/sample-report",
    "published": "2026-10-01",
    "section": "builders",
    "title": "Sample agent-readiness report for tool builders | Anchor Terminal",
    "toc": [
      {
        "id": "summary",
        "level": "h2",
        "text": "Summary"
      },
      {
        "id": "1-where-agents-fail-task-suite",
        "level": "h2",
        "text": "1. Where agents fail (task suite)"
      },
      {
        "id": "2-what-your-schema-costs",
        "level": "h2",
        "text": "2. What your schema costs"
      },
      {
        "id": "3-descriptions-that-mislead",
        "level": "h2",
        "text": "3. Descriptions that mislead"
      },
      {
        "id": "4-errors-agents-can-t-recover-from",
        "level": "h2",
        "text": "4. Errors agents can't recover from"
      },
      {
        "id": "5-onboarding-friction",
        "level": "h2",
        "text": "5. Onboarding friction"
      },
      {
        "id": "6-reliability-and-latency",
        "level": "h2",
        "text": "6. Reliability and latency"
      },
      {
        "id": "7-security-posture-external",
        "level": "h2",
        "text": "7. Security posture (external)"
      },
      {
        "id": "8-prioritised-fixes",
        "level": "h2",
        "text": "8. Prioritised fixes"
      },
      {
        "id": "appendix",
        "level": "h2",
        "text": "Appendix"
      }
    ],
    "updated": "2026-10-04",
    "url": "https://www.anchorterminal.com/builders/sample-report"
  },
  "tokens": {
    "markdown": 1900,
    "slim": 1580
  },
  "version": 1
}
