Sample · fictional tool · full structure
Sample readiness report
What a readiness report looks like end to end, for a fictional web-search MCP server called Harbourline Search. Every number here is invented to show the format, and the grade in it is fictional too. Real reports will come from the published checklist, our probes and task suite run against your tools, and a line-by-line read of the tool definitions. We haven't delivered one for a customer yet.
Summary
| Tool | Harbourline Search MCP (fictional) |
| Assessed | 2026-10-01, methodology v0.3 |
| Current grade | B (63.4), rank 5 of 8 in web search |
| Projected after top five fixes | A (79–81) |
| Biggest single lever | Context cost. 9,400 tokens for tools/list, and four tools would do the work of eleven |
| Biggest risk | search_raw returns unfiltered page HTML including any instructions found on the page |
1. Where agents fail (task suite)
Twelve web-search tasks, four attempts each, with the reference model and harness named in the run notes.
| Task | pass@1 | pass@4 | Median turns | Where it went wrong |
|---|---|---|---|---|
| Find the publication date of a named report | 4/4 | 4/4 | 3 | (none) |
| Find three primary sources for a claim | 2/4 | 4/4 | 7 | Model called search_news for an academic query; description of search_news doesn't say it's limited to the last 30 days |
| Extract a table from a results page | 1/4 | 2/4 | 9 | fetch_page returns HTML; model ran out of context twice at 40k tokens |
| Compare prices across five vendor pages | 0/4 | 1/4 | 11 | Rate limit at request 12 returned 500 Internal Server Error with body "limit"; model retried immediately, four times |
| Answer a question requiring a date filter | 3/4 | 4/4 | 4 | date_from accepts YYYY-MM-DD but the description says "ISO date"; one attempt sent a full timestamp and got an empty result, not an error |
Transcripts for every failed attempt are in the attached JSON with the turn index and the model's visible context at the point of failure.
2. What your schema costs
tools/list in the default configuration comes to 9,412 tokens (reference tokeniser). The full surface with --enable-experimental is 14,880.
| Tool | Description chars | Schema tokens | Note |
|---|---|---|---|
search | 1,842 | 1,310 | Lists every parameter twice, once in prose and once in the schema |
search_news | 1,790 | 1,265 | 90% identical to search |
search_images | 1,650 | 1,180 | Rarely called; could be a parameter of search |
fetch_page | 920 | 740 | Does not mention output size or that HTML is returned |
fetch_page_markdown | 940 | 750 | The tool agents should use; described last |
| (six more) | 4,167 |
The rewrite we'd propose is three tools. search (with type and date_from), fetch (markdown by default, raw: true for HTML) and crawl. Projected 2,100 tokens, no loss of capability. Draft descriptions are in the appendix.
3. Descriptions that mislead
searchsays "Powerful semantic search over the entire web with advanced ranking", which says nothing about when to use it instead ofsearch_news. Models pick by name, because the name is the only signal they have.search_newsdoesn't state the 30-day window. Two of the twelve task failures trace to that omission.fetch_pagedoesn't say it returns HTML, or how large the return can be. Models call it, receive 30k tokens, and lose the thread.date_fromsays "ISO date", which models read as a timestamp. The server accepts onlyYYYY-MM-DDand returns an empty result on anything else. Say the format, or better, accept both.
4. Errors agents can't recover from
| Error observed | Frequency | Recovered without a human? | Should say |
|---|---|---|---|
500 {"error":"limit"} on rate limit | 6.1% of calls in bursts | No: models retry immediately and make it worse | 429, Retry-After: 2, body {"error":"rate_limited","retryAfterMs":2000,"limit":"10/s"} |
| Empty result on bad date format | 1.4% | Rarely | 400 {"error":"invalid_date_from","expected":"YYYY-MM-DD","got":"2026-09-01T00:00:00Z"} |
504 from upstream index | 0.3% | Yes, after a retry | Fine; add Retry-After |
5. Onboarding friction
Time to first successful call for an agent with no prior account is not achievable without a human. Three steps need a person. Create an account (email confirmation), add a card (required even for the free tier), copy a key. With x402 on /search and /fetch, time to first call for a funded agent is one request. With programmatic key issuance behind a card-free tier, it's two.
6. Reliability and latency
30 days, three regions, five-minute probes, plus 41,000 proxied calls.
| Region | Availability | p50 | p95 | Notes |
|---|---|---|---|---|
| London | 99.91% | 610 ms | 2.4 s | Two 12-minute outages, both 03:00–04:00 UTC |
| Virginia | 99.97% | 380 ms | 1.6 s | |
| Singapore | 99.62% | 1,150 ms | 4.9 s | No regional endpoint; every call crosses the Pacific |
Your status page showed 100% for the month. The London outages line up with a nightly index rebuild. A status entry, or a 503 with Retry-After, would have turned a negative event into a documented maintenance window.
7. Security posture (external)
- API key accepted in the query string (
?key=) as well as the header. Query-string keys end up in logs and referrers. Deprecate with a date. - No read-only mode is needed (all tools are read-only), but the tools carry no
readOnlyHintannotation, so hosts can't tell. fetch_pagereturns page content verbatim. Two of our probe pages contained instruction-shaped text and both came back unmarked. Wrap fetched content with a boundary marker and a warning, and document it.- No
/.well-known/security.txt.
8. Prioritised fixes
| # | Fix | Effort | Expected movement |
|---|---|---|---|
| 1 | Return 429 + Retry-After for rate limits; structured 400 for bad parameters | Small | Ergonomics +18, Task success +12 → +4.2 total |
| 2 | Collapse eleven tools into three; rewrite descriptions (drafts attached) | Medium | Schema +22, Ergonomics +10 → +4.2 total |
| 3 | Markdown by default from fetch; document size and add max_length | Small | Ergonomics +8, Task success +8 → +1.8 total |
| 4 | x402 on /search and /fetch at your published prices | Small (Go: ~40 lines) | Payments +40 → +4.0 total |
| 5 | Card-free tier or programmatic key issuance | Medium | Payments +20 → +2.0 total |
| 6 | Deprecate query-string keys; add readOnlyHint; wrap fetched content | Small | Security +14 → +2.0 total |
| 7 | Status page entries or 503s for the nightly rebuild; Singapore endpoint or CDN | Medium | Reliability +6, Performance +10 → +2.0 total |
Projected after fixes 1 to 5, 79 to 81 (A). After all seven, 82 to 84 (A).
Appendix
report.json, every table above as data, plus transcripts and probe logs.descriptions.md, the proposed tool definitions.- A re-run, meaning a reduced-price re-assessment after the changes ship, with a before/after diff.
*Harbourline Search is fictional. Any resemblance to a real search server with eleven tools and a 500 on rate limits is a coincidence the real server should look into.*