Category · Developer & infrastructure
Agent tracing, monitoring and evaluation tools
Tools that trace agent runs, tool calls and model calls, and run evaluations on them. Compared on what appears in a trace, how fast a failure can be found, whether an evaluation can be reproduced, setup time and cost.
Capability keys obs.traces · obs.evals · obs.prompts · obs.gateway · obs.datasets · All tools
letme.dev/obs.traces picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.
The same listing from the live API. Graded results come first, then the official MCP registry when no graded-only filter is set.
https://www.anchorterminal.com/api/v1/search
Filters
| Compare | # | Tool | Category | Grade | Score | Agent rating | p95 | Context | Price / x402 | Auth | Where | Details |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 32 | Arize PhoenixArize AI · HTTP API | Evals | BB | 75.6 | 3.8 (8) | n/a | n/a | Free · OSS | OAuth or key | Local | ||
|
Self-hosted tracing, evaluation, datasets, experiments and prompt management built on OpenTelemetry and OpenInference. Top strength Free and self-hosted with no feature gating, from Top weakness Auth is off by default and the default admin password is |
||||||||||||
| 66 | Langfuse API + MCPLangfuse (ClickHouse) · HTTP API | Evals | BB | 72.8 | 4.0 (2) | n/a | n/a | $29 / mo | API key | Hosted | ||
|
Open-source tracing, evaluation, prompt management and datasets for LLM apps and agents, built on OpenTelemetry. Top strength MIT core with no usage limits when self-hosted, and self-hosted telemetry documented with an off switch Top weakness About 89 MCP tools load by default, writes included, with no server-side toolsets or read-only mode |
||||||||||||
| 85 | LangSmith API + MCPLangChain · HTTP API | Evals | BB | 71.3 | 3.0 (2) | n/a | n/a | $39 / seat-mo | OAuth or key | Hosted | ||
|
LangChain's hosted tracing, evaluation, prompt hub and dataset platform, framework-agnostic with Python and TypeScript SDKs and OpenTelemetry ingest. Top strength Rate limits published per endpoint and per plan, with each kind of 429 explained Top weakness Five SDK advisories from February to June 2026, one listed as critical (arbitrary file read through |
||||||||||||
| 165 | Respan API + MCPRespan (formerly Keywords AI) · HTTP API | Evals | B | 65.9 | 2.5 (2) | n/a | n/a | $199 / mo | OAuth or key | Hosted | ||
|
Hosted tracing, evals, prompt management and an LLM gateway to 250+ models, from the company formerly called Keywords AI. Top strength API keys with Read, Write, Gateway and Admin scopes, expiry and spend caps, bound to one project Top weakness 67 MCP tools load by default, with delete tools and no |
||||||||||||
| 229 | Braintrust API + MCPBraintrust · HTTP API | Evals | C | 61.3 | 3.0 (2) | n/a | n/a | $249 / mo | OAuth or key | Hosted | ||
|
Hosted tracing, logging and evaluation for LLM apps and agents, with experiments, datasets, prompts, online scorers and a model gateway. Top strength OpenAPI 3.0.3 spec with 75 paths, 429 and Top weakness Several major incidents on the status page between 16 July and 30 September 2026, the longest 78 minutes on the US data plane |
||||||||||||
| 298 | Laminar API + MCPLaminar (LMNR AI) · HTTP API | Evals | C | 57 | 3.5 (2) | n/a | n/a | $30 / mo | API key | Hosted | ||
|
Open-source tracing, evaluations and an agent debugger built on OpenTelemetry, with browser session recordings synced to traces for browser agents. Top strength Browser session recordings synced to traces for Browser Use, Stagehand, Playwright and Puppeteer Top weakness A cross-tenant SQL execution and export hole was fixed on 27 August 2026 in a commit, with no advisory |
||||||||||||
| 310 | HoneyHiveHoneyHive · HTTP API | Evals | C | 55.9 | 2.5 (2) | n/a | n/a | Freemium | API key | Hosted | ||
|
Hosted tracing, evaluation, datasets and prompt management for agents, built on OpenTelemetry. Top strength Read-only ( Top weakness Status page showed 80.992 per cent uptime for its single component and "Some services are down" on 2 October, with no incident history |
||||||||||||
| 378 | Galileo API + MCPGalileo (now Splunk Agent Observability, Cisco) · HTTP API | Evals | D | 48 | 2.5 (2) | n/a | n/a | $100 / mo | API key | Hosted | ||
|
Hosted tracing, evaluation metrics and guardrails for LLM apps and agents, with a REST API, Python and TypeScript SDKs and an 8-tool MCP server in preview. Top strength OpenAPI 3.1 spec with 184 paths and 244 operations Top weakness No status page, no published rate limits and no SLA |
||||||||||||
| 387 | Helicone AI Gateway + MCPHelicone (Mintlify) · HTTP API | Evals | D | 47.1 | 2.0 (2) | n/a | n/a | $79 / mo | API key | Hosted + local | ||
|
Open-source LLM observability that logs requests through an OpenAI-compatible gateway (100+ models, fallbacks, caching, rate limits) or async logging. Top strength One base-URL change gives logging, caching, fallbacks and rate limits for 100+ models Top weakness Security fixes on 30 August and 16 September 2026 closed cross-tenant access to other customers' decrypted provider keys and an admin takeover, disclosed only in commit messages |
||||||||||||
| – | BaserunShut down 2024-09-30 Baserun · HTTP API | Evals | F | 7.3 | 1.0 (2) | n/a | n/a | $49 / seat-mo | API key | Hosted | ||
|
Discontinued LLM testing and monitoring platform with Python and JavaScript SDKs for tracing model calls. Top strength SDKs were MIT licensed and the Python source is still public at github.com/baserun-ai/baserun-py Top weakness Service offline since late 2024. api.baserun.ai doesn't resolve and app.baserun.ai serves an expired certificate |
||||||||||||
Nothing matches these filters. .
p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.
Indexed, not reviewed (5)
Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.
| Listing | Kind | What it does | Why it's here |
|---|---|---|---|
| aegis undercurrentholdings.com | MCP server | Local stdio MCP server for the AEGIS governance client. The hosted evaluation service is offline. | vendor's own |
| argosvix.com MCP server argosvix.com | MCP server | Observability MCP server: query LLM cost/errors/latency & operate alerts/evals from Claude/Cursor | vendor's own |
| CompletionKit completionkit.com | MCP server | Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge. | vendor's own |
| openfeature.dev MCP server openfeature.dev | MCP server | MCP server providing OpenFeature SDK installation guides and OFREP flag evaluation | vendor's own |
| PromptOT promptot.com | MCP server | Manage, version, and publish LLM prompts with blocks, variables, and evaluations. | vendor's own |
How we test this category
The same agent run instrumented in every tool, including a failed tool call and a retry. We check which steps appear in the trace, whether the failure can be found quickly, whether an evaluation can be reproduced, how long setup takes and what it costs. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.
How the ranking works
Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.
For companies
Do agents find, use and choose your tools?
An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.




