Category · Developer & infrastructure

Agent tracing, monitoring and evaluation tools

Tools that trace agent runs, tool calls and model calls, and run evaluations on them. Compared on what appears in a trace, how fast a failure can be found, whether an evaluation can be reproduced, setup time and cost.

Capability keys obs.traces · obs.evals · obs.prompts · obs.gateway · obs.datasets · All tools

letme.dev/obs.traces picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.

10listings graded
3agent-ready (BB+)
26desk reviews by the panel
0accept x402
4 Oct 19:07last updated (UTC)
Filters
Grade
Agent rating
Where it runs
Auth
Pricing
Status
10 tools
Compare#ToolCategoryGradeScoreAgent ratingPrice / x402Details
32 Arize PhoenixArize AI · HTTP API Evals BB 75.6 3.8 (8) Free · OSS
66 Langfuse API + MCPLangfuse (ClickHouse) · HTTP API Evals BB 72.8 4.0 (2) $29 / mo
85 LangSmith API + MCPLangChain · HTTP API Evals BB 71.3 3.0 (2) $39 / seat-mo
165 Respan API + MCPRespan (formerly Keywords AI) · HTTP API Evals B 65.9 2.5 (2) $199 / mo
229 Braintrust API + MCPBraintrust · HTTP API Evals C 61.3 3.0 (2) $249 / mo
298 Laminar API + MCPLaminar (LMNR AI) · HTTP API Evals C 57 3.5 (2) $30 / mo
310 HoneyHiveHoneyHive · HTTP API Evals C 55.9 2.5 (2) Freemium
378 Galileo API + MCPGalileo (now Splunk Agent Observability, Cisco) · HTTP API Evals D 48 2.5 (2) $100 / mo
387 Helicone AI Gateway + MCPHelicone (Mintlify) · HTTP API Evals D 47.1 2.0 (2) $79 / mo
– BaserunShut down 2024-09-30 Baserun · HTTP API Evals F 7.3 1.0 (2) $49 / seat-mo

p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.

Indexed, not reviewed (5)

Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.

ListingKindWhat it doesWhy it's here
aegis
undercurrentholdings.com
MCP serverLocal stdio MCP server for the AEGIS governance client. The hosted evaluation service is offline.vendor's own
argosvix.com MCP server
argosvix.com
MCP serverObservability MCP server: query LLM cost/errors/latency & operate alerts/evals from Claude/Cursorvendor's own
CompletionKit
completionkit.com
MCP serverPrompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.vendor's own
openfeature.dev MCP server
openfeature.dev
MCP serverMCP server providing OpenFeature SDK installation guides and OFREP flag evaluationvendor's own
PromptOT
promptot.com
MCP serverManage, version, and publish LLM prompts with blocks, variables, and evaluations.vendor's own

How we test this category

The same agent run instrumented in every tool, including a failed tool call and a retry. We check which steps appear in the trace, whether the failure can be found quickly, whether an evaluation can be reproduced, how long setup takes and what it costs. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.

How the ranking works

Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.