# Agent tracing, monitoring and evaluation tools > 10 agent observability & evals ranked by the Anchor benchmark. Leader Arize Phoenix (BB). Tools that trace agent runs, tool calls and model calls, and run evaluations on them. Compared on what appears in a trace, how fast a failure can be found, whether an evaluation can be reproduced, setup time and cost. - Canonical: https://www.anchorterminal.com/categories/agent-observability - Markdown: https://www.anchorterminal.com/categories/agent-observability.md (~3,050 tokens) - Slim: https://www.anchorterminal.com/categories/agent-observability.min.md (~530 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/categories/agent-observability.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-04 Tools that trace agent runs, tool calls and model calls, and run evaluations on them. Compared on what appears in a trace, how fast a failure can be found, whether an evaluation can be reproduced, setup time and cost. - Tools ranked: 10 · agent-ready (BB or better): 3 · accept x402: 0 · hosted endpoints: 9 · desk reviews by the panel: 26 - JSON: https://www.anchorterminal.com/api/v1/tools.json (list) · https://www.anchorterminal.com/api/v1/rankings.json (ranked) · https://www.anchorterminal.com/api/v1/x402.json (payable) · https://www.anchorterminal.com/api/v1/capabilities.json (by capability) - Grades run AA, A, BB, B, C, D, E, F · methodology: https://www.anchorterminal.com/benchmark/ - Capabilities in this category: obs.traces, obs.evals, obs.prompts, obs.gateway, obs.datasets - https://letme.dev/obs.traces picks the top-graded tool in this list and says how to call it direct; calling through letme comes later (https://www.anchorterminal.com/letme/index.md) ## Ranking | # | Tool | Vendor | Kind | Category | Grade | Score | Confidence | x402 | Auth | Where | Reviews | Page | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 32 | Arize Phoenix | Arize AI | HTTP API | Evals | BB | 75.6 | medium | no | OAuth or key | local | 3.8/5 (8) | https://www.anchorterminal.com/tools/arize-phoenix.md | | 66 | Langfuse API + MCP | Langfuse (ClickHouse) | HTTP API | Evals | BB | 72.8 | high | no | API key | hosted | 4/5 (2) | https://www.anchorterminal.com/tools/langfuse.md | | 85 | LangSmith API + MCP | LangChain | HTTP API | Evals | BB | 71.3 | medium | no | OAuth or key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/langsmith.md | | 165 | Respan API + MCP | Respan (formerly Keywords AI) | HTTP API | Evals | B | 65.9 | medium | no | OAuth or key | hosted | 2.5/5 (2) | https://www.anchorterminal.com/tools/respan.md | | 229 | Braintrust API + MCP | Braintrust | HTTP API | Evals | C | 61.3 | medium | no | OAuth or key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/braintrust.md | | 298 | Laminar API + MCP | Laminar (LMNR AI) | HTTP API | Evals | C | 57 | medium | no | API key | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/laminar.md | | 310 | HoneyHive | HoneyHive | HTTP API | Evals | C | 55.9 | medium | no | API key | hosted | 2.5/5 (2) | https://www.anchorterminal.com/tools/honeyhive.md | | 378 | Galileo API + MCP | Galileo (now Splunk Agent Observability, Cisco) | HTTP API | Evals | D | 48 | medium | no | API key | hosted | 2.5/5 (2) | https://www.anchorterminal.com/tools/galileo.md | | 387 | Helicone AI Gateway + MCP | Helicone (Mintlify) | HTTP API | Evals | D | 47.1 | medium | no | API key | hosted + local | 2/5 (2) | https://www.anchorterminal.com/tools/helicone.md | | not ranked, shut down | Baserun | Baserun | HTTP API | Evals | F | 7.3 | high | no | API key | hosted | 1/5 (2) | https://www.anchorterminal.com/tools/baserun.md | Scores are from public evidence against the published checklist (https://www.anchorterminal.com/benchmark/), with Performance and Task success pending. p95 latency and context cost come from our probes, which haven't run yet. ## Summaries ### 32. Arize Phoenix, BB (75.6) Self-hosted tracing, evaluation, datasets, experiments and prompt management built on OpenTelemetry and OpenInference. Free and self-hosted with no feature gating, from `pip install` to a Helm chart. Auth is off by default and the default admin password is `admin`. - Page: https://www.anchorterminal.com/tools/arize-phoenix · Markdown: https://www.anchorterminal.com/tools/arize-phoenix.md · JSON: https://www.anchorterminal.com/api/v1/tools/arize-phoenix.json - Capabilities: obs.traces, obs.evals, obs.prompts, obs.datasets ### 66. Langfuse API + MCP, BB (72.8) Open-source tracing, evaluation, prompt management and datasets for LLM apps and agents, built on OpenTelemetry. MIT core with no usage limits when self-hosted, and self-hosted telemetry documented with an off switch. About 89 MCP tools load by default, writes included, with no server-side toolsets or read-only mode. - Page: https://www.anchorterminal.com/tools/langfuse · Markdown: https://www.anchorterminal.com/tools/langfuse.md · JSON: https://www.anchorterminal.com/api/v1/tools/langfuse.json - Capabilities: obs.traces, obs.evals, obs.prompts, obs.datasets · endpoint: `https://cloud.langfuse.com/api/public` ### 85. LangSmith API + MCP, BB (71.3) LangChain's hosted tracing, evaluation, prompt hub and dataset platform, framework-agnostic with Python and TypeScript SDKs and OpenTelemetry ingest. Rate limits published per endpoint and per plan, with each kind of 429 explained. Five SDK advisories from February to June 2026, one listed as critical (arbitrary file read through `TracingMiddleware`). - Page: https://www.anchorterminal.com/tools/langsmith · Markdown: https://www.anchorterminal.com/tools/langsmith.md · JSON: https://www.anchorterminal.com/api/v1/tools/langsmith.json - Capabilities: obs.traces, obs.evals, obs.prompts, obs.datasets, obs.gateway · endpoint: `https://api.smith.langchain.com` ### 165. Respan API + MCP, B (65.9) Hosted tracing, evals, prompt management and an LLM gateway to 250+ models, from the company formerly called Keywords AI. API keys with Read, Write, Gateway and Admin scopes, expiry and spend caps, bound to one project. 67 MCP tools load by default, with delete tools and no `readOnlyHint` or `destructiveHint` annotations. - Page: https://www.anchorterminal.com/tools/respan · Markdown: https://www.anchorterminal.com/tools/respan.md · JSON: https://www.anchorterminal.com/api/v1/tools/respan.json - Capabilities: obs.traces, obs.evals, obs.prompts, obs.gateway, obs.datasets · endpoint: `https://api.respan.ai/api` ### 229. Braintrust API + MCP, C (61.3) Hosted tracing, logging and evaluation for LLM apps and agents, with experiments, datasets, prompts, online scorers and a model gateway. OpenAPI 3.0.3 spec with 75 paths, 429 and `Retry-After` declared on 154 operations. Several major incidents on the status page between 16 July and 30 September 2026, the longest 78 minutes on the US data plane. - Page: https://www.anchorterminal.com/tools/braintrust · Markdown: https://www.anchorterminal.com/tools/braintrust.md · JSON: https://www.anchorterminal.com/api/v1/tools/braintrust.json - Capabilities: obs.traces, obs.evals, obs.prompts, obs.gateway, obs.datasets · endpoint: `https://api.braintrust.dev/v1` ### 298. Laminar API + MCP, C (57) Open-source tracing, evaluations and an agent debugger built on OpenTelemetry, with browser session recordings synced to traces for browser agents. Browser session recordings can be matched to agent traces. A cross-tenant SQL execution and export vulnerability was fixed on 27 August 2026, without a separate advisory. - Page: https://www.anchorterminal.com/tools/laminar · Markdown: https://www.anchorterminal.com/tools/laminar.md · JSON: https://www.anchorterminal.com/api/v1/tools/laminar.json - Capabilities: obs.traces, obs.evals, obs.datasets · endpoint: `https://api.lmnr.ai/v1` ### 310. HoneyHive, C (55.9) Hosted tracing, evaluation, datasets and prompt management for agents, built on OpenTelemetry. Read-only (`hh_ro_`), ingestion-only (`hh_ingst_`) and expiring fine-grained (`hh_fgcp_`) API keys. Status page showed 80.992 per cent uptime for its single component and "Some services are down" on 2 October, with no incident history. - Page: https://www.anchorterminal.com/tools/honeyhive · Markdown: https://www.anchorterminal.com/tools/honeyhive.md · JSON: https://www.anchorterminal.com/api/v1/tools/honeyhive.json - Capabilities: obs.traces, obs.evals, obs.prompts, obs.datasets · endpoint: `https://api.dp1.us.honeyhive.ai` ### 378. Galileo API + MCP, D (48) Hosted tracing, evaluation metrics and guardrails for LLM apps and agents, with a REST API, Python and TypeScript SDKs and an 8-tool MCP server in preview. OpenAPI 3.1 spec with 184 paths and 244 operations. No status page, no published rate limits and no SLA. - Page: https://www.anchorterminal.com/tools/galileo · Markdown: https://www.anchorterminal.com/tools/galileo.md · JSON: https://www.anchorterminal.com/api/v1/tools/galileo.json - Capabilities: obs.traces, obs.evals, obs.prompts, obs.datasets · endpoint: `https://api.galileo.ai/v2` ### 387. Helicone AI Gateway + MCP, D (47.1) Open-source LLM observability that logs requests through an OpenAI-compatible gateway (100+ models, fallbacks, caching, rate limits) or async logging. One base-URL change gives logging, caching, fallbacks and rate limits for 100+ models. Security fixes on 30 August and 16 September 2026 closed cross-tenant access to other customers' decrypted provider keys and an admin takeover, disclosed only in commit messages. - Page: https://www.anchorterminal.com/tools/helicone · Markdown: https://www.anchorterminal.com/tools/helicone.md · JSON: https://www.anchorterminal.com/api/v1/tools/helicone.json - Capabilities: obs.traces, obs.gateway, obs.prompts, obs.datasets, obs.evals · endpoint: `https://ai-gateway.helicone.ai` ### Baserun, F (7.3), retired, not ranked Discontinued LLM testing and monitoring platform with Python and JavaScript SDKs for tracing model calls. SDKs were MIT licensed and the Python source is still public at github.com/baserun-ai/baserun-py. Service offline since late 2024. api.baserun.ai doesn't resolve and app.baserun.ai serves an expired certificate. - Page: https://www.anchorterminal.com/tools/baserun · Markdown: https://www.anchorterminal.com/tools/baserun.md · JSON: https://www.anchorterminal.com/api/v1/tools/baserun.json - Capabilities: obs.traces, obs.evals, obs.prompts · endpoint: `https://app.baserun.ai` ## How we test this category The same agent run instrumented in every tool, including a failed tool call and a retry. We check which steps appear in the trace, whether the failure can be found quickly, whether an evaluation can be reproduced, how long setup takes and what it costs. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence. ## Indexed, not reviewed (5) Sorted into this category from public catalogues, with facts and our own checks but no score, grade or rank (https://www.anchorterminal.com/indexed/index.md). | Listing | Kind | What it does | Why it's here | | --- | --- | --- | --- | | [aegis](https://www.anchorterminal.com/tools/undercurrentholdings-aegis.md) | MCP server | Local stdio MCP server for the AEGIS governance client. The hosted evaluation service is offline. | vendor's own | | [argosvix.com MCP server](https://www.anchorterminal.com/tools/argosvix-server.md) | MCP server | Observability MCP server: query LLM cost/errors/latency & operate alerts/evals from Claude/Cursor | vendor's own | | [CompletionKit](https://www.anchorterminal.com/tools/completionkit-evals.md) | MCP server | Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge. | vendor's own | | [openfeature.dev MCP server](https://www.anchorterminal.com/tools/openfeature-mcp.md) | MCP server | MCP server providing OpenFeature SDK installation guides and OFREP flag evaluation | vendor's own | | [PromptOT](https://www.anchorterminal.com/tools/promptot-mcp.md) | MCP server | Manage, version, and publish LLM prompts with blocks, variables, and evaluations. | vendor's own |