# Best agent tracing, monitoring and evaluation tools (slim) > Arize Phoenix (BB), Langfuse API + MCP (BB) and LangSmith API + MCP (BB) lead the 15 ranked agent tracing, monitoring and evaluation tools. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Full: https://www.anchorterminal.com/best/agent-observability/index.md (~6,400 tokens) · this version ~1,580 tokens · JSON https://www.anchorterminal.com/best/agent-observability/index.json · canonical https://www.anchorterminal.com/best/agent-observability/ - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 The 10 highest-scoring of 15 agent tracing, monitoring and evaluation tools on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 15 · agent-ready (BB or better): 3 · accept x402: 0 · hosted endpoints: 12 - Full ranked table: https://www.anchorterminal.com/categories/agent-observability.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/agent-observability/index.md (120) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Arize Phoenix](https://www.anchorterminal.com/tools/arize-phoenix.md) | BB | 75.4 | Teams that want tracing and evals on their own hardware, or air-gapped, with an MCP endpoint an agent can sign into. | Free · OSS | local | | 2 | [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md) | BB | 72.7 | Teams that want open source they can self-host or a predictable cloud bill, across tracing, evals, prompts and datasets. | $29 / mo | hosted | | 3 | [LangSmith API + MCP](https://www.anchorterminal.com/tools/langsmith.md) | BB | 71.1 | Teams on LangChain or LangGraph, and for any team that wants mature hosted evals with clear limits and deprecation rules. | $39 / seat-mo | hosted | | 4 | [W&B Weave](https://www.anchorterminal.com/tools/wandb-weave.md) | B | 66.7 | Teams already on Weights & Biases, or with OpenTelemetry instrumentation, who want traces, evaluations, scorers, datasets and prompts in one place. | $60 / mo | hosted | | 5 | [Prefactor](https://www.anchorterminal.com/tools/prefactor.md) | B | 65.6 | A team that wants an auditable record of agent runs with declared data-risk labels, quality payloads from its own evaluations and a remote stop signal. | $49 / mo | hosted | | 6 | [Respan API + MCP](https://www.anchorterminal.com/tools/respan.md) | B | 65.5 | Teams that want a gateway and tracing in one integration and an agent that can manage prompts, datasets and evaluators over MCP. | $199 / mo | hosted | | 7 | [LangWatch](https://www.anchorterminal.com/tools/langwatch.md) | B | 65.5 | Teams that want tracing, evaluations and simulated-user agent tests in one open-source product, hosted in the EU or self-hosted, and that drive it from a coding assistant through MCP or the CLI. | Freemium | hosted and local | | 8 | [Pydantic Logfire](https://www.anchorterminal.com/tools/pydantic-logfire.md) | B | 64.9 | Teams that want agent traces alongside application logs and metrics in one OpenTelemetry store, queried by SQL from a coding assistant. | $49 / mo | hosted | | 9 | [DeepEval](https://www.anchorterminal.com/tools/deepeval.md) | B | 64.7 | Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics. | $200 / mo | local | | 10 | [MLflow Tracing](https://www.anchorterminal.com/tools/mlflow-tracing.md) | C | 61.2 | Teams that already run MLflow or want Apache-2.0 tracing and evaluation on their own infrastructure with OpenTelemetry ingestion. | Free · OSS | local | ## Picks by need - Highest score overall: [Arize Phoenix](https://www.anchorterminal.com/tools/arize-phoenix.md), BB, 75.4/100 on the benchmark. Also [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md), BB, 72.7/100. - Reliability: [LangSmith API + MCP](https://www.anchorterminal.com/tools/langsmith.md), 81/100 on reliability, against 77 for the overall leader. - Schema & documentation: [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md), 93/100 on schema & documentation, against 88 for the overall leader. - Security & auth: [Pydantic Logfire](https://www.anchorterminal.com/tools/pydantic-logfire.md), 80/100 on security & auth, against 56 for the overall leader. - Maintenance & community: [LangWatch](https://www.anchorterminal.com/tools/langwatch.md), 90/100 on maintenance & community, against 84 for the overall leader. - Transparency & trust: [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md), 85/100 on transparency & trust, against 73 for the overall leader. - Lowest paid price per GB: [Braintrust API + MCP](https://www.anchorterminal.com/tools/braintrust.md), $0.50 per GB, the lowest of the 3 listings here with a paid price in this unit (free allowances aside). Also [Laminar API + MCP](https://www.anchorterminal.com/tools/laminar.md), $1.50 per GB. - A hosted MCP endpoint: [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md), remote MCP server, nothing to install. ## How to choose - Steps and retries in traces: Check whether tool calls, model calls and retries each appear as a separate span, since a merged retry hides the failure the agent is trying to find. - Finding a failed tool call: Check whether a failed tool call can be found by filtering on status or error, since scanning every run by hand slows an agent's debugging down. - Reproducible evaluation runs: Check whether an evaluation can be rerun against a fixed dataset and prompt version, because a score that cannot be reproduced cannot show whether a change helped. - Trace data location and retention: Check the data location, retention period and ownership of traces, since runs can hold prompts, tool outputs and customer data that the operator stays responsible for. - How the benchmark tests this category: The same agent run instrumented in every tool, including a failed tool call and a retry. We check which steps appear in the trace, whether the failure can be found quickly, whether an evaluation can be reproduced, how long setup takes and what it costs. Each listing's verdict, strengths and weaknesses: https://www.anchorterminal.com/best/agent-observability/index.md