# Best agent tracing, monitoring and evaluation tools > Arize Phoenix (BB), Langfuse API + MCP (BB) and LangSmith API + MCP (BB) lead the 15 ranked agent tracing, monitoring and evaluation tools. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Canonical: https://www.anchorterminal.com/best/agent-observability/ - Markdown: https://www.anchorterminal.com/best/agent-observability/index.md (~6,400 tokens) - Slim: https://www.anchorterminal.com/best/agent-observability/index.min.md (~1,580 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/best/agent-observability/index.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 The 10 highest-scoring of 15 agent tracing, monitoring and evaluation tools on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 15 · agent-ready (BB or better): 3 · accept x402: 0 · hosted endpoints: 12 - Full ranked table: https://www.anchorterminal.com/categories/agent-observability.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/agent-observability/index.md (120) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Arize Phoenix](https://www.anchorterminal.com/tools/arize-phoenix.md) | BB | 75.4 | Teams that want tracing and evals on their own hardware, or air-gapped, with an MCP endpoint an agent can sign into. | Free · OSS | local | | 2 | [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md) | BB | 72.7 | Teams that want open source they can self-host or a predictable cloud bill, across tracing, evals, prompts and datasets. | $29 / mo | hosted | | 3 | [LangSmith API + MCP](https://www.anchorterminal.com/tools/langsmith.md) | BB | 71.1 | Teams on LangChain or LangGraph, and for any team that wants mature hosted evals with clear limits and deprecation rules. | $39 / seat-mo | hosted | | 4 | [W&B Weave](https://www.anchorterminal.com/tools/wandb-weave.md) | B | 66.7 | Teams already on Weights & Biases, or with OpenTelemetry instrumentation, who want traces, evaluations, scorers, datasets and prompts in one place. | $60 / mo | hosted | | 5 | [Prefactor](https://www.anchorterminal.com/tools/prefactor.md) | B | 65.6 | A team that wants an auditable record of agent runs with declared data-risk labels, quality payloads from its own evaluations and a remote stop signal. | $49 / mo | hosted | | 6 | [Respan API + MCP](https://www.anchorterminal.com/tools/respan.md) | B | 65.5 | Teams that want a gateway and tracing in one integration and an agent that can manage prompts, datasets and evaluators over MCP. | $199 / mo | hosted | | 7 | [LangWatch](https://www.anchorterminal.com/tools/langwatch.md) | B | 65.5 | Teams that want tracing, evaluations and simulated-user agent tests in one open-source product, hosted in the EU or self-hosted, and that drive it from a coding assistant through MCP or the CLI. | Freemium | hosted and local | | 8 | [Pydantic Logfire](https://www.anchorterminal.com/tools/pydantic-logfire.md) | B | 64.9 | Teams that want agent traces alongside application logs and metrics in one OpenTelemetry store, queried by SQL from a coding assistant. | $49 / mo | hosted | | 9 | [DeepEval](https://www.anchorterminal.com/tools/deepeval.md) | B | 64.7 | Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics. | $200 / mo | local | | 10 | [MLflow Tracing](https://www.anchorterminal.com/tools/mlflow-tracing.md) | C | 61.2 | Teams that already run MLflow or want Apache-2.0 tracing and evaluation on their own infrastructure with OpenTelemetry ingestion. | Free · OSS | local | ## Picks by need - Highest score overall: [Arize Phoenix](https://www.anchorterminal.com/tools/arize-phoenix.md), BB, 75.4/100 on the benchmark. Also [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md), BB, 72.7/100. - Reliability: [LangSmith API + MCP](https://www.anchorterminal.com/tools/langsmith.md), 81/100 on reliability, against 77 for the overall leader. - Schema & documentation: [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md), 93/100 on schema & documentation, against 88 for the overall leader. - Security & auth: [Pydantic Logfire](https://www.anchorterminal.com/tools/pydantic-logfire.md), 80/100 on security & auth, against 56 for the overall leader. - Maintenance & community: [LangWatch](https://www.anchorterminal.com/tools/langwatch.md), 90/100 on maintenance & community, against 84 for the overall leader. - Transparency & trust: [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md), 85/100 on transparency & trust, against 73 for the overall leader. - Lowest paid price per GB: [Braintrust API + MCP](https://www.anchorterminal.com/tools/braintrust.md), $0.50 per GB, the lowest of the 3 listings here with a paid price in this unit (free allowances aside). Also [Laminar API + MCP](https://www.anchorterminal.com/tools/laminar.md), $1.50 per GB. - A hosted MCP endpoint: [Langfuse API + MCP](https://www.anchorterminal.com/tools/langfuse.md), remote MCP server, nothing to install. ## How to choose - Steps and retries in traces: Check whether tool calls, model calls and retries each appear as a separate span, since a merged retry hides the failure the agent is trying to find. - Finding a failed tool call: Check whether a failed tool call can be found by filtering on status or error, since scanning every run by hand slows an agent's debugging down. - Reproducible evaluation runs: Check whether an evaluation can be rerun against a fixed dataset and prompt version, because a score that cannot be reproduced cannot show whether a change helped. - Trace data location and retention: Check the data location, retention period and ownership of traces, since runs can hold prompts, tool outputs and customer data that the operator stays responsible for. - How the benchmark tests this category: The same agent run instrumented in every tool, including a failed tool call and a retry. We check which steps appear in the trace, whether the failure can be found quickly, whether an evaluation can be reproduced, how long setup takes and what it costs. ## Each one in detail ### 1. Arize Phoenix, BB 75.4/100 Self-hosted tracing, evaluation, datasets, experiments and prompt management built on OpenTelemetry and OpenInference. - Verdict: Free and self-hosted with no feature gating, from `pip install` to a Helm chart. Auth is off by default and the default admin password is `admin`. - Choose it for: Teams that want tracing and evals on their own hardware, or air-gapped, with an MCP endpoint an agent can sign into. - Strength: Free and self-hosted with no feature gating, from `pip install` to a Helm chart - Strength: MCP endpoint generated from the OpenAPI, five code-mode tools by default, OAuth 2.1 with PKCE and audience-bound tokens - Strength: Tool annotations derived from HTTP verbs, so a client can auto-approve reads and confirm writes - Weakness: Auth is off by default and the default admin password is `admin` - Weakness: Remote MCP endpoint still labelled beta in the docs - Weakness: Elastic License 2.0 isn't OSI open source and forbids running Phoenix as a managed service - Price: Free · OSS · Auth: OAuth or key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/arize-phoenix.md ### 2. Langfuse API + MCP, BB 72.7/100 Open-source tracing, evaluation, prompt management and datasets for LLM apps and agents, built on OpenTelemetry. - Verdict: MIT core with no usage limits when self-hosted, and self-hosted telemetry documented with an off switch. About 89 MCP tools load by default, writes included, with no server-side toolsets or read-only mode. - Choose it for: Teams that want open source they can self-host or a predictable cloud bill, across tracing, evals, prompts and datasets. - Strength: MIT core with no usage limits when self-hosted, and self-hosted telemetry documented with an off switch - Strength: Rate limits per bucket and plan, with 429 and `Retry-After` documented - Strength: Per-unit cloud pricing at $8 per 100,000 units, and a free Hobby plan with 50,000 units and no card - Weakness: About 89 MCP tools load by default, writes included, with no server-side toolsets or read-only mode - Weakness: 12 degraded incidents on the status page between 3 July and 30 September, mostly ingestion delays - Weakness: No uptime SLA found, only support response targets - Price: $29 / mo · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/langfuse.md - Against #1: https://www.anchorterminal.com/compare/arize-phoenix-vs-langfuse.md ### 3. LangSmith API + MCP, BB 71.1/100 LangChain's hosted tracing, evaluation, prompt hub and dataset platform, framework-agnostic with Python and TypeScript SDKs and OpenTelemetry ingest. - Verdict: Rate limits published per endpoint and per plan, with each kind of 429 explained. Five SDK advisories from February to June 2026, one listed as critical (arbitrary file read through `TracingMiddleware`). - Choose it for: Teams on LangChain or LangGraph, and for any team that wants mature hosted evals with clear limits and deprecation rules. - Strength: Rate limits published per endpoint and per plan, with each kind of 429 explained - Strength: Written deprecation policy, six months on cloud, with `Deprecation` and `Sunset` headers on deprecated endpoints - Strength: Remote MCP with OAuth 2.1, read-only in practice, and `fetch_runs` paging by a 25,000-character budget - Weakness: Five SDK advisories from February to June 2026, one listed as critical (arbitrary file read through `TracingMiddleware`) - Weakness: Base traces keep 14 days, and extended retention was capped at 180 days with about a week's notice - Weakness: Closed-source platform, self-hosting, BYOC and key roles are Enterprise only - Price: $39 / seat-mo · Auth: OAuth or key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/langsmith.md - Against #1: https://www.anchorterminal.com/compare/arize-phoenix-vs-langsmith.md ### 4. W&B Weave, B 66.7/100 W&B Weave is CoreWeave Forge's tracing and evaluation service for LLM applications and agents. It takes traces from Python and TypeScript SDKs or an OTLP endpoint and exposes them through a REST Service API. - Verdict: The Service API at trace.wandb.ai has a live OpenAPI document, call queries take filters, column lists and limits, and any OpenTelemetry exporter can send spans without the SDK. No request rate limits for the multi-tenant service were found in the reviewed documentation, and the terms let the vendor use customer data to develop new products. - Choose it for: Teams already on Weights & Biases, or with OpenTelemetry instrumentation, who want traces, evaluations, scorers, datasets and prompts in one place. - Strength: A live OpenAPI 3.1 document at `https://trace.wandb.ai/openapi.json` describes 127 operations on 114 paths - Strength: Two OTLP endpoints accept protobuf spans from any OpenTelemetry exporter, so tracing needs no Weave SDK - Strength: `/calls/stream_query` takes `filter`, `query`, `sort_by`, `columns`, `limit` and `offset`, so an agent can size each response - Weakness: No request rate limits for the multi-tenant Service API were found in the reviewed documentation - Weakness: The OpenAPI document lists only 200 and 422 responses, and 51 of 127 operations have no description - Weakness: The Master Service Agreement lets W&B use Customer Data to improve its services, develop new products and run AI functions - Price: $60 / mo · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/wandb-weave.md - Against #1: https://www.anchorterminal.com/compare/arize-phoenix-vs-wandb-weave.md ### 5. Prefactor, B 65.6/100 Prefactor is a hosted service that records AI agent runs as spans, classifies each run's data risk and stores quality evaluations. Agents and scripts reach it through an HTTP and WebSocket API, TypeScript and Python SDKs and a CLI. - Verdict: The public OpenAPI document lists 157 operations, each accepting an idempotency key, and tokens can be read-only or bound to one agent deployment. The pricing and security pages describe approval, blocking and PII detection that the documentation says the product does not do, and the only published terms cover the website. - Choose it for: A team that wants an auditable record of agent runs with declared data-risk labels, quality payloads from its own evaluations and a remote stop signal. - Strength: Public OpenAPI 3.0 document with 157 operations and an OpenRPC document with 159 WebSocket methods, both fetched without a login - Strength: Every write accepts an `idempotency_key`, and `POST /bulk` can be retried whole because a reused key fails only its own item - Strength: Tokens are scoped to the account or to one agent deployment, carry a `read_only` or `full_access` role and an expiry, and can be suspended or revoked - Weakness: The pricing page lists hold, approve or block and PII detection on every plan, while the docs say a risk threshold blocks nothing and that Prefactor does not detect sensitive data - Weakness: No approval operation exists in the OpenAPI or OpenRPC documents. The documented control is a cooperative terminate signal the agent's own code must obey - Weakness: The published terms of service cover the website. No service agreement, DPA, sub-processor list or SLA text was found - Price: $49 / mo · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/prefactor.md - Against #1: https://www.anchorterminal.com/compare/arize-phoenix-vs-prefactor.md ### 6. Respan API + MCP, B 65.5/100 Hosted tracing, evals, prompt management and an LLM gateway to 250+ models, from the company formerly called Keywords AI. - Verdict: API keys with Read, Write, Gateway and Admin scopes, expiry and spend caps, bound to one project. 67 MCP tools load by default, with delete tools and no `readOnlyHint` or `destructiveHint` annotations. - Choose it for: Teams that want a gateway and tracing in one integration and an agent that can manage prompts, datasets and evaluators over MCP. - Strength: API keys with Read, Write, Gateway and Admin scopes, expiry and spend caps, bound to one project - Strength: OpenAPI 3.1 spec, llms.txt and Markdown for every docs page - Strength: `Respan-Enabled-Tools` header trims the MCP server's tool list server-side - Weakness: 67 MCP tools load by default, with delete tools and no `readOnlyHint` or `destructiveHint` annotations - Weakness: A 13-hour 47-minute gateway incident on 26 September 2026, while the status bars show 100 per cent - Weakness: Security fixes on 30 July and 28 August 2026 disclosed only as one-line changelog entries - Price: $199 / mo · Auth: OAuth or key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/respan.md - Against #1: https://www.anchorterminal.com/compare/arize-phoenix-vs-respan.md ### 7. LangWatch, B 65.5/100 LangWatch is an open-source platform for tracing, evaluating and testing LLM applications and agents, with prompt management, datasets and an AI gateway. It runs hosted or self-hosted, with a REST API, SDKs, a CLI and an MCP server. - Verdict: API keys can be limited to read or write per permission category, expire, and be revoked, and the remote MCP server uses OAuth with PKCE. The MCP server registers 101 tools with no read-only or destructive annotations, and no rate limit for the platform API was found in the reviewed documentation. - Choose it for: Teams that want tracing, evaluations and simulated-user agent tests in one open-source product, hosted in the EU or self-hosted, and that drive it from a coding assistant through MCP or the CLI. - Strength: API keys take read or write per permission category, a project, team or organisation scope and an expiry. Ingestion keys can only write traces - Strength: A public OpenAPI 3.1 document covers 328 operations and is served without a key, with llms.txt and a Markdown twin of every docs page - Strength: Apache 2.0 platform with MIT SDKs and MCP server. Self-hosting has no volume cap, and every outbound call is documented with its off switch - Weakness: The MCP server registers 101 tools, deletes and key creation among them, with no toolsets and no `readOnlyHint` or `destructiveHint` annotations in the source - Weakness: No request rate limit for the platform API was found in the reviewed documentation. Only two operations in the OpenAPI document declare a 429 - Weakness: The status page shows Scenarios down for 9 hours 34 minutes on 8 September 2026 and trace processing degraded for 14 hours 44 minutes on 31 July - Price: Freemium · Auth: OAuth or key · x402: no · Where: hosted and local - Full assessment: https://www.anchorterminal.com/tools/langwatch.md - Against #1: https://www.anchorterminal.com/compare/arize-phoenix-vs-langwatch.md ### 8. Pydantic Logfire, B 64.9/100 Pydantic Logfire is a hosted observability platform built on OpenTelemetry for traces, logs, metrics, LLM cost tracking, evaluations and prompt management. Agents reach it through a remote MCP server, a SQL query API, a public REST API and SDKs. - Verdict: OAuth with PKCE and dynamic client registration, 44 scopes and a public OpenAPI 3.1 document make access easy to limit and to script, and the free Personal plan needs no card. No status page was found, query limits are published as named levels, not numbers, and the hosted MCP server's tool schemas could not be read. - Choose it for: Teams that want agent traces alongside application logs and metrics in one OpenTelemetry store, queried by SQL from a coding assistant. - Strength: The remote MCP server uses OAuth with PKCE (S256 required) and dynamic client registration, or a Bearer API key with as little as the `project:read` scope - Strength: A public OpenAPI 3.1 document of 68 paths and 106 operations is served without a key, with llms.txt and a Markdown twin of every docs page - Strength: The Personal plan is free with no card, 10 million records a month and a hard cap at $0. Team and Growth charge $2 per million records beyond that - Weakness: No public status page was found on pydantic.dev, in the docs or in the files published for agents - Weakness: Query limits for MCP, read tokens and the public API are published as Low, Standard and High per plan, with no numbers beyond a daily request count on the pricing page - Weakness: The hosted MCP server is closed source and its discovery card lists 51 tools, 31 of them creates, updates or deletes. The card says it is maintained by hand and its tool names differ from the docs - Price: $49 / mo · Auth: OAuth or key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/pydantic-logfire.md - Against #1: https://www.anchorterminal.com/compare/arize-phoenix-vs-pydantic-logfire.md ### 9. DeepEval, B 64.7/100 DeepEval is an open-source Python and TypeScript framework from Confident AI for evaluating LLM applications and agents. Evaluations run locally through a pytest plugin and the `deepeval` CLI, with optional reporting to the hosted Confident AI platform. - Verdict: DeepEval runs evaluations and tracing locally under Apache 2.0 with no account, and writes each test run to JSON or SQLite. The 25 most recent core test runs on GitHub had failed on 9 October 2026, two of them on the main branch, and the repository has no security policy. - Choose it for: Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics. - Strength: Apache-2.0 framework that runs evaluations and tracing locally with no account. Test runs are written to JSON files or a SQLite database on the owner's disk - Strength: Python 4.2.8 was tagged on 2 October 2026, with 14 Python releases tagged since 9 August and TypeScript 0.9.21 on 30 September - Strength: `deepeval test run` takes flags for parallel processes, repeats, a result cache, ignoring errors and skipping cases with missing parameters - Weakness: The 25 most recent Py Core Tests runs had failed when read on 9 October 2026, two of them pushes to main. Issue #3372 of 26 September reports the same - Weakness: No SECURITY.md, no published advisories and no security.txt on deepeval.com or confident-ai.com. The trust centre is drawn by script and was not read - Weakness: Release 4.2.0 reversed the score direction of four safety metrics in a minor version. The changelog marks it as breaking and the code warns at run time - Price: $200 / mo · Auth: OAuth or key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/deepeval.md - Against #1: https://www.anchorterminal.com/compare/arize-phoenix-vs-deepeval.md ### 10. MLflow Tracing, C 61.2/100 Open-source tracing, evaluation and prompt management for LLM applications and agents, part of MLflow, a Linux Foundation project. Owners run the server themselves, and agents read and annotate traces through an experimental MCP server or the `mlflow traces` CLI. - Verdict: Apache-2.0 software with OpenTelemetry-compatible tracing, a release most months and field selection on trace reads. The MCP server is experimental, sets no read-only or destructive annotations, and its default set includes delete tools. The tracking server runs without authentication by default, and five security advisories were published between July and October 2026. - Choose it for: Teams that already run MLflow or want Apache-2.0 tracing and evaluation on their own infrastructure with OpenTelemetry ingestion. - Strength: Apache-2.0 licence, free to self-host, with nothing to buy from the project - Strength: `extract_fields` on `search_traces` and `get_trace` returns only the named fields, with `max_results` and `page_token` for paging - Strength: The server accepts OTLP at `/v1/traces`, so applications in any OpenTelemetry language can send spans - Weakness: The MCP server is marked experimental in the docs and sets no `readOnlyHint` or `destructiveHint` on any tool - Weakness: The tracking server has no authentication unless started with `--app-name basic-auth` - Weakness: Five security advisories published between 27 July and 9 October 2026, one a critical unauthenticated remote code execution fixed in 3.17.0 - Price: Free · OSS · Auth: OAuth or key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/mlflow-tracing.md - Against #1: https://www.anchorterminal.com/compare/arize-phoenix-vs-mlflow-tracing.md 5 more are ranked in the full table: https://www.anchorterminal.com/categories/agent-observability.md ## Head to head - [Arize Phoenix vs Langfuse API + MCP](https://www.anchorterminal.com/compare/arize-phoenix-vs-langfuse.md) - [Arize Phoenix vs LangSmith API + MCP](https://www.anchorterminal.com/compare/arize-phoenix-vs-langsmith.md) - [Arize Phoenix vs W&B Weave](https://www.anchorterminal.com/compare/arize-phoenix-vs-wandb-weave.md) - [Arize Phoenix vs Prefactor](https://www.anchorterminal.com/compare/arize-phoenix-vs-prefactor.md) - [Langfuse API + MCP vs LangSmith API + MCP](https://www.anchorterminal.com/compare/langfuse-vs-langsmith.md) - [Langfuse API + MCP vs W&B Weave](https://www.anchorterminal.com/compare/langfuse-vs-wandb-weave.md) - [Langfuse API + MCP vs Prefactor](https://www.anchorterminal.com/compare/langfuse-vs-prefactor.md) - [LangSmith API + MCP vs W&B Weave](https://www.anchorterminal.com/compare/langsmith-vs-wandb-weave.md) - [LangSmith API + MCP vs Prefactor](https://www.anchorterminal.com/compare/langsmith-vs-prefactor.md) - [Prefactor vs W&B Weave](https://www.anchorterminal.com/compare/prefactor-vs-wandb-weave.md) ## Questions ### What are the highest-rated agent tracing, monitoring and evaluation tools for AI agents? Arize Phoenix has the highest benchmark score of the 15 ranked agent tracing, monitoring and evaluation tools, 75.4 (BB). Langfuse API + MCP is second with 72.7 (BB). ### How many agent tracing, monitoring and evaluation tools are agent-ready? 3 of the 15 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark. ### Which agent tracing, monitoring and evaluation tools accept x402 payments? None of the ranked listings here accepts x402 for its main call yet. ### Which of these agent tracing, monitoring and evaluation tools is cheapest? By published paid prices, Braintrust API + MCP, at $0.50 per GB, the lowest of the 3 listings here with a paid price in this unit (free allowances aside). Plans, volume tiers and free allowances change the sum, so check the listing's price table. ### How is this list ranked? By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026. ## How this list is made The order is the Anchor benchmark score, the same number as on each listing. Each listing is graded from public evidence against the benchmark checklist, and the picks are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.