Best of · Developer & infrastructure
Best agent tracing, monitoring and evaluation tools
The 10 highest-scoring of 15 agent tracing, monitoring and evaluation tools on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 15 ranked
- 3 agent-ready
- 12 hosted endpoints
- Updated 9 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Schema & documentation
93/100 on schema & documentation, against 88 for the overall leader.
Maintenance & community
90/100 on maintenance & community, against 84 for the overall leader.
Transparency & trust
85/100 on transparency & trust, against 73 for the overall leader.
Lowest paid price per GB
$0.50 per GB, the lowest of the 3 listings here with a paid price in this unit (free allowances aside).
Also Laminar API + MCP, $1.50 per GB.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | Arize Phoenix Arize AI |
BB 75.4 | Teams that want tracing and evals on their own hardware, or air-gapped, with an MCP endpoint an agent can sign into. | Free · OSS | local |
| 2 | Langfuse API + MCP Langfuse (ClickHouse) |
BB 72.7 | Teams that want open source they can self-host or a predictable cloud bill, across tracing, evals, prompts and datasets. | $29 / mo | hosted |
| 3 | LangSmith API + MCP LangChain |
BB 71.1 | Teams on LangChain or LangGraph, and for any team that wants mature hosted evals with clear limits and deprecation rules. | $39 / seat-mo | hosted |
| 4 | W&B Weave Weights & Biases (CoreWeave) |
B 66.7 | Teams already on Weights & Biases, or with OpenTelemetry instrumentation, who want traces, evaluations, scorers, datasets and prompts in one place. | $60 / mo | hosted |
| 5 | Prefactor Prefactor Pty Ltd |
B 65.6 | A team that wants an auditable record of agent runs with declared data-risk labels, quality payloads from its own evaluations and a remote stop signal. | $49 / mo | hosted |
| 6 | Respan API + MCP Respan (formerly Keywords AI) |
B 65.5 | Teams that want a gateway and tracing in one integration and an agent that can manage prompts, datasets and evaluators over MCP. | $199 / mo | hosted |
| 7 | LangWatch Reasoning Engine B.V. (LangWatch) |
B 65.5 | Teams that want tracing, evaluations and simulated-user agent tests in one open-source product, hosted in the EU or self-hosted, and that drive it from a coding assistant through MCP or the CLI. | Freemium | hosted and local |
| 8 | Pydantic Logfire Pydantic Services Inc. |
B 64.9 | Teams that want agent traces alongside application logs and metrics in one OpenTelemetry store, queried by SQL from a coding assistant. | $49 / mo | hosted |
| 9 | DeepEval Confident AI, Inc. |
B 64.7 | Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics. | $200 / mo | local |
| 10 | MLflow Tracing MLflow Project (LF Projects, LLC) |
C 61.2 | Teams that already run MLflow or want Apache-2.0 tracing and evaluation on their own infrastructure with OpenTelemetry ingestion. | Free · OSS | local |
5 more are ranked in the full table.
How to choose
- Steps and retries in tracesCheck whether tool calls, model calls and retries each appear as a separate span, since a merged retry hides the failure the agent is trying to find.
- Finding a failed tool callCheck whether a failed tool call can be found by filtering on status or error, since scanning every run by hand slows an agent's debugging down.
- Reproducible evaluation runsCheck whether an evaluation can be rerun against a fixed dataset and prompt version, because a score that cannot be reproduced cannot show whether a change helped.
- Trace data location and retentionCheck the data location, retention period and ownership of traces, since runs can hold prompts, tool outputs and customer data that the operator stays responsible for.
How the benchmark tests this category. The same agent run instrumented in every tool, including a failed tool call and a retry. We check which steps appear in the trace, whether the failure can be found quickly, whether an evaluation can be reproduced, how long setup takes and what it costs.
Each one in detail
Arize Phoenix
BB 75.4/100Self-hosted tracing, evaluation, datasets, experiments and prompt management built on OpenTelemetry and OpenInference.
Verdict Free and self-hosted with no feature gating, from pip install to a Helm chart. Auth is off by default and the default admin password is admin.
Choose it for Teams that want tracing and evals on their own hardware, or air-gapped, with an MCP endpoint an agent can sign into.
Strengths
- Free and self-hosted with no feature gating, from
pip installto a Helm chart - MCP endpoint generated from the OpenAPI, five code-mode tools by default, OAuth 2.1 with PKCE and audience-bound tokens
- Tool annotations derived from HTTP verbs, so a client can auto-approve reads and confirm writes
Weaknesses
- Auth is off by default and the default admin password is
admin - Remote MCP endpoint still labelled beta in the docs
- Elastic License 2.0 isn't OSI open source and forbids running Phoenix as a managed service
Price Free · OSSAuth OAuth or keyx402 nolocal
Langfuse API + MCP
BB 72.7/100Open-source tracing, evaluation, prompt management and datasets for LLM apps and agents, built on OpenTelemetry.
Verdict MIT core with no usage limits when self-hosted, and self-hosted telemetry documented with an off switch. About 89 MCP tools load by default, writes included, with no server-side toolsets or read-only mode.
Choose it for Teams that want open source they can self-host or a predictable cloud bill, across tracing, evals, prompts and datasets.
Strengths
- MIT core with no usage limits when self-hosted, and self-hosted telemetry documented with an off switch
- Rate limits per bucket and plan, with 429 and
Retry-Afterdocumented - Per-unit cloud pricing at $8 per 100,000 units, and a free Hobby plan with 50,000 units and no card
Weaknesses
- About 89 MCP tools load by default, writes included, with no server-side toolsets or read-only mode
- 12 degraded incidents on the status page between 3 July and 30 September, mostly ingestion delays
- No uptime SLA found, only support response targets
Price $29 / moAuth API keyx402 nohosted
LangSmith API + MCP
BB 71.1/100LangChain's hosted tracing, evaluation, prompt hub and dataset platform, framework-agnostic with Python and TypeScript SDKs and OpenTelemetry ingest.
Verdict Rate limits published per endpoint and per plan, with each kind of 429 explained. Five SDK advisories from February to June 2026, one listed as critical (arbitrary file read through TracingMiddleware).
Choose it for Teams on LangChain or LangGraph, and for any team that wants mature hosted evals with clear limits and deprecation rules.
Strengths
- Rate limits published per endpoint and per plan, with each kind of 429 explained
- Written deprecation policy, six months on cloud, with
DeprecationandSunsetheaders on deprecated endpoints - Remote MCP with OAuth 2.1, read-only in practice, and
fetch_runspaging by a 25,000-character budget
Weaknesses
- Five SDK advisories from February to June 2026, one listed as critical (arbitrary file read through
TracingMiddleware) - Base traces keep 14 days, and extended retention was capped at 180 days with about a week's notice
- Closed-source platform, self-hosting, BYOC and key roles are Enterprise only
Price $39 / seat-moAuth OAuth or keyx402 nohosted
W&B Weave
B 66.7/100W&B Weave is CoreWeave Forge's tracing and evaluation service for LLM applications and agents. It takes traces from Python and TypeScript SDKs or an OTLP endpoint and exposes them through a REST Service API.
Verdict The Service API at trace.wandb.ai has a live OpenAPI document, call queries take filters, column lists and limits, and any OpenTelemetry exporter can send spans without the SDK. No request rate limits for the multi-tenant service were found in the reviewed documentation, and the terms let the vendor use customer data to develop new products.
Choose it for Teams already on Weights & Biases, or with OpenTelemetry instrumentation, who want traces, evaluations, scorers, datasets and prompts in one place.
Strengths
- A live OpenAPI 3.1 document at
https://trace.wandb.ai/openapi.jsondescribes 127 operations on 114 paths - Two OTLP endpoints accept protobuf spans from any OpenTelemetry exporter, so tracing needs no Weave SDK
/calls/stream_querytakesfilter,query,sort_by,columns,limitandoffset, so an agent can size each response
Weaknesses
- No request rate limits for the multi-tenant Service API were found in the reviewed documentation
- The OpenAPI document lists only 200 and 422 responses, and 51 of 127 operations have no description
- The Master Service Agreement lets W&B use Customer Data to improve its services, develop new products and run AI functions
Price $60 / moAuth API keyx402 nohosted
Prefactor
B 65.6/100Prefactor is a hosted service that records AI agent runs as spans, classifies each run's data risk and stores quality evaluations. Agents and scripts reach it through an HTTP and WebSocket API, TypeScript and Python SDKs and a CLI.
Verdict The public OpenAPI document lists 157 operations, each accepting an idempotency key, and tokens can be read-only or bound to one agent deployment. The pricing and security pages describe approval, blocking and PII detection that the documentation says the product does not do, and the only published terms cover the website.
Choose it for A team that wants an auditable record of agent runs with declared data-risk labels, quality payloads from its own evaluations and a remote stop signal.
Strengths
- Public OpenAPI 3.0 document with 157 operations and an OpenRPC document with 159 WebSocket methods, both fetched without a login
- Every write accepts an
idempotency_key, andPOST /bulkcan be retried whole because a reused key fails only its own item - Tokens are scoped to the account or to one agent deployment, carry a
read_onlyorfull_accessrole and an expiry, and can be suspended or revoked
Weaknesses
- The pricing page lists hold, approve or block and PII detection on every plan, while the docs say a risk threshold blocks nothing and that Prefactor does not detect sensitive data
- No approval operation exists in the OpenAPI or OpenRPC documents. The documented control is a cooperative terminate signal the agent's own code must obey
- The published terms of service cover the website. No service agreement, DPA, sub-processor list or SLA text was found
Price $49 / moAuth API keyx402 nohosted
Respan API + MCP
B 65.5/100Hosted tracing, evals, prompt management and an LLM gateway to 250+ models, from the company formerly called Keywords AI.
Verdict API keys with Read, Write, Gateway and Admin scopes, expiry and spend caps, bound to one project. 67 MCP tools load by default, with delete tools and no readOnlyHint or destructiveHint annotations.
Choose it for Teams that want a gateway and tracing in one integration and an agent that can manage prompts, datasets and evaluators over MCP.
Strengths
- API keys with Read, Write, Gateway and Admin scopes, expiry and spend caps, bound to one project
- OpenAPI 3.1 spec, llms.txt and Markdown for every docs page
Respan-Enabled-Toolsheader trims the MCP server's tool list server-side
Weaknesses
- 67 MCP tools load by default, with delete tools and no
readOnlyHintordestructiveHintannotations - A 13-hour 47-minute gateway incident on 26 September 2026, while the status bars show 100 per cent
- Security fixes on 30 July and 28 August 2026 disclosed only as one-line changelog entries
Price $199 / moAuth OAuth or keyx402 nohosted
LangWatch
B 65.5/100LangWatch is an open-source platform for tracing, evaluating and testing LLM applications and agents, with prompt management, datasets and an AI gateway. It runs hosted or self-hosted, with a REST API, SDKs, a CLI and an MCP server.
Verdict API keys can be limited to read or write per permission category, expire, and be revoked, and the remote MCP server uses OAuth with PKCE. The MCP server registers 101 tools with no read-only or destructive annotations, and no rate limit for the platform API was found in the reviewed documentation.
Choose it for Teams that want tracing, evaluations and simulated-user agent tests in one open-source product, hosted in the EU or self-hosted, and that drive it from a coding assistant through MCP or the CLI.
Strengths
- API keys take read or write per permission category, a project, team or organisation scope and an expiry. Ingestion keys can only write traces
- A public OpenAPI 3.1 document covers 328 operations and is served without a key, with llms.txt and a Markdown twin of every docs page
- Apache 2.0 platform with MIT SDKs and MCP server. Self-hosting has no volume cap, and every outbound call is documented with its off switch
Weaknesses
- The MCP server registers 101 tools, deletes and key creation among them, with no toolsets and no
readOnlyHintordestructiveHintannotations in the source - No request rate limit for the platform API was found in the reviewed documentation. Only two operations in the OpenAPI document declare a 429
- The status page shows Scenarios down for 9 hours 34 minutes on 8 September 2026 and trace processing degraded for 14 hours 44 minutes on 31 July
Price FreemiumAuth OAuth or keyx402 nohosted and local
Pydantic Logfire
B 64.9/100Pydantic Logfire is a hosted observability platform built on OpenTelemetry for traces, logs, metrics, LLM cost tracking, evaluations and prompt management. Agents reach it through a remote MCP server, a SQL query API, a public REST API and SDKs.
Verdict OAuth with PKCE and dynamic client registration, 44 scopes and a public OpenAPI 3.1 document make access easy to limit and to script, and the free Personal plan needs no card. No status page was found, query limits are published as named levels, not numbers, and the hosted MCP server's tool schemas could not be read.
Choose it for Teams that want agent traces alongside application logs and metrics in one OpenTelemetry store, queried by SQL from a coding assistant.
Strengths
- The remote MCP server uses OAuth with PKCE (S256 required) and dynamic client registration, or a Bearer API key with as little as the
project:readscope - A public OpenAPI 3.1 document of 68 paths and 106 operations is served without a key, with llms.txt and a Markdown twin of every docs page
- The Personal plan is free with no card, 10 million records a month and a hard cap at $0. Team and Growth charge $2 per million records beyond that
Weaknesses
- No public status page was found on pydantic.dev, in the docs or in the files published for agents
- Query limits for MCP, read tokens and the public API are published as Low, Standard and High per plan, with no numbers beyond a daily request count on the pricing page
- The hosted MCP server is closed source and its discovery card lists 51 tools, 31 of them creates, updates or deletes. The card says it is maintained by hand and its tool names differ from the docs
Price $49 / moAuth OAuth or keyx402 nohosted
DeepEval
B 64.7/100DeepEval is an open-source Python and TypeScript framework from Confident AI for evaluating LLM applications and agents. Evaluations run locally through a pytest plugin and the deepeval CLI, with optional reporting to the hosted Confident AI platform.
Verdict DeepEval runs evaluations and tracing locally under Apache 2.0 with no account, and writes each test run to JSON or SQLite. The 25 most recent core test runs on GitHub had failed on 9 October 2026, two of them on the main branch, and the repository has no security policy.
Choose it for Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics.
Strengths
- Apache-2.0 framework that runs evaluations and tracing locally with no account. Test runs are written to JSON files or a SQLite database on the owner's disk
- Python 4.2.8 was tagged on 2 October 2026, with 14 Python releases tagged since 9 August and TypeScript 0.9.21 on 30 September
deepeval test runtakes flags for parallel processes, repeats, a result cache, ignoring errors and skipping cases with missing parameters
Weaknesses
- The 25 most recent Py Core Tests runs had failed when read on 9 October 2026, two of them pushes to main. Issue #3372 of 26 September reports the same
- No SECURITY.md, no published advisories and no security.txt on deepeval.com or confident-ai.com. The trust centre is drawn by script and was not read
- Release 4.2.0 reversed the score direction of four safety metrics in a minor version. The changelog marks it as breaking and the code warns at run time
Price $200 / moAuth OAuth or keyx402 nolocal
MLflow Tracing
C 61.2/100Open-source tracing, evaluation and prompt management for LLM applications and agents, part of MLflow, a Linux Foundation project. Owners run the server themselves, and agents read and annotate traces through an experimental MCP server or the mlflow traces CLI.
Verdict Apache-2.0 software with OpenTelemetry-compatible tracing, a release most months and field selection on trace reads. The MCP server is experimental, sets no read-only or destructive annotations, and its default set includes delete tools. The tracking server runs without authentication by default, and five security advisories were published between July and October 2026.
Choose it for Teams that already run MLflow or want Apache-2.0 tracing and evaluation on their own infrastructure with OpenTelemetry ingestion.
Strengths
- Apache-2.0 licence, free to self-host, with nothing to buy from the project
extract_fieldsonsearch_tracesandget_tracereturns only the named fields, withmax_resultsandpage_tokenfor paging- The server accepts OTLP at
/v1/traces, so applications in any OpenTelemetry language can send spans
Weaknesses
- The MCP server is marked experimental in the docs and sets no
readOnlyHintordestructiveHinton any tool - The tracking server has no authentication unless started with
--app-name basic-auth - Five security advisories published between 27 July and 9 October 2026, one a critical unauthenticated remote code execution fixed in 3.17.0
Price Free · OSSAuth OAuth or keyx402 nolocal
Head to head
- Arize Phoenix vs Langfuse API + MCP BB 75.4 vs BB 72.7
- Arize Phoenix vs LangSmith API + MCP BB 75.4 vs BB 71.1
- Arize Phoenix vs W&B Weave BB 75.4 vs B 66.7
- Arize Phoenix vs Prefactor BB 75.4 vs B 65.6
- Langfuse API + MCP vs LangSmith API + MCP BB 72.7 vs BB 71.1
- Langfuse API + MCP vs W&B Weave BB 72.7 vs B 66.7
- Langfuse API + MCP vs Prefactor BB 72.7 vs B 65.6
- LangSmith API + MCP vs W&B Weave BB 71.1 vs B 66.7
- LangSmith API + MCP vs Prefactor BB 71.1 vs B 65.6
- Prefactor vs W&B Weave B 65.6 vs B 66.7
Questions
What are the highest-rated agent tracing, monitoring and evaluation tools for AI agents?
Arize Phoenix has the highest benchmark score of the 15 ranked agent tracing, monitoring and evaluation tools, 75.4 (BB). Langfuse API + MCP is second with 72.7 (BB).
How many agent tracing, monitoring and evaluation tools are agent-ready?
3 of the 15 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which agent tracing, monitoring and evaluation tools accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
Which of these agent tracing, monitoring and evaluation tools is cheapest?
By published paid prices, Braintrust API + MCP, at $0.50 per GB, the lowest of the 3 listings here with a paid price in this unit (free allowances aside). Plans, volume tiers and free allowances change the sum, so check the listing's price table.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 120 head-to-head comparisons · Best tools in every category