Best of · Developer & infrastructure

Best agent tracing, monitoring and evaluation tools

The 10 highest-scoring of 15 agent tracing, monitoring and evaluation tools on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.

  • 15 ranked
  • 3 agent-ready
  • 12 hosted endpoints
  • Updated 9 October 2026

Top three

Picks by need

Worked out from the scores, prices and facts, so they change when the research does.

Highest score overall

Arize Phoenix BB

BB, 75.4/100 on the benchmark.

Also Langfuse API + MCP, BB, 72.7/100.

Reliability

LangSmith API + MCP BB

81/100 on reliability, against 77 for the overall leader.

Schema & documentation

Langfuse API + MCP BB

93/100 on schema & documentation, against 88 for the overall leader.

Security & auth

Pydantic Logfire B

80/100 on security & auth, against 56 for the overall leader.

Maintenance & community

LangWatch B

90/100 on maintenance & community, against 84 for the overall leader.

Transparency & trust

Langfuse API + MCP BB

85/100 on transparency & trust, against 73 for the overall leader.

Lowest paid price per GB

Braintrust API + MCP C

$0.50 per GB, the lowest of the 3 listings here with a paid price in this unit (free allowances aside).

Also Laminar API + MCP, $1.50 per GB.

A hosted MCP endpoint

Langfuse API + MCP BB

remote MCP server, nothing to install.

The shortlist

#ToolGradeBest forPriceWhere
1 Arize Phoenix
Arize AI
BB 75.4 Teams that want tracing and evals on their own hardware, or air-gapped, with an MCP endpoint an agent can sign into. Free · OSS local
2 Langfuse API + MCP
Langfuse (ClickHouse)
BB 72.7 Teams that want open source they can self-host or a predictable cloud bill, across tracing, evals, prompts and datasets. $29 / mo hosted
3 LangSmith API + MCP
LangChain
BB 71.1 Teams on LangChain or LangGraph, and for any team that wants mature hosted evals with clear limits and deprecation rules. $39 / seat-mo hosted
4 W&B Weave
Weights & Biases (CoreWeave)
B 66.7 Teams already on Weights & Biases, or with OpenTelemetry instrumentation, who want traces, evaluations, scorers, datasets and prompts in one place. $60 / mo hosted
5 Prefactor
Prefactor Pty Ltd
B 65.6 A team that wants an auditable record of agent runs with declared data-risk labels, quality payloads from its own evaluations and a remote stop signal. $49 / mo hosted
6 Respan API + MCP
Respan (formerly Keywords AI)
B 65.5 Teams that want a gateway and tracing in one integration and an agent that can manage prompts, datasets and evaluators over MCP. $199 / mo hosted
7 LangWatch
Reasoning Engine B.V. (LangWatch)
B 65.5 Teams that want tracing, evaluations and simulated-user agent tests in one open-source product, hosted in the EU or self-hosted, and that drive it from a coding assistant through MCP or the CLI. Freemium hosted and local
8 Pydantic Logfire
Pydantic Services Inc.
B 64.9 Teams that want agent traces alongside application logs and metrics in one OpenTelemetry store, queried by SQL from a coding assistant. $49 / mo hosted
9 DeepEval
Confident AI, Inc.
B 64.7 Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics. $200 / mo local
10 MLflow Tracing
MLflow Project (LF Projects, LLC)
C 61.2 Teams that already run MLflow or want Apache-2.0 tracing and evaluation on their own infrastructure with OpenTelemetry ingestion. Free · OSS local

5 more are ranked in the full table.

How to choose

  1. Steps and retries in tracesCheck whether tool calls, model calls and retries each appear as a separate span, since a merged retry hides the failure the agent is trying to find.
  2. Finding a failed tool callCheck whether a failed tool call can be found by filtering on status or error, since scanning every run by hand slows an agent's debugging down.
  3. Reproducible evaluation runsCheck whether an evaluation can be rerun against a fixed dataset and prompt version, because a score that cannot be reproduced cannot show whether a change helped.
  4. Trace data location and retentionCheck the data location, retention period and ownership of traces, since runs can hold prompts, tool outputs and customer data that the operator stays responsible for.

How the benchmark tests this category. The same agent run instrumented in every tool, including a failed tool call and a retry. We check which steps appear in the trace, whether the failure can be found quickly, whether an evaluation can be reproduced, how long setup takes and what it costs.

Each one in detail

#1

Arize Phoenix

BB 75.4/100

Self-hosted tracing, evaluation, datasets, experiments and prompt management built on OpenTelemetry and OpenInference.

Verdict Free and self-hosted with no feature gating, from pip install to a Helm chart. Auth is off by default and the default admin password is admin.

Choose it for Teams that want tracing and evals on their own hardware, or air-gapped, with an MCP endpoint an agent can sign into.

Strengths

  • Free and self-hosted with no feature gating, from pip install to a Helm chart
  • MCP endpoint generated from the OpenAPI, five code-mode tools by default, OAuth 2.1 with PKCE and audience-bound tokens
  • Tool annotations derived from HTTP verbs, so a client can auto-approve reads and confirm writes

Weaknesses

  • Auth is off by default and the default admin password is admin
  • Remote MCP endpoint still labelled beta in the docs
  • Elastic License 2.0 isn't OSI open source and forbids running Phoenix as a managed service

Price Free · OSSAuth OAuth or keyx402 nolocal

Full assessment

#2

Langfuse API + MCP

BB 72.7/100

Open-source tracing, evaluation, prompt management and datasets for LLM apps and agents, built on OpenTelemetry.

Verdict MIT core with no usage limits when self-hosted, and self-hosted telemetry documented with an off switch. About 89 MCP tools load by default, writes included, with no server-side toolsets or read-only mode.

Choose it for Teams that want open source they can self-host or a predictable cloud bill, across tracing, evals, prompts and datasets.

Strengths

  • MIT core with no usage limits when self-hosted, and self-hosted telemetry documented with an off switch
  • Rate limits per bucket and plan, with 429 and Retry-After documented
  • Per-unit cloud pricing at $8 per 100,000 units, and a free Hobby plan with 50,000 units and no card

Weaknesses

  • About 89 MCP tools load by default, writes included, with no server-side toolsets or read-only mode
  • 12 degraded incidents on the status page between 3 July and 30 September, mostly ingestion delays
  • No uptime SLA found, only support response targets

Price $29 / moAuth API keyx402 nohosted

Full assessment · Against #1, Arize Phoenix

#3

LangSmith API + MCP

BB 71.1/100

LangChain's hosted tracing, evaluation, prompt hub and dataset platform, framework-agnostic with Python and TypeScript SDKs and OpenTelemetry ingest.

Verdict Rate limits published per endpoint and per plan, with each kind of 429 explained. Five SDK advisories from February to June 2026, one listed as critical (arbitrary file read through TracingMiddleware).

Choose it for Teams on LangChain or LangGraph, and for any team that wants mature hosted evals with clear limits and deprecation rules.

Strengths

  • Rate limits published per endpoint and per plan, with each kind of 429 explained
  • Written deprecation policy, six months on cloud, with Deprecation and Sunset headers on deprecated endpoints
  • Remote MCP with OAuth 2.1, read-only in practice, and fetch_runs paging by a 25,000-character budget

Weaknesses

  • Five SDK advisories from February to June 2026, one listed as critical (arbitrary file read through TracingMiddleware)
  • Base traces keep 14 days, and extended retention was capped at 180 days with about a week's notice
  • Closed-source platform, self-hosting, BYOC and key roles are Enterprise only

Price $39 / seat-moAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Arize Phoenix

#4

W&B Weave

B 66.7/100

W&B Weave is CoreWeave Forge's tracing and evaluation service for LLM applications and agents. It takes traces from Python and TypeScript SDKs or an OTLP endpoint and exposes them through a REST Service API.

Verdict The Service API at trace.wandb.ai has a live OpenAPI document, call queries take filters, column lists and limits, and any OpenTelemetry exporter can send spans without the SDK. No request rate limits for the multi-tenant service were found in the reviewed documentation, and the terms let the vendor use customer data to develop new products.

Choose it for Teams already on Weights & Biases, or with OpenTelemetry instrumentation, who want traces, evaluations, scorers, datasets and prompts in one place.

Strengths

  • A live OpenAPI 3.1 document at https://trace.wandb.ai/openapi.json describes 127 operations on 114 paths
  • Two OTLP endpoints accept protobuf spans from any OpenTelemetry exporter, so tracing needs no Weave SDK
  • /calls/stream_query takes filter, query, sort_by, columns, limit and offset, so an agent can size each response

Weaknesses

  • No request rate limits for the multi-tenant Service API were found in the reviewed documentation
  • The OpenAPI document lists only 200 and 422 responses, and 51 of 127 operations have no description
  • The Master Service Agreement lets W&B use Customer Data to improve its services, develop new products and run AI functions

Price $60 / moAuth API keyx402 nohosted

Full assessment · Against #1, Arize Phoenix

#5

Prefactor

B 65.6/100

Prefactor is a hosted service that records AI agent runs as spans, classifies each run's data risk and stores quality evaluations. Agents and scripts reach it through an HTTP and WebSocket API, TypeScript and Python SDKs and a CLI.

Verdict The public OpenAPI document lists 157 operations, each accepting an idempotency key, and tokens can be read-only or bound to one agent deployment. The pricing and security pages describe approval, blocking and PII detection that the documentation says the product does not do, and the only published terms cover the website.

Choose it for A team that wants an auditable record of agent runs with declared data-risk labels, quality payloads from its own evaluations and a remote stop signal.

Strengths

  • Public OpenAPI 3.0 document with 157 operations and an OpenRPC document with 159 WebSocket methods, both fetched without a login
  • Every write accepts an idempotency_key, and POST /bulk can be retried whole because a reused key fails only its own item
  • Tokens are scoped to the account or to one agent deployment, carry a read_only or full_access role and an expiry, and can be suspended or revoked

Weaknesses

  • The pricing page lists hold, approve or block and PII detection on every plan, while the docs say a risk threshold blocks nothing and that Prefactor does not detect sensitive data
  • No approval operation exists in the OpenAPI or OpenRPC documents. The documented control is a cooperative terminate signal the agent's own code must obey
  • The published terms of service cover the website. No service agreement, DPA, sub-processor list or SLA text was found

Price $49 / moAuth API keyx402 nohosted

Full assessment · Against #1, Arize Phoenix

#6

Respan API + MCP

B 65.5/100

Hosted tracing, evals, prompt management and an LLM gateway to 250+ models, from the company formerly called Keywords AI.

Verdict API keys with Read, Write, Gateway and Admin scopes, expiry and spend caps, bound to one project. 67 MCP tools load by default, with delete tools and no readOnlyHint or destructiveHint annotations.

Choose it for Teams that want a gateway and tracing in one integration and an agent that can manage prompts, datasets and evaluators over MCP.

Strengths

  • API keys with Read, Write, Gateway and Admin scopes, expiry and spend caps, bound to one project
  • OpenAPI 3.1 spec, llms.txt and Markdown for every docs page
  • Respan-Enabled-Tools header trims the MCP server's tool list server-side

Weaknesses

  • 67 MCP tools load by default, with delete tools and no readOnlyHint or destructiveHint annotations
  • A 13-hour 47-minute gateway incident on 26 September 2026, while the status bars show 100 per cent
  • Security fixes on 30 July and 28 August 2026 disclosed only as one-line changelog entries

Price $199 / moAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Arize Phoenix

#7

LangWatch

B 65.5/100

LangWatch is an open-source platform for tracing, evaluating and testing LLM applications and agents, with prompt management, datasets and an AI gateway. It runs hosted or self-hosted, with a REST API, SDKs, a CLI and an MCP server.

Verdict API keys can be limited to read or write per permission category, expire, and be revoked, and the remote MCP server uses OAuth with PKCE. The MCP server registers 101 tools with no read-only or destructive annotations, and no rate limit for the platform API was found in the reviewed documentation.

Choose it for Teams that want tracing, evaluations and simulated-user agent tests in one open-source product, hosted in the EU or self-hosted, and that drive it from a coding assistant through MCP or the CLI.

Strengths

  • API keys take read or write per permission category, a project, team or organisation scope and an expiry. Ingestion keys can only write traces
  • A public OpenAPI 3.1 document covers 328 operations and is served without a key, with llms.txt and a Markdown twin of every docs page
  • Apache 2.0 platform with MIT SDKs and MCP server. Self-hosting has no volume cap, and every outbound call is documented with its off switch

Weaknesses

  • The MCP server registers 101 tools, deletes and key creation among them, with no toolsets and no readOnlyHint or destructiveHint annotations in the source
  • No request rate limit for the platform API was found in the reviewed documentation. Only two operations in the OpenAPI document declare a 429
  • The status page shows Scenarios down for 9 hours 34 minutes on 8 September 2026 and trace processing degraded for 14 hours 44 minutes on 31 July

Price FreemiumAuth OAuth or keyx402 nohosted and local

Full assessment · Against #1, Arize Phoenix

#8

Pydantic Logfire

B 64.9/100

Pydantic Logfire is a hosted observability platform built on OpenTelemetry for traces, logs, metrics, LLM cost tracking, evaluations and prompt management. Agents reach it through a remote MCP server, a SQL query API, a public REST API and SDKs.

Verdict OAuth with PKCE and dynamic client registration, 44 scopes and a public OpenAPI 3.1 document make access easy to limit and to script, and the free Personal plan needs no card. No status page was found, query limits are published as named levels, not numbers, and the hosted MCP server's tool schemas could not be read.

Choose it for Teams that want agent traces alongside application logs and metrics in one OpenTelemetry store, queried by SQL from a coding assistant.

Strengths

  • The remote MCP server uses OAuth with PKCE (S256 required) and dynamic client registration, or a Bearer API key with as little as the project:read scope
  • A public OpenAPI 3.1 document of 68 paths and 106 operations is served without a key, with llms.txt and a Markdown twin of every docs page
  • The Personal plan is free with no card, 10 million records a month and a hard cap at $0. Team and Growth charge $2 per million records beyond that

Weaknesses

  • No public status page was found on pydantic.dev, in the docs or in the files published for agents
  • Query limits for MCP, read tokens and the public API are published as Low, Standard and High per plan, with no numbers beyond a daily request count on the pricing page
  • The hosted MCP server is closed source and its discovery card lists 51 tools, 31 of them creates, updates or deletes. The card says it is maintained by hand and its tool names differ from the docs

Price $49 / moAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Arize Phoenix

#9

DeepEval

B 64.7/100

DeepEval is an open-source Python and TypeScript framework from Confident AI for evaluating LLM applications and agents. Evaluations run locally through a pytest plugin and the deepeval CLI, with optional reporting to the hosted Confident AI platform.

Verdict DeepEval runs evaluations and tracing locally under Apache 2.0 with no account, and writes each test run to JSON or SQLite. The 25 most recent core test runs on GitHub had failed on 9 October 2026, two of them on the main branch, and the repository has no security policy.

Choose it for Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics.

Strengths

  • Apache-2.0 framework that runs evaluations and tracing locally with no account. Test runs are written to JSON files or a SQLite database on the owner's disk
  • Python 4.2.8 was tagged on 2 October 2026, with 14 Python releases tagged since 9 August and TypeScript 0.9.21 on 30 September
  • deepeval test run takes flags for parallel processes, repeats, a result cache, ignoring errors and skipping cases with missing parameters

Weaknesses

  • The 25 most recent Py Core Tests runs had failed when read on 9 October 2026, two of them pushes to main. Issue #3372 of 26 September reports the same
  • No SECURITY.md, no published advisories and no security.txt on deepeval.com or confident-ai.com. The trust centre is drawn by script and was not read
  • Release 4.2.0 reversed the score direction of four safety metrics in a minor version. The changelog marks it as breaking and the code warns at run time

Price $200 / moAuth OAuth or keyx402 nolocal

Full assessment · Against #1, Arize Phoenix

#10

MLflow Tracing

C 61.2/100

Open-source tracing, evaluation and prompt management for LLM applications and agents, part of MLflow, a Linux Foundation project. Owners run the server themselves, and agents read and annotate traces through an experimental MCP server or the mlflow traces CLI.

Verdict Apache-2.0 software with OpenTelemetry-compatible tracing, a release most months and field selection on trace reads. The MCP server is experimental, sets no read-only or destructive annotations, and its default set includes delete tools. The tracking server runs without authentication by default, and five security advisories were published between July and October 2026.

Choose it for Teams that already run MLflow or want Apache-2.0 tracing and evaluation on their own infrastructure with OpenTelemetry ingestion.

Strengths

  • Apache-2.0 licence, free to self-host, with nothing to buy from the project
  • extract_fields on search_traces and get_trace returns only the named fields, with max_results and page_token for paging
  • The server accepts OTLP at /v1/traces, so applications in any OpenTelemetry language can send spans

Weaknesses

  • The MCP server is marked experimental in the docs and sets no readOnlyHint or destructiveHint on any tool
  • The tracking server has no authentication unless started with --app-name basic-auth
  • Five security advisories published between 27 July and 9 October 2026, one a critical unauthenticated remote code execution fixed in 3.17.0

Price Free · OSSAuth OAuth or keyx402 nolocal

Full assessment · Against #1, Arize Phoenix

Head to head

All 120 comparisons in this category

Questions

What are the highest-rated agent tracing, monitoring and evaluation tools for AI agents?

Arize Phoenix has the highest benchmark score of the 15 ranked agent tracing, monitoring and evaluation tools, 75.4 (BB). Langfuse API + MCP is second with 72.7 (BB).

How many agent tracing, monitoring and evaluation tools are agent-ready?

3 of the 15 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.

Which agent tracing, monitoring and evaluation tools accept x402 payments?

None of the ranked listings here accepts x402 for its main call yet.

Which of these agent tracing, monitoring and evaluation tools is cheapest?

By published paid prices, Braintrust API + MCP, at $0.50 per GB, the lowest of the 3 listings here with a paid price in this unit (free allowances aside). Plans, volume tiers and free allowances change the sum, so check the listing's price table.

How is this list ranked?

By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.

How this list is made

The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.

Full ranked table · 120 head-to-head comparisons · Best tools in every category

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.