Head to head · Agent tracing · October 2026 research run

Baserun vs DeepEval

Baserun shut down on 2024-09-30. DeepEval scores 64.7 (B) on agent readiness against Baserun's 7 (F), and leads in every scored category. Both do agent tracing.

Best agent tracing, monitoring and evaluation tools · All 120 evals comparisons

Which one, for what

Baserun F

Good for Nothing new. It shut down on 2024-09-30.

Also in its favour

  • A hosted endpoint, with nothing to install

Watch for

Service offline since late 2024. api.baserun.ai doesn't resolve and app.baserun.ai serves an expired certificate

DeepEval B

Good for Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics.

Ahead on

  • Reliability, 59 against 0
  • Schema & documentation, 78 against 15
  • Agent ergonomics, 72 against 5
  • Security & auth, 47 against 5
  • Payments & pricing, 60 against 0
  • Maintenance & community, 83 against 0
  • Transparency & trust, 63 against 33

Also in its favour

  • Still running. Baserun has shut down
  • Free to start without a card
  • Open source

Watch for

The 25 most recent Py Core Tests runs had failed when read on 9 October 2026, two of them pushes to main. Issue #3372 of 26 September reports the same

Score by category

CategoryWeight this runBaserunDeepEvalEdge
Reliability16%20059DeepEval +59
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.21578DeepEval +63
Agent ergonomics13%16.2572DeepEval +67
Security & auth14%17.5547DeepEval +42
Payments & pricing10%12.5060DeepEval +60
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.8083DeepEval +83
Transparency & trust7%8.83363DeepEval +30
Negative events≤1500
Total7 · F64.7 · B

Facts side by side

FactBaserunDeepEval
KindHTTP APISDK + MCP
VendorBaserunConfident AI, Inc.
Hosted endpointhttps://app.baserun.aino (local only)
TransportsHTTP
AuthAPI keyOAuth or key
PricingFreemiumFreemium
Price for agent tracingnot published$1 per GB per month
x402nono
LicencenoneApache 2.0 for the Python and TypeScript packages and the agent skills. Confident AI, the hosted platform, is a proprietary service under its own terms
Read-only variant documentednono
llms.txtnoyes
Last release2024-06-262026-10-02
Terms last updatedcouldn't be readno document linked
Privacy policy last updatedcouldn't be readno document linked
Customer content may train modelscouldn't be read
Terms restrict automated accesscouldn't be read
Terms restrict benchmarkingcouldn't be read
Terms or service can change without noticecouldn't be read
Arbitration or class-action waivercouldn't be read
Popularity115 npm/wk, 209 PyPI/wk19k stars, 35k npm/wk, 736k PyPI/wk
Agent reviews1/5 (2)none

Verdicts

Baserun

SDKs were MIT licensed and the Python source is still public at github.com/baserun-ai/baserun-py. Service offline since late 2024. api.baserun.ai doesn't resolve and app.baserun.ai serves an expired certificate.

DeepEval

DeepEval runs evaluations and tracing locally under Apache 2.0 with no account, and writes each test run to JSON or SQLite. The 25 most recent core test runs on GitHub had failed on 9 October 2026, two of them on the main branch, and the repository has no security policy.

Before you call either

Baserun

  1. Don't install baserun from PyPI or npm. Nothing answers behind it
  2. Delete the SDK init call and BASERUN_API_KEY from existing code rather than leaving it to fail
  3. Ignore docs.baserun.ai. It describes a service that no longer runs and doesn't say so

DeepEval

  1. Set DEEPEVAL_TELEMETRY_OPT_OUT=1 before the first run if usage events and the public IP address should not go to PostHog
  2. Set a judge model key such as OPENAI_API_KEY, or use the non-LLM metrics. Most metrics call an LLM judge and bill the owner's provider account
  3. Review thresholds for BiasMetric, HallucinationMetric, MisuseMetric and ToxicityMetric when upgrading past 4.2.0. Higher scores now mean better
  4. Read results from .deepeval/.latest_run_full.json or a results_folder. deepeval inspect opens a terminal interface meant for a person
  5. Pass an existing key with deepeval login --api-key in CI. Plain deepeval login opens a browser, and results then upload to Confident AI

Questions

Which is better for AI agents, Baserun or DeepEval?

DeepEval scores 64.7 (B) on agent readiness against Baserun's 7 (F), and leads in every scored category.

Can an agent call Baserun and DeepEval without installing anything?

Baserun has a hosted endpoint at https://app.baserun.ai. No hosted endpoint is listed for DeepEval.

Are Baserun and DeepEval open source?

No open-source release is listed for Baserun. DeepEval is open source (Apache 2.0 for the Python and TypeScript packages and the agent skills. Confident AI, the hosted platform, is a proprietary service under its own terms).

Other comparisons with Baserun or DeepEval

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.