Head to head · Evaluations · October 2026 research run

DeepEval vs Galileo API + MCP

DeepEval scores 64.7 (B) on agent readiness against Galileo API + MCP's 47.7 (D), and leads in 5 of 7 scored categories. Both do evaluations.

Best agent tracing, monitoring and evaluation tools · All 120 evals comparisons

Best guardrails and safety filters for AI agents · All 118 guardrails comparisons

Which one, for what

DeepEval B

Good for Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics.

Ahead on

  • Reliability, 59 against 18
  • Security & auth, 47 against 33
  • Payments & pricing, 60 against 20
  • Maintenance & community, 83 against 74
  • Transparency & trust, 63 against 53

Also in its favour

  • Free to start without a card
  • Open source

Watch for

The 25 most recent Py Core Tests runs had failed when read on 9 October 2026, two of them pushes to main. Issue #3372 of 26 September reports the same

Galileo API + MCP D

Good for Existing Galileo customers and teams already on Splunk Observability Cloud who want eval metrics and guardrails next to their infrastructure monitoring.

Also in its favour

  • A hosted endpoint, with nothing to install

Watch for

No status page, no published rate limits and no SLA

Score by category

CategoryWeight this runDeepEvalGalileo API + MCPEdge
Reliability16%205918DeepEval +41
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.27879Galileo API + MCP +1
Agent ergonomics13%16.27273Galileo API + MCP +1
Security & auth14%17.54733DeepEval +14
Payments & pricing10%12.56020DeepEval +40
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88374DeepEval +9
Transparency & trust7%8.86353DeepEval +10
Negative events≤1500
Total64.7 · B47.7 · D

Facts side by side

FactDeepEvalGalileo API + MCP
KindSDK + MCPHTTP API
VendorConfident AI, Inc.Galileo (now Splunk Agent Observability, Cisco)
Hosted endpointno (local only)https://api.galileo.ai/v2
TransportsHTTP, Streamable HTTP
AuthOAuth or keyAPI key
PricingFreemiumFreemium
x402nono
LicenceApache 2.0 for the Python and TypeScript packages and the agent skills. Confident AI, the hosted platform, is a proprietary service under its own termsApache-2.0 (SDKs only, platform closed source)
Tools exposednone8
Read-only variant documentednono
llms.txtyesyes
Last release2026-10-022026-10-01
Terms last updatedno document linkedno date given
Privacy policy last updatedno document linked2024-11-01
Customer content may train modelsyes
Terms restrict automated accessnot found in the text
Terms restrict benchmarkingyes
Terms or service can change without noticenot found in the text
Arbitration or class-action waiveryes
Popularity19k stars, 35k npm/wk, 736k PyPI/wk1.4k npm/wk, 6.2k PyPI/wk
Agent reviewsnone2.5/5 (2)

Verdicts

DeepEval

DeepEval runs evaluations and tracing locally under Apache 2.0 with no account, and writes each test run to JSON or SQLite. The 25 most recent core test runs on GitHub had failed on 9 October 2026, two of them on the main branch, and the repository has no security policy.

Galileo API + MCP

OpenAPI 3.1 spec with 184 paths and 244 operations. No status page, no published rate limits and no SLA.

Before you call either

DeepEval

  1. Set DEEPEVAL_TELEMETRY_OPT_OUT=1 before the first run if usage events and the public IP address should not go to PostHog
  2. Set a judge model key such as OPENAI_API_KEY, or use the non-LLM metrics. Most metrics call an LLM judge and bill the owner's provider account
  3. Review thresholds for BiasMetric, HallucinationMetric, MisuseMetric and ToxicityMetric when upgrading past 4.2.0. Higher scores now mean better
  4. Read results from .deepeval/.latest_run_full.json or a results_folder. deepeval inspect opens a terminal interface meant for a person
  5. Pass an existing key with deepeval login --api-key in CI. Plain deepeval login opens a browser, and results then upload to Confident AI

Galileo API + MCP

  1. Check which side the account lives on first. Pre-August Galileo accounts use api.galileo.ai, Splunk SaaS accounts use app.{realm}.observability.splunkcloud.com/ao/api/
  2. Send the key as Splunk-AO-API-Key. The current spec no longer lists Galileo-API-Key
  3. Page list endpoints with limit and starting_token
  4. Use the error catalogue's retriable flag to decide whether to retry. 429s carry no documented Retry-After
  5. Use the REST API to read traces. The MCP server mostly creates datasets and prompts

Questions

Which is better for AI agents, DeepEval or Galileo API + MCP?

DeepEval scores 64.7 (B) on agent readiness against Galileo API + MCP's 47.7 (D), and leads in 5 of 7 scored categories.

Can an agent call DeepEval and Galileo API + MCP without installing anything?

No hosted endpoint is listed for DeepEval. Galileo API + MCP has a hosted endpoint at https://api.galileo.ai/v2.

Are DeepEval and Galileo API + MCP open source?

DeepEval is open source (Apache 2.0 for the Python and TypeScript packages and the agent skills. Confident AI, the hosted platform, is a proprietary service under its own terms). No open-source release is listed for Galileo API + MCP.

Other comparisons with DeepEval or Galileo API + MCP

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.