# DeepEval (slim) > DeepEval is an open-source Python and TypeScript framework from Confident AI for evaluating LLM applications and agents. Evaluations run locally through a pytest plugin and the deepeval CLI, with optional reporting to the hosted Confident AI platform. - Full: https://www.anchorterminal.com/tools/deepeval.md (~8,150 tokens) · this version ~1,680 tokens · JSON https://www.anchorterminal.com/tools/deepeval.json · canonical https://www.anchorterminal.com/tools/deepeval - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **B · 64.7/100 · rank #343 of 950 · #9 in Agent observability & evals · not agent-ready · confidence medium** Assessment: DeepEval runs evaluations and tracing locally under Apache 2.0 with no account, and writes each test run to JSON or SQLite. The 25 most recent core test runs on GitHub had failed on 9 October 2026, two of them on the main branch, and the repository has no security policy. ## Facts - Kind: SDK + MCP · vendor: Confident AI, Inc. · category: Agent observability & evals · legal entity: Confident AI, Inc. · provenance 52/100 - Packages: pypi `deepeval`, npm `deepeval` - Auth: OAuth or key · pricing: Freemium · x402: no · licence: Apache 2.0 for the Python and TypeScript packages and the agent skills. Confident AI, the hosted platform, is a proprietary service under its own terms - Probe metrics: not measured yet (probes haven't run) - Surface graded: The DeepEval framework and CLI, run by the owner. Confident AI's hosted REST API and MCP server are separate surfaces and are not graded here - Packages: `deepeval` on PyPI, 4.2.8 (2 October 2026), Python 3.9 or later. `deepeval` on npm, 0.9.21 (30 September 2026). Both Apache-2.0 - CLI: `deepeval test run`, `generate`, `inspect`, `diagnose`, `login`, `logout`, `view`, `gate`, `settings`, `set-local-store`, `set-confident-region`, `set-mode`, `set-eval-mode`, `set-debug`, plus model provider `set-*` commands - Test run flags: `-n` processes, `-r` repeat, `-c` cache, `-i` ignore errors, `-s` skip on missing parameters, `-x` stop at first failure, `-id` identifier, `-m` pytest marker, `-d` display - Local storage: JSON by default (`.deepeval/.latest_run_full.json` and `test_run_.json` in a results folder) or SQLite (`deepeval.db`). Location set with `DEEPEVAL_CACHE_FOLDER` - Credentials: None for local runs. LLM judge keys come from the environment or a dotenv file. `CONFIDENT_API_KEY` is a project key from Confident AI, written to `.env.local` by `deepeval login` - Retries: `DEEPEVAL_RETRY_MAX_ATTEMPTS` defaults to 2 with exponential backoff capped at 5 seconds, for LLM provider calls. Retries can be handed to the provider SDK - Telemetry: PostHog at us.i.posthog.com, on by default. Opt out with `DEEPEVAL_TELEMETRY_OPT_OUT=1` - Agent files: llms.txt on deepeval.com. Skills `deepeval`, `deepeval-tracing` and `deepeval-otel` in the repository, with `.claude-plugin` and `.cursor-plugin` manifests - Hosted option: Confident AI. Free with 5 test runs a week, 2 seats, 1 project and 1 GB-month of trace spans. Starter $200 a month, Team $2,000 a month, then $1 per GB-month. Enterprise by quote - Hosted regions: US by default at api.confident-ai.com, EU at eu.api.confident-ai.com with `deepeval set-confident-region EU`. Other regions and on-premises hosting through sales - Repository: 18.7k stars, 322 open issues and 387 open pull requests on 9 October 2026. 17 workflow files and 332 Python test files - Prices: Confident AI Starter plan $200 per month (plan); Confident AI Team plan $2000 per month (plan); Confident AI trace spans beyond the plan $1 per GB per month - Scores: Reliability 59, Performance pending, Schema & documentation 78, Agent ergonomics 72, Security & auth 47, Payments & pricing 60, Task success pending, Maintenance & community 83, Transparency & trust 63 · total over the 7 assessed categories - Why: Reliability, Scored on the local-software lines, since DeepEval runs where the owner installs it. · Schema & documentation, Read on the framework lines. · Agent ergonomics, Read on the framework lines. · Security & auth, Local runs need no credential. · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, Python 4.2.8 was tagged on 2 October 2026, seven days before the check (30). · Transparency & trust, Apache-2.0 for the whole repository (30). - Sources: 36, open questions: 11, both in the full twin - Capabilities: obs.evals, obs.traces, obs.datasets, obs.prompts - JSON: https://www.anchorterminal.com/api/v1/tools/deepeval.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/deepeval.svg` or a link to https://www.anchorterminal.com/tools/deepeval from a page on confident-ai.com or one of its subdomains, or the README of github.com/confident-ai/deepeval, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Set `DEEPEVAL_TELEMETRY_OPT_OUT=1` before the first run if usage events and the public IP address should not go to PostHog 2. Set a judge model key such as `OPENAI_API_KEY`, or use the non-LLM metrics. Most metrics call an LLM judge and bill the owner's provider account 3. Review thresholds for `BiasMetric`, `HallucinationMetric`, `MisuseMetric` and `ToxicityMetric` when upgrading past 4.2.0. Higher scores now mean better 4. Read results from `.deepeval/.latest_run_full.json` or a `results_folder`. `deepeval inspect` opens a terminal interface meant for a person 5. Pass an existing key with `deepeval login --api-key` in CI. Plain `deepeval login` opens a browser, and results then upload to Confident AI ## Connect ```bash pip install -U deepeval ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/deepeval ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Arize Phoenix | BB | 75.4 | obs.traces, obs.evals, obs.prompts, obs.datasets | https://www.anchorterminal.com/tools/arize-phoenix.min.md | | Langfuse API + MCP | BB | 72.7 | obs.traces, obs.evals, obs.prompts, obs.datasets | https://www.anchorterminal.com/tools/langfuse.min.md | | LangSmith API + MCP | BB | 71.1 | obs.traces, obs.evals, obs.prompts, obs.datasets | https://www.anchorterminal.com/tools/langsmith.min.md | | W&B Weave | B | 66.7 | obs.traces, obs.evals, obs.prompts, obs.datasets | https://www.anchorterminal.com/tools/wandb-weave.min.md | | Respan API + MCP | B | 65.5 | obs.traces, obs.evals, obs.prompts, obs.datasets | https://www.anchorterminal.com/tools/respan.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)