# DeepEval vs LangWatch (slim) > LangWatch and DeepEval score within a point of each other for evaluations, 65.5 and 64.7 out of 100. Prices, MCP, x402, uptime and agent notes side by side. - Full: https://www.anchorterminal.com/compare/deepeval-vs-langwatch.md (~2,750 tokens) · this version ~780 tokens · JSON https://www.anchorterminal.com/compare/deepeval-vs-langwatch.json · canonical https://www.anchorterminal.com/compare/deepeval-vs-langwatch - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 LangWatch and DeepEval score within a point of each other on agent readiness, 65.5 (B) and 64.7 (B). DeepEval leads on reliability, agent ergonomics and payments & pricing. Both do evaluations. - DeepEval: B 64.7, rank #343 of 950 · https://www.anchorterminal.com/tools/deepeval.min.md - LangWatch: B 65.5, rank #318 of 950 · https://www.anchorterminal.com/tools/langwatch.min.md - DeepEval, good for Teams that want evaluations in pytest or a CLI on their own machines, with agent, RAG, multi-turn and MCP metrics. Ahead on Reliability, 59 against 53; Agent ergonomics, 72 against 67; Payments & pricing, 60 against 40. - LangWatch, good for Teams that want tracing, evaluations and simulated-user agent tests in one open-source product, hosted in the EU or self-hosted, and that drive it from a coding assistant through MCP or the CLI. Ahead on Schema & documentation, 87 against 78; Security & auth, 71 against 47; Maintenance & community, 90 against 83; Transparency & trust, 75 against 63. Also A hosted endpoint, with nothing to install; Runs on your own machine. | Category | DeepEval | LangWatch | | --- | --- | --- | | Reliability (16%) | 59 | 53 | | Performance (10%) | pending | pending | | Schema & documentation (13%) | 78 | 87 | | Agent ergonomics (13%) | 72 | 67 | | Security & auth (14%) | 47 | 71 | | Payments & pricing (10%) | 60 | 40 | | Task success (10%) | pending | pending | | Maintenance & community (7%) | 83 | 90 | | Transparency & trust (7%) | 63 | 75 | | Fact (where they differ) | DeepEval | LangWatch | | --- | --- | --- | | Kind | SDK + MCP | HTTP API | | Vendor | Confident AI, Inc. | Reasoning Engine B.V. (LangWatch) | | Hosted endpoint | no (local only) | `https://app.langwatch.ai` | | Transports | | HTTP, Streamable HTTP, SSE (legacy), stdio | | Price for evaluations | not published | $0.0546 per 1M tokens | | Licence | Apache 2.0 for the Python and TypeScript packages and the agent skills. Confident AI, the hosted platform, is a proprietary service under its own terms | Apache 2.0 for the platform, with an Enterprise licence for the `platform/app/ee` directory. The SDKs and the MCP server are MIT | | Tools exposed | none | 101 | | Terms last updated | no document linked | 2026-09-22 | | Privacy policy last updated | no document linked | 2026-09-29 | | Customer content may train models | | not found in the text | | Terms restrict automated access | | yes | | Terms restrict benchmarking | | not found in the text | | Terms or service can change without notice | | yes | | Arbitration or class-action waiver | | not found in the text | | Popularity | 19k stars, 35k npm/wk, 736k PyPI/wk | 4.9k stars, 50k npm/wk, 87k PyPI/wk |