Best of · Work & business apps

Best company knowledge search and data catalogues for AI agents

The 10 highest-scoring of 13 company knowledge search and data catalogues on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.

  • 13 ranked
  • 0 agent-ready
  • 5 hosted endpoints
  • Updated 9 October 2026

Top three

Picks by need

Worked out from the scores, prices and facts, so they change when the research does.

Highest score overall

Glean B

B, 69.8/100 on the benchmark.

Also OpenMetadata, B, 66.6/100.

Reliability

OpenMetadata B

84/100 on reliability, against 60 for the overall leader.

Agent ergonomics

OpenMetadata B

85/100 on agent ergonomics, against 80 for the overall leader.

Maintenance & community

OpenMetadata B

89/100 on maintenance & community, against 85 for the overall leader.

Transparency & trust

Collibra B

82/100 on transparency & trust, against 80 for the overall leader.

A hosted MCP endpoint

Onyx B

remote MCP server, nothing to install.

Self-hosting under an open licence

OpenMetadata B

self-hosted, Apache-2 licence.

Also Onyx, self-hosted, MIT licence.

The shortlist

#ToolGradeBest forPriceWhere
1 Glean
Glean Technologies, Inc.
B 69.8 Agents working inside a company that already runs Glean and need answers from its documents, chats, tickets, code and people with each user's permissions applied. Paid local
2 OpenMetadata
Collate, Inc.
B 66.6 Teams running OpenMetadata who want agents to search the catalogue, read lineage and data-quality results, find root causes and create tests or lineage under the user's own permissions. Freemium local
3 Onyx
DanswerAI, Inc. (Onyx, formerly Danswer)
B 65 Teams that want one permission-aware index over Slack, Drive, Confluence, Jira, GitHub and the rest, self-hosted or air-gapped, with agents searching it through MCP. $20 / seat-mo hosted
4 Collibra
Collibra
B 64.6 An organisation that already runs Collibra and wants an agent to look up assets, glossary terms, lineage and data quality, or to write assets and classifications under a governed role. Paid local
5 Marmot
Marmot Data
B 64.4 Teams that want agents to find which table or topic holds something, who owns it, what feeds it and what a business term means, from a catalogue they host. $49 / mo local
6 Dust
Permutation Labs
B 62.8 A team that already works in Dust and wants an outside agent to search its spaces, read documents and call its agents with a user's own access. Freemium hosted
7 Atlan
Atlan
B 62.5 A company already on Atlan that wants agents to find governed tables, owners, lineage, glossary terms and data quality rules, and to run read-only SQL, under the same permissions people have. Paid hosted
8 Alation
Alation, Inc.
C 60.3 A company already on Alation Cloud Service that wants agents to find governed tables, columns, queries, BI reports and lineage, and to run read-only SQL against data products, under an Alation role. Paid local
9 DataHub
Acryl Data, Inc. (DataHub)
C 59.4 Teams already running DataHub who want agents to find tables, trace column lineage, read the SQL people run and see owners and glossary terms before touching data. Freemium hosted and local
10 Cohere North
Cohere Inc.
C 55.1 Regulated companies that want an agent and search platform over their own documents, mail and tickets, run privately or air-gapped and called through a REST or OpenAI-compatible API. Paid local

3 more are ranked in the full table.

How to choose

  1. Source permissions kept intactCheck whether each connected source keeps its own access controls, since an agent that surfaces a restricted document through search has leaked it.
  2. Answers cite their sourceConfirm that answers cite the document or table they came from, so an agent's claim about a metric can be traced back to its lineage.
  3. Owner and lineage per assetCheck whether ownership and lineage come back with each table or metric, so an agent can say who to ask before it acts.
  4. Index lag for new documentsMeasure how long a newly added document takes to become findable, since an agent answering from stale indexes gives wrong answers about current policy.

How the benchmark tests this category. A fixed set of questions about one test company's documents and data (who owns a table, where a metric comes from, what a policy says), asked through each listing's agent interface as users with different permissions. We check whether answers cite the right source, whether a user ever sees what they shouldn't, and how long a new document takes to become findable.

Each one in detail

#1

Glean

B 69.8/100

Enterprise search and AI assistant from Glean Technologies in San Francisco.

Verdict Three public OpenAPI specs (Client, Indexing, Platform) regenerated almost daily, plus llms.txt and Markdown docs. No public price, trial or self-serve signup. Access starts with a demo request.

Choose it for Agents working inside a company that already runs Glean and need answers from its documents, chats, tickets, code and people with each user's permissions applied.

Strengths

  • Three public OpenAPI specs (Client, Indexing, Platform) regenerated almost daily, plus llms.txt and Markdown docs
  • OAuth with dynamic client registration, or Glean-issued tokens with 19 scopes, user-scoped or global, and optional expiry
  • Source-system permissions enforced on every search, chat and document read through the MCP server

Weaknesses

  • No public price, trial or self-serve signup. Access starts with a demo request
  • Eight incidents on status.glean.com between 10 July and 3 September 2026, seven marked major, most on Chat and the Assistant
  • No idempotency keys, and no documented confirmation step for MCP tools that write

Price PaidAuth OAuth or keyx402 nolocal

Full assessment

#2

OpenMetadata

B 66.6/100

Open-source data catalogue for discovery, lineage, data quality and governance, with connectors to external data systems.

Verdict MCP built into every instance at /mcp and on by default since 2.0, with nothing extra to install. The 16 default tools take about 50,000 characters of definitions, 12,285 of them for search_metadata alone.

Choose it for Teams running OpenMetadata who want agents to search the catalogue, read lineage and data-quality results, find root causes and create tests or lineage under the user's own permissions.

Strengths

  • MCP built into every instance at /mcp and on by default since 2.0, with nothing extra to install
  • OAuth 2.0 with PKCE and dynamic client registration through the instance's own SSO or basic auth, 1-hour access tokens and rotating 7-day refresh tokens, plus personal access tokens and bot JWTs
  • Every tool carries readOnlyHint and destructiveHint, and descriptions say when to choose one tool over another and warn where an answer can be silently wrong

Weaknesses

  • The 16 default tools take about 50,000 characters of definitions, 12,285 of them for search_metadata alone
  • Three advisories in 2026, a critical template injection and two leaks of bot JWTs to low-privilege users
  • An open report from 2 October 2026 says get_entity_details returns service connections (host, user, the stored password field) that the REST API masks, and we found no masking step in the MCP read path

Price FreemiumAuth OAuth or keyx402 nolocal

Full assessment

#3

Onyx

B 65/100

Open-source enterprise search and chat platform, formerly Danswer.

Verdict MIT Community Edition, MCP server included, run by Docker Compose, Helm or Terraform, with a two-week Cloud trial that needs no card. GHSA-q62f-rv3h-f822 (critical, CVSS 9.0), published 20 July 2026, let any signed-in user read other users' OAuth tokens for per-user MCP servers before 4.0.0.

Choose it for Teams that want one permission-aware index over Slack, Drive, Confluence, Jira, GitHub and the rest, self-hosted or air-gapped, with agents searching it through MCP.

Strengths

  • MIT Community Edition, MCP server included, run by Docker Compose, Helm or Terraform, with a two-week Cloud trial that needs no card
  • Three read-only MCP tools in 3,360 characters, whose filters return close matches instead of searching unscoped
  • Personal access tokens limited to read:search, with 7, 30 or 365-day expiry, hashed storage and revocation one by one

Weaknesses

  • GHSA-q62f-rv3h-f822 (critical, CVSS 9.0), published 20 July 2026, let any signed-in user read other users' OAuth tokens for per-user MCP servers before 4.0.0
  • Telemetry is on by default and documented as anonymous, while its events carry user IDs and Enterprise builds send the first user's email domain
  • Document search has no result limit or paging, and an unparseable time_cutoff is dropped with only a server log line

Price $20 / seat-moAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Glean

#4

Collibra

B 64.6/100

Collibra is a hosted data governance platform with a catalogue, business glossary, lineage and data quality. Agents use REST and GraphQL APIs, a bulk Import API, a hosted MCP server on each cloud instance and an open-source local one.

Verdict Each REST API has an OpenAPI 3.0 definition, the developer portal publishes llms.txt and Markdown pages, and every MCP tool carries read-only and destructive annotations. Access needs a sales contract, since no price, trial or self-serve sign-up is published, and no general API rate limit was found in the reviewed documentation.

Choose it for An organisation that already runs Collibra and wants an agent to look up assets, glossary terms, lineage and data quality, or to write assets and classifications under a governed role.

Strengths

  • OpenAPI 3.0 definitions for each REST API, printed on every reference page of developer.collibra.com and served by each instance at /docs/index.html
  • The developer portal publishes llms.txt, a Markdown copy of every page and an MCP server for searching the documentation
  • Every tool in the open-source MCP server declares readOnlyHint, destructiveHint and idempotentHint, and the 18 data quality tools stay unregistered until a flag is set

Weaknesses

  • No published price, trial or self-serve sign-up. The pricing address returns 404 and the site links a demo request and a product tour
  • No general request rate limit or Retry-After guidance found. Only the OAuth token endpoint has numbers, a pool of 100 tokens refilled at 30 a minute
  • The Developer Terms of 22 August 2023 forbid publishing benchmarks and calls not driven by bona fide end user requests, except reasonable testing

Price PaidAuth OAuth or keyx402 nolocal

Full assessment · Against #1, Glean

#5

Marmot

B 64.4/100

Open-source data catalogue from Marmot Data Ltd in London, MIT licensed and shipped as one Go binary on Postgres.

Verdict MIT licence, one Go binary on Postgres, with Docker images, a Helm chart and Linux and macOS builds for amd64 and arm64. Pre-1.0 (0.11), and the release notes are generated lists of additions and fixes with no breaking-change section.

Choose it for Teams that want agents to find which table or topic holds something, who owns it, what feeds it and what a business term means, from a catalogue they host.

Strengths

  • MIT licence, one Go binary on Postgres, with Docker images, a Helm chart and Linux and macOS builds for amd64 and arm64
  • Nine MCP tools in 6,962 characters of descriptions, each saying when to use it, with limit and offset paging (default 20, max 100)
  • The three MCP write tools return a preview first, apply only on a second call with confirm set to true, and refuse without assets:manage

Weaknesses

  • Pre-1.0 (0.11), and the release notes are generated lists of additions and fixes with no breaking-change section
  • The MCP docs page lists 3 tools while v0.11.0 registers 9, and none carries readOnlyHint or destructiveHint
  • Telemetry is on by default, and the per-channel lookup counts in its daily report aren't on the telemetry page's list

Price $49 / moAuth OAuth or keyx402 nolocal

Full assessment

#6

Dust

B 62.8/100

Dust is a platform for building AI agents on a company's documents and connected tools, from Permutation Labs in Paris. Outside agents reach it through a REST API, a JavaScript SDK, a CLI and a remote MCP server with OAuth.

Verdict The whole platform is MIT on GitHub, with a public OpenAPI document of 68 operations, llms.txt and a remote MCP server that acts with the signed-in user's access. API calls that run agents draw only on a paid workspace credit pool, and the terms of service could not be read.

Choose it for A team that already works in Dust and wants an outside agent to search its spaces, read documents and call its agents with a user's own access.

Strengths

  • The platform's source is public under the MIT licence at github.com/dust-tt/dust, with commits on the day of the check
  • Public OpenAPI 3.0 document with 68 operations, plus llms.txt and a Markdown copy of each docs page
  • Remote MCP server at https://dust.tt/mcp and https://eu.dust.tt/mcp, with OAuth, dynamic client registration and the signed-in user's own access

Weaknesses

  • Programmatic usage has no free credits. It draws only on the workspace credit pool, which Free seats cannot use
  • Rate limits are published for document upserts and app runs only, and no Retry-After guidance was found in the docs
  • The terms link redirects to a Notion site drawn by script, so the service terms, any SLA and the DPA went unread

Price FreemiumAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Glean

#7

Atlan

B 62.5/100

Hosted data catalogue and metadata platform for finding and managing enterprise data.

Verdict OAuth per user at mcp.atlan.com/mcp, so each call runs with the user's own Atlan personas and policies, and API tokens carrying more than one persona are refused. Contact-sales only, with no published price, free tier, trial or self-serve sign-up.

Choose it for A company already on Atlan that wants agents to find governed tables, owners, lineage, glossary terms and data quality rules, and to run read-only SQL, under the same permissions people have.

Strengths

  • OAuth per user at mcp.atlan.com/mcp, so each call runs with the user's own Atlan personas and policies, and API tokens carrying more than one persona are refused
  • Write tools return a preview and wait for approval, and the SQL tool refuses anything but SELECT, WITH, SHOW, DESCRIBE and EXPLAIN
  • Ten coded MCP errors, each with a category and a recovery step, and search paging of 20 by default and 100 at most, or a count alone

Weaknesses

  • Contact-sales only, with no published price, free tier, trial or self-serve sign-up
  • 39 tools on one endpoint, and the read-only mode that cuts them to 15 is set by Atlan on request
  • status.atlan.com created its four components on 28 September 2026, none for MCP, and its feed holds one incident

Price PaidAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Glean

#8

Alation

C 60.3/100

Alation is a hosted data catalogue for tables, queries, BI reports, lineage and data quality. Agents reach it through REST APIs, a natural-language Aggregated Context API, a Python agent SDK and an MCP server on each cloud instance.

Verdict Each API has an OpenAPI 3.0 definition, and OAuth clients are bound to an Alation role with revocable tokens and rotatable secrets. Access needs a sales contract, since no price, trial or self-serve sign-up is published. The agent SDK and local MCP server have stayed at a release candidate since 14 January 2026.

Choose it for A company already on Alation Cloud Service that wants agents to find governed tables, columns, queries, BI reports and lineage, and to run read-only SQL against data products, under an Alation role.

Strengths

  • OpenAPI 3.0 definitions for each API, shown on every reference page and served by each instance under /openapi/
  • OAuth 2.0 client credentials tied to a system user with one Alation role, with token revocation, secret rotation and offline JWT verification
  • The AI API answers 429 with Retry-After and three Ai-RateLimit headers, and the REST 429 body states the seconds to wait

Weaknesses

  • No published price, trial or self-serve sign-up. The pricing address redirects to a sales form
  • Four incidents between 10 August and 10 September 2026, two of them regional outages longer than an hour
  • alation-ai-agent-sdk and alation-ai-agent-mcp were last published on 14 January 2026 as 1.0.0rc3, and a plain pip install still resolves to 0.12.0

Price PaidAuth OAuth or keyx402 nolocal

Full assessment · Against #1, Glean

#9

DataHub

C 59.4/100

Open-source data catalogue and metadata platform from Acryl Data, trading as DataHub.

Verdict Apache-2.0 platform and MCP server, run with uvx mcp-server-datahub@latest or the acryldata/mcp-server-datahub Docker image against DataHub Core or DataHub Cloud. The eight default tools carry about 26,000 characters of descriptions, and search and get_lineage each repeat the same 3,063-character filter grammar.

Choose it for Teams already running DataHub who want agents to find tables, trace column lineage, read the SQL people run and see owners and glossary terms before touching data.

Strengths

  • Apache-2.0 platform and MCP server, run with uvx mcp-server-datahub@latest or the acryldata/mcp-server-datahub Docker image against DataHub Core or DataHub Cloud
  • Write tools stay off until TOOLS_IS_MUTATION_ENABLED=true, and all 10 read tools carry readOnlyHint
  • One filter string on search and lineage (platform = snowflake AND env = PROD), paging capped at 50, facet-only searches and an 80,000-token response budget

Weaknesses

  • The eight default tools carry about 26,000 characters of descriptions, and search and get_lineage each repeat the same 3,063-character filter grammar
  • Every tool call sends a Mixpanel event by default with the tool name, the client and up to 500 characters of any error message, and no docs page mentions it
  • The MCP server is pre-1.0 (0.7.1), and CHANGELOG.md stops at 0.5.3 with a breaking HTTP change still under Unreleased

Price FreemiumAuth OAuth or keyx402 nohosted and local

Full assessment

#10

Cohere North

C 55.1/100

Cohere North is an enterprise AI agent and search platform from Cohere in Toronto. It indexes connected work apps into Compass for permission-aware search and chat, with a REST API per instance. North 2 was announced on 5 October 2026.

Verdict North has a public OpenAPI 3.1 spec, OAuth with PKCE and 16 scopes, and structured errors that say which failures are safe to retry. Access is the main limitation. Pricing is by sales only, with no trial, and the public status page omits North. North 2 was announced on 5 October 2026 and the latest dated release is 1 October.

Choose it for Regulated companies that want an agent and search platform over their own documents, mail and tickets, run privately or air-gapped and called through a REST or OpenAI-compatible API.

Strengths

  • Public OpenAPI 3.1 spec for the North API (128 operations) and for Compass search, plus llms.txt and Markdown docs
  • OAuth 2.0 with PKCE, 16 scopes, refresh tokens and a revoke endpoint, or token exchange from the company's identity provider
  • Errors carry a stable error_code, an is_retryable flag and retry_after_seconds, and 429 and 503 send Retry-After

Weaknesses

  • No public price, trial or self-serve signup. Access starts with a sales conversation
  • status.cohere.com has no North or Compass component, and no North SLA was found
  • Rate limits are set per customer through Flow Control, which is marked Alpha, and no default numbers are published

Price PaidAuth OAuthx402 nolocal

Full assessment · Against #1, Glean

Head to head

All 63 comparisons in this category

Questions

What are the highest-rated company knowledge search and data catalogues for AI agents?

Glean has the highest benchmark score of the 13 ranked company knowledge search and data catalogues, 69.8 (B). OpenMetadata is second with 66.6 (B).

How many company knowledge search and data catalogues are agent-ready?

0 of the 13 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.

Which company knowledge search and data catalogues accept x402 payments?

None of the ranked listings here accepts x402 for its main call yet.

How is this list ranked?

By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.

How this list is made

The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.

Full ranked table · 63 head-to-head comparisons · Best tools in every category

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.