Best of · Models & inference

Best embedding and reranking APIs for AI agents

All 10 ranked embedding and reranking APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.

  • 10 ranked
  • 4 agent-ready
  • 9 hosted endpoints
  • Updated 8 October 2026

Top three

Picks by need

Worked out from the scores, prices and facts, so they change when the research does.

Highest score overall

Amazon Nova Multimodal Embeddings BB

BB, 75/100 on the benchmark.

Also OpenAI embeddings, BB, 73.2/100.

Schema & documentation

Cohere Embed and Rerank BB

92/100 on schema & documentation, against 76 for the overall leader.

Agent ergonomics

Voyage AI embeddings and rerankers C

98/100 on agent ergonomics, against 78 for the overall leader.

Security & auth

OpenAI embeddings BB

95/100 on security & auth, against 91 for the overall leader.

Maintenance & community

Cohere Embed and Rerank BB

90/100 on maintenance & community, against 50 for the overall leader.

Transparency & trust

OpenAI embeddings BB

85/100 on transparency & trust, against 79 for the overall leader.

A hosted MCP endpoint

Jina Embeddings and Reranker C

remote MCP server, nothing to install.

Self-hosting under an open licence

NVIDIA NeMo Retriever Embedding and Reranking NIMs C

self-hosted, Proprietary containers under the NVIDIA Software Licence Agreement and Product-Specific Terms for AI Products licence.

The shortlist

#ToolGradeBest forPriceWhere
1 Amazon Nova Multimodal Embeddings
Amazon Web Services
BB 75 Suited to mixed-media retrieval for teams already on AWS, especially video and audio archives processed through S3. Pay per use hosted
2 OpenAI embeddings
OpenAI
BB 73.2 An agent already on OpenAI that needs cheap general-purpose text retrieval with a small index. Pay per use hosted
3 Cohere Embed and Rerank
Cohere
BB 72.5 Best when reranking is the job, or for long multilingual documents and image-heavy material where a 128K embedding context helps, with a cheaper Fast model for queries against a Pro index. $2 / 1k req hosted
4 Gemini Embedding
Google
BB 70.6 Multimodal corpora, especially video and audio, and for agents already on Google Cloud. Freemium hosted
5 Jina Embeddings and Reranker
Jina AI (Elastic)
C 61 Reranking large candidate sets and multimodal corpora with audio or video. Freemium hosted
6 NVIDIA NeMo Retriever Embedding and Reranking NIMs
NVIDIA
C 61 Teams that already run NVIDIA GPUs and need embedding and reranking inside their own network, including page-image retrieval with the VL models. Freemium local
7 Voyage AI embeddings and rerankers
Voyage AI (MongoDB)
C 58.8 Retrieval quality across domains, code and long documents, with a reranker from the same key. Freemium hosted
8 Mistral Embed and Codestral Embed
Mistral AI
C 57.9 EU data residency, Mistral-only stacks and code retrieval with small vectors. Freemium hosted
9 Nomic Embed
Nomic, Inc.
D 49.2 Teams that want a hosted endpoint for an open-weight model they can also run themselves, with the same vectors either way. Freemium hosted
10 ZeroEntropy zerank and zembed
ZeroEntropy
F 13.7 Only as open weights for teams that can self-host a reranker or embedding model. Pay per use hosted

How to choose

  1. Vector dimensions and sizeCheck the vector dimensions and whether they can be shortened, because the index stores and scans every dimension for every query.
  2. Query and document purposesCheck whether queries and documents can be embedded with different purpose settings, since the setting can change which results rank highest.
  3. Languages and input typesConfirm which languages and modalities the model supports, including images and code, because retrieval quality can fall outside those.
  4. Maximum input lengthCheck the maximum input length and whether long documents are segmented for you, because a truncated chunk silently loses the text that answers a query.

How the benchmark tests this category. The same corpus and queries embedded with each model, then the top results reranked. We check retrieval quality against labelled answers, latency and the cost per million tokens.

Each one in detail

#1

Amazon Nova Multimodal Embeddings

BB 75/100

Amazon Nova Multimodal Embeddings is an AWS model on Amazon Bedrock that turns text, images, document images, video and audio into vectors in one space, at 256, 384, 1024 or 3072 dimensions, through synchronous and asynchronous calls.

Verdict One model embeds text, images, document images, video and audio into a shared space, with nine documented purpose settings and published per-unit prices. It runs in US East (N. Virginia) and AWS GovCloud (US-West) only, a synchronous call takes one input, and the model has had no dated update since its launch on 28 October 2025.

Choose it for Suited to mixed-media retrieval for teams already on AWS, especially video and audio archives processed through S3.

Strengths

  • Text, images, document images, video and audio share one vector space, with four output sizes from 256 to 3072
  • embeddingPurpose has nine documented values, with separate settings for indexing and for each retrieval type
  • Published quotas of 2,000 requests a minute and 30 concurrent asynchronous jobs per Region

Weaknesses

  • In-Region inference in us-east-1 and us-gov-west-1 only, with no cross-Region inference profile
  • A synchronous request embeds one item, with at most 8,192 characters of inline text or 30 seconds of audio or video
  • The Bedrock model card marks Invoke as unsupported while the Nova guide documents InvokeModel for synchronous calls

Price Pay per useAuth API keyx402 nohosted

Full assessment

#2

OpenAI embeddings

BB 73.2/100

OpenAI's text embedding API, with adjustable output dimensions for search and retrieval applications.

Verdict text-embedding-3-small at $0.02 per million tokens, $0.01 through the Batch API. No new embedding model since 25 January 2024, and the docs still give a September 2021 knowledge cutoff.

Choose it for An agent already on OpenAI that needs cheap general-purpose text retrieval with a small index.

Strengths

  • text-embedding-3-small at $0.02 per million tokens, $0.01 through the Batch API
  • Restricted project keys are set per endpoint, so an agent's key can be cut down to read and model calls
  • Up to 2,048 inputs and 300,000 tokens in one request

Weaknesses

  • No new embedding model since 25 January 2024, and the docs still give a September 2021 knowledge cutoff
  • Text only, 8,192 tokens an input, and no reranker
  • Over-long inputs fail rather than being truncated, and output is float or base64 only

Price Pay per useAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Nova Multimodal Embeddings

#3

Cohere Embed and Rerank

BB 72.5/100

Cohere's Embed API turns text, images and mixed text-and-image inputs such as PDF pages into vectors with Embed 5 Pro and Fast, and its Rerank API reorders search results with Rerank 4.

Verdict Embed 5 Pro and Fast share one embedding space with 128K context and compressed outputs, and embed and rerank prices are public. Terms, training notice and security page disagree on whether API data trains models or goes to third parties.

Choose it for Best when reranking is the job, or for long multilingual documents and image-heavy material where a 128K embedding context helps, with a cheaper Fast model for queries against a Pro index.

Strengths

  • Embed 5 Pro and Fast share one embedding space, so a Pro index answers Fast queries
  • 128K context with six output sizes and int8, binary and base64 output
  • Rerank 4 Pro and Fast with 32K context, top_n and published per-search prices

Weaknesses

  • Terms, training notice and security page disagree on whether API data trains models or goes to third parties
  • A Google Cloud outage degraded embed and rerank for about four hours on 1 September 2026, and Embed 5 isn't yet a status component
  • 96 inputs a call, and input_type is required

Price $2 / 1k reqAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Nova Multimodal Embeddings

#4

Gemini Embedding

BB 70.6/100

gemini-embedding-2, Google's multimodal embedding model, takes text, images, video, audio and PDFs into one 3072-dimension space (truncatable to 128) at 8,192 input tokens in 100+ languages.

Verdict Text, images, video, audio and PDFs interleaved in one request and one vector space. $0.20 per million text tokens, against $0.02 for OpenAI's small model.

Choose it for Multimodal corpora, especially video and audio, and for agents already on Google Cloud.

Strengths

  • Text, images, video, audio and PDFs interleaved in one request and one vector space
  • Any output size from 128 to 3072, with truncated vectors returned normalised
  • Batch API at half the standard embedding price

Weaknesses

  • $0.20 per million text tokens, against $0.02 for OpenAI's small model
  • 8,192 input tokens and float output only
  • No reranker on the Gemini API

Price FreemiumAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Nova Multimodal Embeddings

#5

Jina Embeddings and Reranker

C 61/100

jina-embeddings-v5 in text and omni (text, image, audio, video, PDF) variants at up to 32,768 tokens, plus the jina-reranker-v3.5 at 131,072 tokens a call.

Verdict jina-reranker-v3.5 (20 July 2026) with a 131,072-token window and no document cap. No price per token in any currency on the public pages.

Choose it for Reranking large candidate sets and multimodal corpora with audio or video.

Strengths

  • jina-reranker-v3.5 (20 July 2026) with a 131,072-token window and no document cap
  • v5-omni embeds text, images, audio, video and PDFs into one space
  • OpenAPI 3.1 file with enums for model, task and embedding_type, and error responses from 400 to 504

Weaknesses

  • No price per token in any currency on the public pages
  • One prepaid balance shared with Reader and Search, so a scraping job can drain the embedding budget
  • 26 automated incidents on the status feed from 15 September to 1 October 2026, and no status component for v5-omni or reranker v3.5

Price FreemiumAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Nova Multimodal Embeddings

#6

NVIDIA NeMo Retriever Embedding and Reranking NIMs

C 61/100

NVIDIA's NeMo Retriever Embedding and Reranking NIMs are GPU containers that run text and image embedding models and rerankers behind a local REST API, with /v1/embeddings in the OpenAI shape and /v1/ranking.

Verdict Self-hosted containers with OpenAPI 3.1 files, typed request fields, five embedding output types and a dated end-of-life list. The API has no authentication or rate limiting of its own, production use needs an NVIDIA AI Enterprise licence at $4,500 a GPU a year, and the release notes carry no dates.

Choose it for Teams that already run NVIDIA GPUs and need embedding and reranking inside their own network, including page-image retrieval with the VL models.

Strengths

  • OpenAPI 3.1 files for both services, with enums for input_type, modality, embedding_type and truncate and no extra properties allowed
  • /v1/embeddings follows the OpenAI shape, and a -query or -passage model suffix replaces input_type for OpenAI clients
  • Output can be shrunk with dimensions from 128 to 2048 on three models, or with int8, uint8, binary and ubinary types

Weaknesses

  • The API has no authentication and no rate limiting. The security page leaves both to a proxy the deployer runs
  • Production use needs an NVIDIA AI Enterprise licence, $4,500 a GPU a year or $1 a GPU-hour on cloud marketplaces, plus the GPU
  • Release notes carry no dates, and environment variable names changed between 2.0 and 2.3 without a note in the 2.2 or 2.3 notes we read

Price FreemiumAuth Nonex402 nolocal

Full assessment · Against #1, Amazon Nova Multimodal Embeddings

#7

Voyage AI embeddings and rerankers

C 58.8/100

Embedding and reranking models for text, code and multimodal retrieval from MongoDB-owned Voyage AI.

Verdict 200 million free tokens per current model, then $0.02 to $0.12 per million. Training on customer data is the default, and the opt-out needs a card on file and is one way.

Choose it for Retrieval quality across domains, code and long documents, with a reranker from the same key.

Strengths

  • 200 million free tokens per current model, then $0.02 to $0.12 per million
  • Domain models for code, finance and law, a multimodal model and contextualised chunk embeddings
  • Output in float, int8, uint8, binary or ubinary at 256 to 2048 dimensions, per request

Weaknesses

  • Training on customer data is the default, and the opt-out needs a card on file and is one way
  • No security.txt, and no status page linked or reachable
  • Rate-limit tiers only begin once a payment method is added

Price FreemiumAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Nova Multimodal Embeddings

#8

Mistral Embed and Codestral Embed

C 57.9/100

Mistral's API for generating text and code embeddings.

Verdict EU and US regional endpoints and a French legal entity. Embedding API uptime of 94.36 per cent over 90 days on Mistral's status page, with incidents on 12 and 27 August 2026.

Choose it for EU data residency, Mistral-only stacks and code retrieval with small vectors.

Strengths

  • EU and US regional endpoints and a French legal entity
  • codestral-embed with up to 3072 dimensions, first-n truncation and int8 or binary output
  • OpenAPI document and llms.txt for the whole API

Weaknesses

  • Embedding API uptime of 94.36 per cent over 90 days on Mistral's status page, with incidents on 12 and 27 August 2026
  • 8k context on both models, and text or code only
  • mistral-embed dates from December 2023 with fixed 1024-dimension float output, and nothing new since May 2025

Price FreemiumAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Nova Multimodal Embeddings

#9

Nomic Embed

D 49.2/100

Nomic's hosted embedding endpoints on the Atlas API turn text and images into vectors with the open-weight Nomic Embed models. Agents call them over HTTP with an API key or through the Python and TypeScript clients.

Verdict The text models have Apache-2.0 weights and a public OpenAPI 3.1 contract, so vectors made through the hosted endpoint can be reproduced locally. Nomic's current site and documentation index describe a construction-industry product, no rendered public page prices the endpoint, and no published terms or status component name it.

Choose it for Teams that want a hosted endpoint for an open-weight model they can also run themselves, with the same vectors either way.

Strengths

  • Weights for nomic-embed-text-v1, v1.5, v2-moe, nomic-embed-code and nomic-embed-vision-v1.5 are Apache-2.0 on Hugging Face
  • Public OpenAPI 3.1 document for the Atlas API (v0.57.0) with typed request and response schemas for both embedding endpoints
  • nomic-embed-text-v1.5 accepts a dimensionality from 64 to 768, and inputs up to 8,192 tokens per text

Weaknesses

  • docs.nomic.ai/llms.txt and www.nomic.ai now describe a product for architecture, engineering and construction firms, and the documentation index no longer lists the embedding pages
  • No rendered public page states a price for the endpoint. The $1 per 10M tokens figure comes from the Atlas web app's script
  • status.nomic.ai has no component for api-atlas.nomic.ai, and no SLA was found

Price FreemiumAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Nova Multimodal Embeddings

#10

ZeroEntropy zerank and zembed

F 13.7/100

Discontinued retrieval API acquired by Notion. Its embedding and reranking models remain available as open weights for self-hosting.

Verdict All four models now open weights under Apache 2.0 on Hugging Face. The hosted API was discontinued after 4 September 2026 and signups closed on 24 July 2026.

Choose it for Only as open weights for teams that can self-host a reranker or embedding model.

Strengths

  • All four models now open weights under Apache 2.0 on Hugging Face
  • A migration guide with self-hosting recipes for Baseten and Modal and named hosted alternatives
  • 42 days' notice before the API was retired

Weaknesses

  • The hosted API was discontinued after 4 September 2026 and signups closed on 24 July 2026
  • The docs and pricing page still advertise per-token API prices without mentioning the shutdown
  • Nothing published on what happens to customer documents after the shutdown

Price Pay per useAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Nova Multimodal Embeddings

Head to head

All 45 comparisons in this category

Questions

What are the highest-rated embedding and reranking APIs for AI agents?

Amazon Nova Multimodal Embeddings has the highest benchmark score of the 10 ranked embedding and reranking APIs, 75 (BB). OpenAI embeddings is second with 73.2 (BB).

How many embedding and reranking APIs are agent-ready?

4 of the 10 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.

Which embedding and reranking APIs accept x402 payments?

None of the ranked listings here accepts x402 for its main call yet.

How is this list ranked?

By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 8 October 2026.

How this list is made

The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.

Full ranked table · 45 head-to-head comparisons · Best tools in every category

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.