Best of · Models & inference
Best embedding and reranking APIs for AI agents
All 10 ranked embedding and reranking APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 10 ranked
- 4 agent-ready
- 9 hosted endpoints
- Updated 8 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Highest score overall
Amazon Nova Multimodal Embeddings BB
BB, 75/100 on the benchmark.
Also OpenAI embeddings, BB, 73.2/100.
Schema & documentation
92/100 on schema & documentation, against 76 for the overall leader.
Agent ergonomics
Voyage AI embeddings and rerankers C
98/100 on agent ergonomics, against 78 for the overall leader.
Maintenance & community
90/100 on maintenance & community, against 50 for the overall leader.
Transparency & trust
85/100 on transparency & trust, against 79 for the overall leader.
Self-hosting under an open licence
NVIDIA NeMo Retriever Embedding and Reranking NIMs C
self-hosted, Proprietary containers under the NVIDIA Software Licence Agreement and Product-Specific Terms for AI Products licence.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | Amazon Nova Multimodal Embeddings Amazon Web Services |
BB 75 | Suited to mixed-media retrieval for teams already on AWS, especially video and audio archives processed through S3. | Pay per use | hosted |
| 2 | OpenAI embeddings OpenAI |
BB 73.2 | An agent already on OpenAI that needs cheap general-purpose text retrieval with a small index. | Pay per use | hosted |
| 3 | Cohere Embed and Rerank Cohere |
BB 72.5 | Best when reranking is the job, or for long multilingual documents and image-heavy material where a 128K embedding context helps, with a cheaper Fast model for queries against a Pro index. | $2 / 1k req | hosted |
| 4 | Gemini Embedding |
BB 70.6 | Multimodal corpora, especially video and audio, and for agents already on Google Cloud. | Freemium | hosted |
| 5 | Jina Embeddings and Reranker Jina AI (Elastic) |
C 61 | Reranking large candidate sets and multimodal corpora with audio or video. | Freemium | hosted |
| 6 | NVIDIA NeMo Retriever Embedding and Reranking NIMs NVIDIA |
C 61 | Teams that already run NVIDIA GPUs and need embedding and reranking inside their own network, including page-image retrieval with the VL models. | Freemium | local |
| 7 | Voyage AI embeddings and rerankers Voyage AI (MongoDB) |
C 58.8 | Retrieval quality across domains, code and long documents, with a reranker from the same key. | Freemium | hosted |
| 8 | Mistral Embed and Codestral Embed Mistral AI |
C 57.9 | EU data residency, Mistral-only stacks and code retrieval with small vectors. | Freemium | hosted |
| 9 | Nomic Embed Nomic, Inc. |
D 49.2 | Teams that want a hosted endpoint for an open-weight model they can also run themselves, with the same vectors either way. | Freemium | hosted |
| 10 | ZeroEntropy zerank and zembed ZeroEntropy |
F 13.7 | Only as open weights for teams that can self-host a reranker or embedding model. | Pay per use | hosted |
How to choose
- Vector dimensions and sizeCheck the vector dimensions and whether they can be shortened, because the index stores and scans every dimension for every query.
- Query and document purposesCheck whether queries and documents can be embedded with different purpose settings, since the setting can change which results rank highest.
- Languages and input typesConfirm which languages and modalities the model supports, including images and code, because retrieval quality can fall outside those.
- Maximum input lengthCheck the maximum input length and whether long documents are segmented for you, because a truncated chunk silently loses the text that answers a query.
How the benchmark tests this category. The same corpus and queries embedded with each model, then the top results reranked. We check retrieval quality against labelled answers, latency and the cost per million tokens.
Each one in detail
Amazon Nova Multimodal Embeddings
BB 75/100Amazon Nova Multimodal Embeddings is an AWS model on Amazon Bedrock that turns text, images, document images, video and audio into vectors in one space, at 256, 384, 1024 or 3072 dimensions, through synchronous and asynchronous calls.
Verdict One model embeds text, images, document images, video and audio into a shared space, with nine documented purpose settings and published per-unit prices. It runs in US East (N. Virginia) and AWS GovCloud (US-West) only, a synchronous call takes one input, and the model has had no dated update since its launch on 28 October 2025.
Choose it for Suited to mixed-media retrieval for teams already on AWS, especially video and audio archives processed through S3.
Strengths
- Text, images, document images, video and audio share one vector space, with four output sizes from 256 to 3072
embeddingPurposehas nine documented values, with separate settings for indexing and for each retrieval type- Published quotas of 2,000 requests a minute and 30 concurrent asynchronous jobs per Region
Weaknesses
- In-Region inference in us-east-1 and us-gov-west-1 only, with no cross-Region inference profile
- A synchronous request embeds one item, with at most 8,192 characters of inline text or 30 seconds of audio or video
- The Bedrock model card marks Invoke as unsupported while the Nova guide documents
InvokeModelfor synchronous calls
Price Pay per useAuth API keyx402 nohosted
OpenAI embeddings
BB 73.2/100OpenAI's text embedding API, with adjustable output dimensions for search and retrieval applications.
Verdict text-embedding-3-small at $0.02 per million tokens, $0.01 through the Batch API. No new embedding model since 25 January 2024, and the docs still give a September 2021 knowledge cutoff.
Choose it for An agent already on OpenAI that needs cheap general-purpose text retrieval with a small index.
Strengths
- text-embedding-3-small at $0.02 per million tokens, $0.01 through the Batch API
- Restricted project keys are set per endpoint, so an agent's key can be cut down to read and model calls
- Up to 2,048 inputs and 300,000 tokens in one request
Weaknesses
- No new embedding model since 25 January 2024, and the docs still give a September 2021 knowledge cutoff
- Text only, 8,192 tokens an input, and no reranker
- Over-long inputs fail rather than being truncated, and output is float or base64 only
Price Pay per useAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Nova Multimodal Embeddings
Cohere Embed and Rerank
BB 72.5/100Cohere's Embed API turns text, images and mixed text-and-image inputs such as PDF pages into vectors with Embed 5 Pro and Fast, and its Rerank API reorders search results with Rerank 4.
Verdict Embed 5 Pro and Fast share one embedding space with 128K context and compressed outputs, and embed and rerank prices are public. Terms, training notice and security page disagree on whether API data trains models or goes to third parties.
Choose it for Best when reranking is the job, or for long multilingual documents and image-heavy material where a 128K embedding context helps, with a cheaper Fast model for queries against a Pro index.
Strengths
- Embed 5 Pro and Fast share one embedding space, so a Pro index answers Fast queries
- 128K context with six output sizes and int8, binary and base64 output
- Rerank 4 Pro and Fast with 32K context, top_n and published per-search prices
Weaknesses
- Terms, training notice and security page disagree on whether API data trains models or goes to third parties
- A Google Cloud outage degraded embed and rerank for about four hours on 1 September 2026, and Embed 5 isn't yet a status component
- 96 inputs a call, and input_type is required
Price $2 / 1k reqAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Nova Multimodal Embeddings
Gemini Embedding
BB 70.6/100gemini-embedding-2, Google's multimodal embedding model, takes text, images, video, audio and PDFs into one 3072-dimension space (truncatable to 128) at 8,192 input tokens in 100+ languages.
Verdict Text, images, video, audio and PDFs interleaved in one request and one vector space. $0.20 per million text tokens, against $0.02 for OpenAI's small model.
Choose it for Multimodal corpora, especially video and audio, and for agents already on Google Cloud.
Strengths
- Text, images, video, audio and PDFs interleaved in one request and one vector space
- Any output size from 128 to 3072, with truncated vectors returned normalised
- Batch API at half the standard embedding price
Weaknesses
- $0.20 per million text tokens, against $0.02 for OpenAI's small model
- 8,192 input tokens and float output only
- No reranker on the Gemini API
Price FreemiumAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Nova Multimodal Embeddings
Jina Embeddings and Reranker
C 61/100jina-embeddings-v5 in text and omni (text, image, audio, video, PDF) variants at up to 32,768 tokens, plus the jina-reranker-v3.5 at 131,072 tokens a call.
Verdict jina-reranker-v3.5 (20 July 2026) with a 131,072-token window and no document cap. No price per token in any currency on the public pages.
Choose it for Reranking large candidate sets and multimodal corpora with audio or video.
Strengths
- jina-reranker-v3.5 (20 July 2026) with a 131,072-token window and no document cap
- v5-omni embeds text, images, audio, video and PDFs into one space
- OpenAPI 3.1 file with enums for model, task and embedding_type, and error responses from 400 to 504
Weaknesses
- No price per token in any currency on the public pages
- One prepaid balance shared with Reader and Search, so a scraping job can drain the embedding budget
- 26 automated incidents on the status feed from 15 September to 1 October 2026, and no status component for v5-omni or reranker v3.5
Price FreemiumAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Nova Multimodal Embeddings
NVIDIA NeMo Retriever Embedding and Reranking NIMs
C 61/100NVIDIA's NeMo Retriever Embedding and Reranking NIMs are GPU containers that run text and image embedding models and rerankers behind a local REST API, with /v1/embeddings in the OpenAI shape and /v1/ranking.
Verdict Self-hosted containers with OpenAPI 3.1 files, typed request fields, five embedding output types and a dated end-of-life list. The API has no authentication or rate limiting of its own, production use needs an NVIDIA AI Enterprise licence at $4,500 a GPU a year, and the release notes carry no dates.
Choose it for Teams that already run NVIDIA GPUs and need embedding and reranking inside their own network, including page-image retrieval with the VL models.
Strengths
- OpenAPI 3.1 files for both services, with enums for
input_type,modality,embedding_typeandtruncateand no extra properties allowed /v1/embeddingsfollows the OpenAI shape, and a-queryor-passagemodel suffix replacesinput_typefor OpenAI clients- Output can be shrunk with
dimensionsfrom 128 to 2048 on three models, or withint8,uint8,binaryandubinarytypes
Weaknesses
- The API has no authentication and no rate limiting. The security page leaves both to a proxy the deployer runs
- Production use needs an NVIDIA AI Enterprise licence, $4,500 a GPU a year or $1 a GPU-hour on cloud marketplaces, plus the GPU
- Release notes carry no dates, and environment variable names changed between 2.0 and 2.3 without a note in the 2.2 or 2.3 notes we read
Price FreemiumAuth Nonex402 nolocal
Full assessment · Against #1, Amazon Nova Multimodal Embeddings
Voyage AI embeddings and rerankers
C 58.8/100Embedding and reranking models for text, code and multimodal retrieval from MongoDB-owned Voyage AI.
Verdict 200 million free tokens per current model, then $0.02 to $0.12 per million. Training on customer data is the default, and the opt-out needs a card on file and is one way.
Choose it for Retrieval quality across domains, code and long documents, with a reranker from the same key.
Strengths
- 200 million free tokens per current model, then $0.02 to $0.12 per million
- Domain models for code, finance and law, a multimodal model and contextualised chunk embeddings
- Output in float, int8, uint8, binary or ubinary at 256 to 2048 dimensions, per request
Weaknesses
- Training on customer data is the default, and the opt-out needs a card on file and is one way
- No security.txt, and no status page linked or reachable
- Rate-limit tiers only begin once a payment method is added
Price FreemiumAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Nova Multimodal Embeddings
Mistral Embed and Codestral Embed
C 57.9/100Mistral's API for generating text and code embeddings.
Verdict EU and US regional endpoints and a French legal entity. Embedding API uptime of 94.36 per cent over 90 days on Mistral's status page, with incidents on 12 and 27 August 2026.
Choose it for EU data residency, Mistral-only stacks and code retrieval with small vectors.
Strengths
- EU and US regional endpoints and a French legal entity
- codestral-embed with up to 3072 dimensions, first-n truncation and int8 or binary output
- OpenAPI document and llms.txt for the whole API
Weaknesses
- Embedding API uptime of 94.36 per cent over 90 days on Mistral's status page, with incidents on 12 and 27 August 2026
- 8k context on both models, and text or code only
- mistral-embed dates from December 2023 with fixed 1024-dimension float output, and nothing new since May 2025
Price FreemiumAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Nova Multimodal Embeddings
Nomic Embed
D 49.2/100Nomic's hosted embedding endpoints on the Atlas API turn text and images into vectors with the open-weight Nomic Embed models. Agents call them over HTTP with an API key or through the Python and TypeScript clients.
Verdict The text models have Apache-2.0 weights and a public OpenAPI 3.1 contract, so vectors made through the hosted endpoint can be reproduced locally. Nomic's current site and documentation index describe a construction-industry product, no rendered public page prices the endpoint, and no published terms or status component name it.
Choose it for Teams that want a hosted endpoint for an open-weight model they can also run themselves, with the same vectors either way.
Strengths
- Weights for nomic-embed-text-v1, v1.5, v2-moe, nomic-embed-code and nomic-embed-vision-v1.5 are Apache-2.0 on Hugging Face
- Public OpenAPI 3.1 document for the Atlas API (v0.57.0) with typed request and response schemas for both embedding endpoints
- nomic-embed-text-v1.5 accepts a dimensionality from 64 to 768, and inputs up to 8,192 tokens per text
Weaknesses
- docs.nomic.ai/llms.txt and www.nomic.ai now describe a product for architecture, engineering and construction firms, and the documentation index no longer lists the embedding pages
- No rendered public page states a price for the endpoint. The $1 per 10M tokens figure comes from the Atlas web app's script
- status.nomic.ai has no component for api-atlas.nomic.ai, and no SLA was found
Price FreemiumAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Nova Multimodal Embeddings
ZeroEntropy zerank and zembed
F 13.7/100Discontinued retrieval API acquired by Notion. Its embedding and reranking models remain available as open weights for self-hosting.
Verdict All four models now open weights under Apache 2.0 on Hugging Face. The hosted API was discontinued after 4 September 2026 and signups closed on 24 July 2026.
Choose it for Only as open weights for teams that can self-host a reranker or embedding model.
Strengths
- All four models now open weights under Apache 2.0 on Hugging Face
- A migration guide with self-hosting recipes for Baseten and Modal and named hosted alternatives
- 42 days' notice before the API was retired
Weaknesses
- The hosted API was discontinued after 4 September 2026 and signups closed on 24 July 2026
- The docs and pricing page still advertise per-token API prices without mentioning the shutdown
- Nothing published on what happens to customer documents after the shutdown
Price Pay per useAuth API keyx402 nohosted
Full assessment · Against #1, Amazon Nova Multimodal Embeddings
Head to head
- Amazon Nova Multimodal Embeddings vs OpenAI embeddings BB 75 vs BB 73.2
- Amazon Nova Multimodal Embeddings vs Cohere Embed and Rerank BB 75 vs BB 72.5
- Amazon Nova Multimodal Embeddings vs Gemini Embedding BB 75 vs BB 70.6
- Amazon Nova Multimodal Embeddings vs Jina Embeddings and Reranker BB 75 vs C 61
- Cohere Embed and Rerank vs OpenAI embeddings BB 72.5 vs BB 73.2
- Gemini Embedding vs OpenAI embeddings BB 70.6 vs BB 73.2
- Jina Embeddings and Reranker vs OpenAI embeddings C 61 vs BB 73.2
- Cohere Embed and Rerank vs Gemini Embedding BB 72.5 vs BB 70.6
- Cohere Embed and Rerank vs Jina Embeddings and Reranker BB 72.5 vs C 61
- Gemini Embedding vs Jina Embeddings and Reranker BB 70.6 vs C 61
Questions
What are the highest-rated embedding and reranking APIs for AI agents?
Amazon Nova Multimodal Embeddings has the highest benchmark score of the 10 ranked embedding and reranking APIs, 75 (BB). OpenAI embeddings is second with 73.2 (BB).
How many embedding and reranking APIs are agent-ready?
4 of the 10 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which embedding and reranking APIs accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 8 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 45 head-to-head comparisons · Best tools in every category