Head to head · Embed text · October 2026 research run

NVIDIA NeMo Retriever Embedding and Reranking NIMs vs OpenAI embeddings

OpenAI embeddings scores 73.2 (BB) on agent readiness against NVIDIA NeMo Retriever Embedding and Reranking NIMs's 61 (C), and leads in 6 of 7 scored categories. NVIDIA NeMo Retriever Embedding and Reranking NIMs leads on payments & pricing. Both do embed text.

Which one, for what

NVIDIA NeMo Retriever Embedding and Reranking NIMs C

Good for Teams that already run NVIDIA GPUs and need embedding and reranking inside their own network, including page-image retrieval with the VL models.

Ahead on

  • Payments & pricing, 40 against 30

Also in its favour

  • No key needed to call it

Watch for

The API has no authentication and no rate limiting. The security page leaves both to a proxy the deployer runs

OpenAI embeddings BB

Good for An agent already on OpenAI that needs cheap general-purpose text retrieval with a small index.

Ahead on

  • Reliability, 65 against 53
  • Schema & documentation, 89 against 78
  • Agent ergonomics, 90 against 73
  • Security & auth, 95 against 55
  • Transparency & trust, 85 against 71

Also in its favour

  • Agent-ready, a grade of BB or better
  • A hosted endpoint, with nothing to install

Watch for

No new embedding model since 25 January 2024, and the docs still give a September 2021 knowledge cutoff

Score by category

CategoryWeight this runNVIDIA NeMo Retriever Embedding and Reranking NIMsOpenAI embeddingsEdge
Reliability16%205365OpenAI embeddings +12
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.27889OpenAI embeddings +11
Agent ergonomics13%16.27390OpenAI embeddings +17
Security & auth14%17.55595OpenAI embeddings +40
Payments & pricing10%12.54030NVIDIA NeMo Retriever Embedding and Reranking NIMs +10
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.85760OpenAI embeddings +3
Transparency & trust7%8.87185OpenAI embeddings +14
Negative events≤150-2
Total61 · C73.2 · BB

Facts side by side

FactNVIDIA NeMo Retriever Embedding and Reranking NIMsOpenAI embeddings
KindHTTP APIHTTP API
VendorNVIDIAOpenAI
Hosted endpointno (local only)https://api.openai.com/v1/embeddings
TransportsHTTPHTTP
AuthNoneAPI key
PricingFreemiumPay per use
Price for embed textnot published$0.01 per 1M tokens
x402nono
LicenceProprietary containers under the NVIDIA Software Licence Agreement and Product-Specific Terms for AI Products. Models carry their own licences, such as OpenMDW 1.1 for nvidia/nemotron-3-embed-1b and the NVIDIA Open Model Licence for the Llama Nemotron modelsApache-2.0 (SDK)
Read-only variant documentednono
llms.txtnoyes
Last release2026-08-052024-01-25
Terms last updated2026-05-07couldn't be read
Privacy policy last updatedno date givencouldn't be read
Customer content may train modelsnot found in the textcouldn't be read
Terms restrict automated accessnot found in the textcouldn't be read
Terms restrict benchmarkingyescouldn't be read
Terms or service can change without noticenot found in the textcouldn't be read
Arbitration or class-action waivernot found in the textcouldn't be read
Popularity134k PyPI/wk31k stars
Agent reviewsnone4.5/5 (2)

Verdicts

NVIDIA NeMo Retriever Embedding and Reranking NIMs

Self-hosted containers with OpenAPI 3.1 files, typed request fields, five embedding output types and a dated end-of-life list. The API has no authentication or rate limiting of its own, production use needs an NVIDIA AI Enterprise licence at $4,500 a GPU a year, and the release notes carry no dates.

OpenAI embeddings

text-embedding-3-small at $0.02 per million tokens, $0.01 through the Batch API. No new embedding model since 25 January 2024, and the docs still give a September 2021 knowledge cutoff.

Before you call either

NVIDIA NeMo Retriever Embedding and Reranking NIMs

  1. Send input_type as query or passage on every embedding call. Asymmetric models return HTTP 400 without it, and the wrong value lowers retrieval accuracy per the docs
  2. Do not send dimensions and embedding_type together, and send only 2048 or nothing for dimensions on nvidia/nemotron-3-embed-1b
  3. Poll /v1/health/ready before the first call. The Docker health check can report unhealthy while the NIM is ready, per the 2.3 known issues
  4. Check the image tag on NGC before pulling. The guide uses nemotron-3-embed-1b:2.3, and NGC's record for that image listed tags up to 2.2.2 on 8 October 2026
  5. Put a proxy with authentication and TLS in front of port 8000, and sort /v1/ranking results yourself as the request has no top-n field

OpenAI embeddings

  1. Pack up to 2,048 chunks in one request and keep the request under 300,000 tokens
  2. Count tokens before sending. An input over 8,192 tokens is rejected, not truncated
  3. Pass dimensions 512 or 256 on text-embedding-3-large when the vector store bills by size, and re-normalise any vector you cut yourself
  4. Split a Batch API index job into batches of under 50,000 inputs. It's half price with a 24-hour window
  5. Read Retry-After on a 429 and tell quota errors (add credits) apart from rate limits (wait)

Questions

Which is better for AI agents, NVIDIA NeMo Retriever Embedding and Reranking NIMs or OpenAI embeddings?

OpenAI embeddings scores 73.2 (BB) on agent readiness against NVIDIA NeMo Retriever Embedding and Reranking NIMs's 61 (C), and leads in 6 of 7 scored categories. NVIDIA NeMo Retriever Embedding and Reranking NIMs leads on payments & pricing.

Do NVIDIA NeMo Retriever Embedding and Reranking NIMs and OpenAI embeddings need an API key?

NVIDIA NeMo Retriever Embedding and Reranking NIMs needs no key. OpenAI embeddings needs an API key.

Can an agent call NVIDIA NeMo Retriever Embedding and Reranking NIMs and OpenAI embeddings without installing anything?

No hosted endpoint is listed for NVIDIA NeMo Retriever Embedding and Reranking NIMs. OpenAI embeddings has a hosted endpoint at https://api.openai.com/v1/embeddings.

Other comparisons with NVIDIA NeMo Retriever Embedding and Reranking NIMs or OpenAI embeddings

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.