# NVIDIA NeMo Retriever Embedding and Reranking NIMs (slim) > NVIDIA's NeMo Retriever Embedding and Reranking NIMs are GPU containers that run text and image embedding models and rerankers behind a local REST API, with /v1/embeddings in the OpenAI shape and /v1/ranking. - Full: https://www.anchorterminal.com/tools/nvidia-nemo-retriever.md (~8,450 tokens) · this version ~1,980 tokens · JSON https://www.anchorterminal.com/tools/nvidia-nemo-retriever.json · canonical https://www.anchorterminal.com/tools/nvidia-nemo-retriever - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **C · 61/100 · rank #439 of 842 · #6 in Embeddings & rerankers · not agent-ready · confidence medium** Assessment: Self-hosted containers with OpenAPI 3.1 files, typed request fields, five embedding output types and a dated end-of-life list. The API has no authentication or rate limiting of its own, production use needs an NVIDIA AI Enterprise licence at $4,500 a GPU a year, and the release notes carry no dates. ## Facts - Kind: HTTP API · vendor: NVIDIA · category: Embeddings & rerankers · legal entity: NVIDIA Corporation · provenance 83/100 - Local only (HTTP): pypi `langchain-nvidia-ai-endpoints` - Auth: None · pricing: Freemium · x402: no · licence: Proprietary containers under the NVIDIA Software Licence Agreement and Product-Specific Terms for AI Products. Models carry their own licences, such as OpenMDW 1.1 for `nvidia/nemotron-3-embed-1b` and the NVIDIA Open Model Licence for the Llama Nemotron models - Probe metrics: not measured yet (probes haven't run) - Surface graded: The self-hosted containers from `nvcr.io/nim/nvidia/`, version 2.3. The hosted trial endpoints on build.nvidia.com are recorded and not graded - Endpoints: Embedding NIM: POST `/v1/embeddings`, GET `/v1/models`, `/v1/health/ready`, `/v1/health/live`, `/v1/metrics`, `/v1/metadata`, `/v1/manifest`, `/v1/version` and a licence endpoint. Reranking NIM: POST `/v1/ranking` and the same GET set. Optional KServe V2 gRPC - Embedding models: `nvidia/nemotron-3-embed-1b` (4,096 tokens, 2048 dimensions), `nvidia/llama-nemotron-embed-vl-1b-v2` (2,048 tokens, text and image), `nvidia/llama-nemotron-embed-1b-v2` and `nvidia/llama-nemotron-embed-300m-v2` (8,192 tokens), `nvidia/nv-embedqa-e5-v5`, `baai/bge-m3`, `baai/bge-large-zh-v1.5` - Reranking models: `nvidia/llama-nemotron-rerank-vl-1b-v2`, `nvidia/llama-nemotron-rerank-1b-v2`, `nvidia/llama-nemotron-rerank-500m-v2`, each at 8,192 tokens - Request limits: Up to 8,192 inputs a call on `/v1/embeddings` and 512 passages on `/v1/ranking` per the OpenAPI files. Inline images up to 5 MiB by default and 8192 x 16384 pixels - Output sizing: `dimensions` of 128, 256, 384, 512, 768, 1024, 1536 or 2048 on models with dynamic embeddings. `embedding_type` of float, int8, uint8, binary or ubinary. `encoding_format` float or base64 - Errors: JSON with `object: "error"`, a message and a type. 400, 404, 415, 422 and 503 documented, with nine example messages - Hardware: NVIDIA GPUs from A10G and L4 to H100, H200, B200 and GB200, plus DGX Spark on arm64 from 2.3. x86 hosts need at least 8 cores. Multi-instance GPU mode and multi-GPU deployment are not supported - Licence to run: Free for research, development and testing on up to 16 GPUs through the NVIDIA developer programme. Production needs NVIDIA AI Enterprise, with a 90-day trial licence - Observability: Prometheus metrics at `/v1/metrics`, OpenTelemetry metrics and traces over OTLP HTTP with `NIM_ENABLE_OTEL=1`, and logs in pretty, JSON or compact format - Releases: Embedding 2.0 (image pushed 1 June 2026), 2.2 and 2.2.2, 2.3 (3 August 2026). Reranking 2.0 (1 June 2026) and 2.3 (27 July 2026). Dates are from the NGC registry, as the release notes carry none - Support branches: Feature branch releases are supported for one month and production branches for nine months, per the NVIDIA AI Enterprise lifecycle policy - Prices: NVIDIA AI Enterprise on a cloud marketplace, per GPU $1 per GPU-hour - Scores: Reliability 53, Performance pending, Schema & documentation 78, Agent ergonomics 73, Security & auth 55, Payments & pricing 40, Task success pending, Maintenance & community 57, Transparency & trust 71 · total over the 7 assessed categories - Why: Reliability, Local-software reading, for the containers the owner runs. · Schema & documentation, OpenAPI 3.1 files to download for both services. · Agent ergonomics, Vectors can be shortened with `dimensions` on three models or packed as int8 or binary, and returned as base64 (20 of 25). · Security & auth, The API takes no credential. · Payments & pricing, Proprietary software with a paid licence, so scored on that licence and not by the free self-hosted rule. · Maintenance & community, The newest image we could date is `nemotron-3-embed-1b`, updated on NGC on 5 August 2026, 64 days before the check, with Embedding 2.3 pushe… · Transparency & trust, Closed runtime under published terms, with model licences named per model and weights for `nvidia/nemotron-3-embed-1b` on Hugging Face (18 o… - Sources: 30, open questions: 8, both in the full twin - Capabilities: embed.text, embed.multimodal, embed.multilingual, rerank - JSON: https://www.anchorterminal.com/api/v1/tools/nvidia-nemo-retriever.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/nvidia-nemo-retriever.svg` or a link to https://www.anchorterminal.com/tools/nvidia-nemo-retriever from a page on nvidia.com or one of its subdomains, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Send `input_type` as `query` or `passage` on every embedding call. Asymmetric models return HTTP 400 without it, and the wrong value lowers retrieval accuracy per the docs 2. Do not send `dimensions` and `embedding_type` together, and send only 2048 or nothing for `dimensions` on `nvidia/nemotron-3-embed-1b` 3. Poll `/v1/health/ready` before the first call. The Docker health check can report unhealthy while the NIM is ready, per the 2.3 known issues 4. Check the image tag on NGC before pulling. The guide uses `nemotron-3-embed-1b:2.3`, and NGC's record for that image listed tags up to 2.2.2 on 8 October 2026 5. Put a proxy with authentication and TLS in front of port 8000, and sort `/v1/ranking` results yourself as the request has no top-n field ## Connect ```bash echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin docker run -it --rm --runtime=nvidia --gpus all --shm-size=16GB -e HF_TOKEN -v ~/.cache/nim/cache:/opt/cache -v ~/.cache/nim/weights:/model -u $(id -u) -p 8000:8000 nvcr.io/nim/nvidia/nemotron-3-embed-1b:2.3 ``` ```bash curl -X POST http://localhost:8000/v1/embeddings \ -H 'accept: application/json' -H 'Content-Type: application/json' \ -d '{"input":["What is NVIDIA?"],"model":"nvidia/nemotron-3-embed-1b","input_type":"query","modality":"text","embedding_type":"float","encoding_format":"float"}' ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/nvidia-nemo-retriever ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Cohere Embed and Rerank | BB | 72.5 | embed.text, embed.multimodal, embed.multilingual, rerank | https://www.anchorterminal.com/tools/cohere-embed.min.md | | Jina Embeddings and Reranker | C | 61 | embed.text, embed.multimodal, embed.multilingual, rerank | https://www.anchorterminal.com/tools/jina-embeddings.min.md | | Voyage AI embeddings and rerankers | C | 58.8 | embed.text, embed.multimodal, embed.multilingual, rerank | https://www.anchorterminal.com/tools/voyage-ai.min.md | | Gemini Embedding | BB | 70.6 | embed.text, embed.multimodal, embed.multilingual | https://www.anchorterminal.com/tools/gemini-embedding.min.md | | Nomic Embed | D | 49.2 | embed.text, embed.multimodal, embed.multilingual | https://www.anchorterminal.com/tools/nomic-embed.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)