# DeepInfra (slim) > DeepInfra is a hosted inference API for open-weight and some third-party models, covering chat, embeddings, reranking, image, video and speech. It answers OpenAI-style and Anthropic-style calls at api.deepinfra.com with a Bearer key. - Full: https://www.anchorterminal.com/tools/deepinfra.md (~8,350 tokens) · this version ~2,130 tokens · JSON https://www.anchorterminal.com/tools/deepinfra.json · canonical https://www.anchorterminal.com/tools/deepinfra - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **B · 63/100 · rank #371 of 842 · #9 in Model APIs & inference · not agent-ready · confidence medium** Assessment: The model list, context sizes and per-token prices are readable without a key, and keys can carry an IP allowlist, a monthly spending cap and model-limited JWTs. Deprecated models get one week's notice and are then redirected to another model, there is no changelog or SLA, and an account needs a card or prepayment before any call. ## Facts - Kind: Model API · vendor: Deep Infra Inc. · category: Model APIs & inference · legal entity: Deep Infra Inc. · provenance 74/100 - Endpoint: `https://api.deepinfra.com/v1/openai` (HTTP) - Auth: API key · pricing: Pay per use · x402: no · licence: Proprietary service under the DeepInfra Terms of Service. The Python and Node SDKs and the docs repository are MIT - Probe metrics: not measured yet (probes haven't run) - Endpoints: OpenAI-style chat, completions, embeddings, images, audio, videos, files and batches under https://api.deepinfra.com/v1/openai, Anthropic-style `/anthropic/v1/messages` and `/anthropic/v1/messages/count_tokens`, and a native `/v1/inference/{model_name}` route for every model type - Models on 8 October 2026: 181 ids from `/v1/openai/models`. Tags count 98 chat, 51 vision, 28 image generation, 25 embedding, 13 text to speech, 10 video and 7 speech to text. 47 are tagged for prompt caching - Rate limits: 200 concurrent requests per model per account, with increases requested in the dashboard. 429 `Rate limited` over the limit, and 429 `engine_overloaded` when a model is busy - Service tiers: Standard by default. `service_tier` of `priority` at 1.5x the base price or `flex` at 0.8x on tagged models. `fail_fast` returns 429 at once when a model is at capacity, and `models` names up to four fallbacks tried server-side - Batch: OpenAI-style files and batches for `/v1/chat/completions`, `/v1/completions` and `/v1/embeddings`, one model per file, 24-hour window, 20 per cent below real-time prices. Files expire after 30 days by default - Prompt caching: Automatic prefix caching, an optional `prompt_cache_key`, and paid retention for 5 minutes or 1 hour through `prompt_cache_options`. Cached tokens are reported in `usage.prompt_tokens_details.cached_tokens` - Structured output: `response_format` takes `json_object` and `json_schema` with `strict`. The OpenAPI document also lists a regex format - Tool calling: `tool_choice` of none, auto, required or a named function, with streaming. Parallel calls are supported with the note that quality may vary. Nested calls are not - Output cap: 16,384 tokens for most models in one response, with a documented way to continue a response up to the context window - Credentials: Bearer API keys, named and deletable by id, with `allowed_ips` and a monthly USD limit per key. Scoped JWTs limited by model, expiry (one year at most) and spending limit, counted against the signing key - Deprecation notice: At least one week by email to recent users. After the date, requests are forwarded to a replacement model. Scheduled deprecations are listed on the status page with UTC times - Data use in the terms: No sale of Customer Data and no use to train or improve any model. Zero Data Retention, except data kept at the customer's written request for support (deleted within 30 days of resolution), operational metadata and records required by law or for fraud and abuse - Certifications: The privacy policy says security measures comply with SOC 2 and ISO 27001, with measures for GDPR and HIPAA. The trust centre answered 403 and was not read - SDKs: The docs use the OpenAI and Anthropic clients with a changed base URL. Python `deepinfra` 0.3.0 (12 August 2026) and npm `deepinfra` 2.0.2 (8 May 2024), both MIT. The Node repository is at 2.1.0, which npm did not list - Other products on the same key: Private model deployments on dedicated GPUs from $0.89 a GPU-hour, GPU instances, sandboxes and hosted agents. Not graded in this listing - Models ($/1M in/out): `deepseek-ai/DeepSeek-V4-Flash-0731` $0.06/$0.18; `moonshotai/Kimi-K3` $2.85/$14.25; `meta-llama/Llama-3.3-70B-Instruct-Turbo` $0.10/$0.32 - 2026-10-08 Shutdown: `Hy3` scheduled to redirect to `tencent/Hy4-preview` - 2026-10-08 Shutdown: `Nemotron-3-Nano-30B-A3B` scheduled to redirect to `nvidia/NVIDIA-Nemotron-3.5-Lightning` - 2026-10-09 Shutdown: `Ling-3.0-flash-Fin` scheduled to redirect to `inclusionAI/Ling-3.0-flash-VL` - 2026-10-12 Shutdown: `Qwen3.8-2.4T-A95B` scheduled to redirect to `zai-org/GLM-5.3` - 2026-10-25 Shutdown: `Nemotron-3-Diarization-preview` scheduled to redirect to `nvidia/Nemotron-3-Diarization` - Scores: Reliability 70, Performance pending, Schema & documentation 69, Agent ergonomics 77, Security & auth 64, Payments & pricing 20, Task success pending, Maintenance & community 65, Transparency & trust 67 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, Model reading. · Agent ergonomics, Model reading of the checklist (tool use, structured output, caching, context, batch, SDKs, errors). · Security & auth, Model reading. · Payments & pricing, No machine payment protocol (0). · Maintenance & community, Model reading. · Transparency & trust, Closed service with terms that name the entity and California law, and an MIT licence on the SDKs and docs (15). - Sources: 31, open questions: 10, both in the full twin - Capabilities: inference.llm, inference.open-weights, embed.text, rerank, image.generate, speech.stt, speech.tts, compute.batch - JSON: https://www.anchorterminal.com/api/v1/tools/deepinfra.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/deepinfra.svg` or a link to https://www.anchorterminal.com/tools/deepinfra from a page on deepinfra.com or one of its subdomains, or the README of github.com/deepinfra/deepinfra-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Call `GET https://api.deepinfra.com/v1/openai/models` at start-up for ids, context sizes and prices. No key is needed 2. Check the `model` field of each response. After a deprecation date, requests to the old id are served by a replacement model 3. Ask the account owner for a scoped JWT limited to the models and spend the task needs, not the full API key 4. Stay under 200 concurrent requests per model. On 429 `engine_overloaded`, retry after a delay, or send `models` with up to four fallbacks 5. Never inspect a JWT with `GET /v1/scoped-jwt?jwtoken=`, which puts the token in the URL. Keep credentials in the `Authorization` header ## Connect ```bash pip install openai # or: npm install openai ``` ```bash curl "https://api.deepinfra.com/v1/openai/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $DEEPINFRA_API_KEY" \ -d '{"model":"deepseek-ai/DeepSeek-V4-Flash-0731","messages":[{"role":"user","content":"Hello"}]}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | LocalAI | B | 68 | inference.open-weights, embed.text, rerank, speech.stt, speech.tts, image.generate | https://www.anchorterminal.com/tools/localai.min.md | | Lemonade | B | 63.8 | inference.open-weights, embed.text, rerank, speech.stt, speech.tts, image.generate | https://www.anchorterminal.com/tools/lemonade.min.md | | KoboldCpp | C | 60.5 | inference.open-weights, embed.text, image.generate, speech.stt, speech.tts | https://www.anchorterminal.com/tools/koboldcpp.min.md | | Docker Model Runner | C | 57.1 | inference.open-weights, inference.llm, embed.text, rerank, image.generate | https://www.anchorterminal.com/tools/docker-model-runner.min.md | | Foundry Local | C | 60.5 | inference.open-weights, embed.text, speech.stt | https://www.anchorterminal.com/tools/foundry-local.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)