{
  "data": {
    "a": {
      "slug": "deepinfra",
      "name": "DeepInfra",
      "vendor": "Deep Infra Inc.",
      "vendorUrl": "https://deepinfra.com",
      "kind": "model",
      "category": "inference",
      "summary": "DeepInfra is a hosted inference API for open-weight and some third-party models, covering chat, embeddings, reranking, image, video and speech. It answers OpenAI-style and Anthropic-style calls at api.deepinfra.com with a Bearer key.",
      "url": "https://www.anchorterminal.com/tools/deepinfra",
      "markdownUrl": "https://www.anchorterminal.com/tools/deepinfra.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/deepinfra.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/deepinfra.json",
      "repo": "https://github.com/deepinfra/deepinfra-python",
      "license": "Proprietary service under the DeepInfra Terms of Service. The Python and Node SDKs and the docs repository are MIT",
      "transports": [
        "http"
      ],
      "remoteUrl": "https://api.deepinfra.com/v1/openai",
      "packages": [
        {
          "registry": "pypi",
          "name": "deepinfra"
        },
        {
          "registry": "npm",
          "name": "deepinfra"
        }
      ],
      "auth": "api-key",
      "authNotes": "Self-serve Bearer key from the dashboard at https://deepinfra.com/dash/api_keys after a browser sign-up with Google, GitHub, email or Okta SSO. The Anthropic-style routes also take the key in `x-api-key`. Keys can carry an IP allowlist and a monthly spending limit, and a key can mint scoped JWTs limited by model, expiry and spend.",
      "pricing": "usage",
      "pricingNotes": "Pay per token, image, audio minute or GPU-hour, with no free tier. An account must add a card or prepay. DeepSeek-V4-Flash-0731 costs $0.06 in and $0.18 out per 1M tokens, the priority tier is 1.5x, flex 0.8x and batch 20 per cent off (https://deepinfra.com/pricing, checked 2026-10-08).",
      "priceSummary": "from $0.06 / 1M in",
      "where": "hosted",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs index, the docs repository, the OpenAPI document or the pricing page (checked 2026-10-08).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 21,
        "npmWeekly": 1250,
        "pypiWeekly": 50,
        "asOf": "2026-10-08"
      },
      "docsUrl": "https://docs.deepinfra.com/",
      "rateLimitsUrl": "https://docs.deepinfra.com/account/rate-limits",
      "llmsTxt": "https://docs.deepinfra.com/llms.txt",
      "openapi": "https://api.deepinfra.com/openapi.json",
      "capabilities": [
        "inference.llm",
        "inference.open-weights",
        "embed.text",
        "rerank",
        "image.generate",
        "speech.stt",
        "speech.tts",
        "compute.batch"
      ],
      "tags": [
        "hosted",
        "model",
        "open-weights",
        "usage-based",
        "card-required",
        "openapi",
        "llms-txt",
        "openai-compatible",
        "anthropic-compatible",
        "batch",
        "prompt-caching",
        "python",
        "typescript",
        "status-page",
        "soc2"
      ],
      "lastRelease": "2026-10-07",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 63,
        "grade": "B",
        "agentReady": false,
        "rank": 371,
        "ranked": true,
        "rankOf": 842,
        "categoryRank": 9,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 77,
          "maintenance": 65,
          "payments": 20,
          "reliability": 70,
          "schema": 69,
          "security": 64,
          "transparency": 67
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-08"
        },
        "negative": 0,
        "verdict": "The model list, context sizes and per-token prices are readable without a key, and keys can carry an IP allowlist, a monthly spending cap and model-limited JWTs. Deprecated models get one week's notice and are then redirected to another model, there is no changelog or SLA, and an account needs a card or prepayment before any call.",
        "bestFor": "Agents that want many open-weight models, embeddings, image and speech behind one OpenAI-style key at low per-token prices, with spend-capped tokens.",
        "strengths": [
          "`GET /v1/openai/models` answered without a key on 8 October 2026 with 181 models, each with context length, output cap and per-token prices",
          "Scoped JWTs limit a token to named models, an expiry and a USD spending limit, and each API key can carry an IP allowlist and a monthly cap",
          "The terms commit to zero data retention and no training on Customer Data, and say that clause controls over the privacy policy and docs",
          "Batch API at 20 per cent off, a flex tier at 0.8x, prompt caching with cached-input prices and server-side fallback across up to four models",
          "Status page with 90-day history for the API, the website and 154 models, plus a dated list of scheduled deprecations"
        ],
        "weaknesses": [
          "A deprecated model gets at least one week's notice, and requests are then forwarded to a replacement model under the old id",
          "The status feed lists 50 automated major-outage observations on single models since 3 August 2026, several longer than 24 hours, with no written incident notes",
          "No free tier. The pricing page says an account must add a card or prepay before using the service",
          "No changelog, SLA, error reference or `Retry-After` header was found in the reviewed documentation",
          "The sub-processor list names three companies and is dated 6 September 2024, while the docs say Google and Anthropic receive data for their models"
        ],
        "agentNotes": [
          "Call `GET https://api.deepinfra.com/v1/openai/models` at start-up for ids, context sizes and prices. No key is needed",
          "Check the `model` field of each response. After a deprecation date, requests to the old id are served by a replacement model",
          "Ask the account owner for a scoped JWT limited to the models and spend the task needs, not the full API key",
          "Stay under 200 concurrent requests per model. On 429 `engine_overloaded`, retry after a delay, or send `models` with up to four fallbacks",
          "Never inspect a JWT with `GET /v1/scoped-jwt?jwtoken=`, which puts the token in the URL. Keep credentials in the `Authorization` header"
        ],
        "metrics": {
          "kind": "remote",
          "measured": false
        },
        "reviewCount": 0,
        "avgRating": 0,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "B",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 63
          }
        ],
        "editorialScores": {
          "ergonomics": 77,
          "maintenance": 65,
          "payments": 20,
          "reliability": 70,
          "schema": 69,
          "security": 64,
          "transparency": 60
        },
        "provenanceScore": 74
      },
      "connect": {
        "install": "pip install openai   # or: npm install openai",
        "http": "curl \"https://api.deepinfra.com/v1/openai/chat/completions\" \\\n  -H \"Content-Type: application/json\" \\\n  -H \"Authorization: Bearer $DEEPINFRA_API_KEY\" \\\n  -d '{\"model\":\"deepseek-ai/DeepSeek-V4-Flash-0731\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'"
      },
      "letme": {
        "capability": "https://letme.dev/inference.llm",
        "tool": "https://letme.dev/deepinfra"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "Deep Infra Inc.",
        "domain": "deepinfra.com",
        "domainRegistered": "2017-12-08",
        "endpointOnVendorDomain": true,
        "terms": "https://deepinfra.com/terms",
        "privacy": "https://deepinfra.com/privacy",
        "statusPage": "https://status.deepinfra.com",
        "changelog": "",
        "securityTxt": "none",
        "checked": "2026-10-08",
        "notes": [
          "The Terms of Service, last modified 17 August 2026, name Deep Infra Inc., a Delaware corporation, with California law and JAMS arbitration in San Francisco. They are written around Service Orders and govern the services, the API included.",
          "The Privacy Policy, last modified 15 August 2026, names DeepInfra, Inc., a Delaware corporation, covers the website and the APIs, and has a section on data sent to and returned by the inference service.",
          "No changelog or release notes page was found in the docs index or the site map.",
          "https://deepinfra.com/.well-known/security.txt returns 404.",
          "RDAP for deepinfra.com gives a registration date of 2017-12-08.",
          "The API answers at api.deepinfra.com, the docs at docs.deepinfra.com and the status page at status.deepinfra.com.",
          "No SLA or published DPA was found. The terms mention a DPA only where the parties execute one."
        ],
        "score": 74
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/deepinfra.json",
      "live": {
        "slug": "deepinfra",
        "probe": {
          "target": "https://api.deepinfra.com/v1/openai",
          "method": "get",
          "lastAt": "2026-10-09T10:42:41.174794945Z",
          "lastOk": true,
          "lastStatus": 404,
          "lastMs": 424,
          "authRequired": false,
          "uptime24h": 100,
          "uptime30d": 100,
          "p50ms24h": 377,
          "p95ms24h": 506,
          "samples24h": 33,
          "samples30d": 33,
          "days": [
            {
              "date": "2026-10-09",
              "probes": 33,
              "ok": 33
            }
          ]
        },
        "vendorStatus": {
          "page": "https://status.deepinfra.com",
          "indicator": "unknown",
          "summary": "no machine-readable status found",
          "checkedAt": "2026-10-09T07:57:47.5406023Z"
        },
        "updatedAt": "2026-10-09T10:42:41.174794945Z"
      }
    },
    "answer": "DeepInfra scores 63 (B) on agent readiness against Prism Inference's 60.1 (C), and leads in 4 of 7 scored categories. Prism Inference leads on schema \u0026 documentation and payments \u0026 pricing.",
    "b": {
      "slug": "prism-inference",
      "name": "Prism Inference",
      "vendor": "Prism Technologies Inc",
      "vendorUrl": "https://prisminference.com",
      "kind": "model",
      "category": "inference",
      "summary": "Prism is a hosted inference API from Prism Technologies Inc for open-weight models, aimed at coding agents. It accepts OpenAI Chat Completions, OpenAI Responses and Anthropic Messages requests at api.prisminference.com. It launched on 24 September 2026.",
      "url": "https://www.anchorterminal.com/tools/prism-inference",
      "markdownUrl": "https://www.anchorterminal.com/tools/prism-inference.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/prism-inference.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/prism-inference.json",
      "repo": "https://github.com/prismhq/hermes-prism-provider",
      "license": "Proprietary service under Prism's terms of service. The OpenAPI file declares `LicenseRef-Proprietary`. The Hermes provider plugin repository carries no licence file",
      "transports": [
        "http"
      ],
      "remoteUrl": "https://api.prisminference.com/v1",
      "packages": [],
      "auth": "api-key",
      "authNotes": "`Authorization: Bearer` with a key issued from account settings after sign-up at prisminference.com/signup. The inference endpoints also accept the key in `x-api-key`. A missing, invalid, expired or revoked key returns 401. No key scopes were found in the reviewed documentation. An agent can call `POST https://prisminference.com/api/agent-signups` with the owner's email and a username and receive a key once, but that key can't run inference until the owner supplies a six-digit emailed code, and it expires after 30 days. `GET /v1/models` needs no key.",
      "pricing": "usage",
      "pricingNotes": "Prepaid per-token pricing with no minimum. DeepSeek-V4.1-Flash is $0.09 in, $1.20 out and $0.06 cache read per million tokens, and Gemma 4 31B is $0.30, $0.40 and $0.15 (https://prisminference.com/pricing, matched by https://api.prisminference.com/v1/models). No free tier or trial credit was found, and a workspace without credit gets 402. A person funds the workspace in a browser. Elastic endpoints, dedicated deployments and batch are sold through sales with no published price.",
      "priceSummary": "from $0.09 / 1M in",
      "where": "hosted",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in llms.txt, the docs index, the OpenAPI file or the pricing page, read 2026-10-08. The docs describe prepaid credit funded by a person in the browser.",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": null,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-08"
      },
      "docsUrl": "https://docs.prisminference.com",
      "rateLimitsUrl": "https://docs.prisminference.com/rate-limits",
      "llmsTxt": "https://prisminference.com/llms.txt",
      "openapi": "https://docs.prisminference.com/openapi.yaml",
      "capabilities": [
        "inference.fast",
        "inference.open-weights",
        "inference.llm"
      ],
      "tags": [
        "hosted",
        "model",
        "open-weights",
        "fast",
        "usage-priced",
        "prepaid",
        "openapi",
        "llms-txt",
        "openai-compatible",
        "anthropic-compatible",
        "zero-retention",
        "status-page",
        "new"
      ],
      "lastRelease": "2026-10-06",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 60.1,
        "grade": "C",
        "agentReady": false,
        "rank": 479,
        "ranked": true,
        "rankOf": 842,
        "categoryRank": 10,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 68,
          "maintenance": 49,
          "payments": 30,
          "reliability": 65,
          "schema": 82,
          "security": 65,
          "transparency": 61
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-08"
        },
        "negative": -2,
        "negativeNotes": [
          "2026-10-08. The home page shows a '99.99% Uptime SLA' tile, and the pricing page says 'No minimums, no rate limits' above the per-token table. The terms of 9 September 2026 say the services have no guaranteed uptime or service credit unless a separate written agreement says otherwise, and the docs describe per-key rate limits that return 429. No SLA document was found. A misleading claim, with the smallest deduction because the terms and docs state the real position (https://prisminference.com/, https://prisminference.com/pricing, https://prisminference.com/terms, https://docs.prisminference.com/rate-limits)."
        ],
        "verdict": "Three wire formats, a public OpenAPI 3.1 file, per-token prices in a keyless catalogue and zero data retention by default on every tier. The service launched on 24 September 2026 with two models, one of them by request, from a two-person company. No rate-limit numbers, SLA document, free tier or deprecation policy was found.",
        "bestFor": "Coding agents that want DeepSeek-V4.1-Flash at a low input price, with no retention, through whichever of the three wire formats the harness already speaks.",
        "strengths": [
          "One key works across OpenAI Chat Completions, OpenAI Responses and Anthropic Messages, with a public OpenAPI 3.1 file, llms.txt and Markdown docs",
          "Zero data retention is the default on every tier, and the privacy policy, terms and docs all say inputs and outputs are never used for training",
          "Every error carries a stable `code`, a `retryable` flag, a `fix` hint and a `docs_url`, and 429 carries `Retry-After` in seconds",
          "`GET /v1/models` answers without a key and returns context length, maximum output and per-token prices for each model",
          "DeepSeek-V4.1-Flash is listed with a 1M-token context and 384,000 output tokens at $0.09 in and $1.20 out per million"
        ],
        "weaknesses": [
          "Two models. The docs mark Gemma 4 31B as request access per organisation, while llms.txt and the keyless catalogue list it as available",
          "No rate-limit numbers are published. The docs say per-key limits exist, and the pricing page says 'no rate limits'",
          "The home page shows a '99.99% Uptime SLA' tile, while the terms say there is no guaranteed uptime without a separate written agreement",
          "No free tier found. Billing is prepaid, a person funds the workspace in a browser, and the agent sign-up key needs an emailed code",
          "The service launched on 24 September 2026. The changelog has one entry, and no deprecation policy, sub-processor list or certification was found"
        ],
        "agentNotes": [
          "Call `GET https://api.prisminference.com/v1/models` at start-up, with no key, and use only ids it returns. Expect 403 on `gemma-4-31b` without organisation access",
          "Use base URL `https://api.prisminference.com/v1` for OpenAI clients and `https://api.prisminference.com` with no `/v1` for Anthropic clients",
          "Read `error.retryable` before retrying, and wait for `Retry-After` on 429, which covers both key limits and model capacity",
          "Send `reasoning_effort: \"none\"` or `low` when latency matters. Reasoning is on by default and its tokens are billed as output",
          "Keep conversation state yourself and send `store: false` on Responses. `previous_response_id`, stored responses and hosted tools aren't supported"
        ],
        "metrics": {
          "kind": "remote",
          "measured": false
        },
        "reviewCount": 0,
        "avgRating": 0,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "C",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 60.1
          }
        ],
        "editorialScores": {
          "ergonomics": 68,
          "maintenance": 49,
          "payments": 30,
          "reliability": 65,
          "schema": 82,
          "security": 65,
          "transparency": 43
        },
        "provenanceScore": 79
      },
      "connect": {
        "install": "pip install openai   # or: npm install openai, base URL https://api.prisminference.com/v1. Anthropic SDKs use https://api.prisminference.com with no /v1",
        "http": "curl \"https://api.prisminference.com/v1/chat/completions\" \\\n  -H \"Authorization: Bearer $PRISM_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"deepseek-v4.1-flash\",\"messages\":[{\"role\":\"user\",\"content\":\"Return pong.\"}]}'",
        "claudeCode": "export ANTHROPIC_BASE_URL=https://api.prisminference.com\nexport ANTHROPIC_AUTH_TOKEN=$PRISM_API_KEY\nexport ANTHROPIC_MODEL=deepseek-v4.1-flash\nexport ANTHROPIC_SMALL_FAST_MODEL=gemma-4-31b"
      },
      "letme": {
        "capability": "https://letme.dev/inference.fast",
        "tool": "https://letme.dev/prism-inference"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "Prism Technologies Inc",
        "domain": "prisminference.com",
        "domainRegistered": "2026-09-09",
        "endpointOnVendorDomain": true,
        "terms": "https://prisminference.com/terms",
        "privacy": "https://prisminference.com/privacy",
        "statusPage": "https://status.prisminference.com",
        "changelog": "https://docs.prisminference.com/changelog",
        "securityTxt": "valid",
        "checked": "2026-10-08",
        "notes": [
          "The terms and the privacy policy, both last updated 9 September 2026, name Prism Technologies Inc. The terms are governed by California law with arbitration in San Francisco.",
          "RDAP gives a registration date of 2026-09-09 for prisminference.com, with Name.com as registrar.",
          "security.txt has a Contact line (founders@prisminference.com), an Expires date of 2027-10-06 and a Canonical line. No disclosure policy or bug bounty is named.",
          "The status page runs on incident.io with one component for each model and no component for the API or the website. Its incidents feed was empty on 8 October 2026.",
          "The YC directory lists Prism in the Spring 2025 batch, in San Francisco, with a team of 2. The home page links an X account named prism_videos, from the company's earlier video product."
        ],
        "score": 79
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/prism-inference.json",
      "live": {
        "slug": "prism-inference",
        "probe": {
          "target": "https://api.prisminference.com/v1",
          "method": "get",
          "lastAt": "2026-10-09T10:42:53.406705714Z",
          "lastOk": true,
          "lastStatus": 200,
          "lastMs": 162,
          "authRequired": false,
          "uptime24h": 100,
          "uptime30d": 100,
          "p50ms24h": 139,
          "p95ms24h": 443,
          "samples24h": 207,
          "samples30d": 207,
          "days": [
            {
              "date": "2026-10-08",
              "probes": 93,
              "ok": 93
            },
            {
              "date": "2026-10-09",
              "probes": 114,
              "ok": 114
            }
          ]
        },
        "vendorStatus": {
          "page": "https://status.prisminference.com",
          "indicator": "none",
          "summary": "All Systems Operational",
          "checkedAt": "2026-10-09T10:41:59.970329293Z"
        },
        "versions": [
          {
            "registry": "github",
            "name": "prismhq/hermes-prism-provider",
            "version": "v1.0.2",
            "released": "2026-09-16",
            "seenAt": "2026-10-08T16:26:20.638377705Z"
          }
        ],
        "githubStars": 0,
        "securityTxt": {
          "url": "https://prisminference.com/.well-known/security.txt",
          "state": "valid",
          "expires": "2027-10-06T00:00:00.000Z",
          "checkedAt": "2026-10-08T15:38:58.310470594Z"
        },
        "pages": [
          {
            "url": "https://docs.prisminference.com/changelog",
            "kind": "changelog",
            "status": 200,
            "checkedAt": "2026-10-08T18:19:15.357311357Z",
            "changedAt": "0001-01-01T00:00:00Z",
            "fingerprint": "184a0fafdb8c"
          },
          {
            "url": "https://prisminference.com/pricing",
            "kind": "pricing",
            "status": 200,
            "checkedAt": "2026-10-08T18:23:20.103073255Z",
            "changedAt": "0001-01-01T00:00:00Z",
            "fingerprint": "dffb5d16a3c8"
          },
          {
            "url": "https://prisminference.com/privacy",
            "kind": "privacy",
            "status": 200,
            "checkedAt": "2026-10-08T18:23:22.266572016Z",
            "changedAt": "0001-01-01T00:00:00Z",
            "fingerprint": "64ea7b8f1b7b"
          },
          {
            "url": "https://prisminference.com/terms",
            "kind": "terms",
            "status": 200,
            "checkedAt": "2026-10-08T18:23:24.296415884Z",
            "changedAt": "0001-01-01T00:00:00Z",
            "fingerprint": "27dd15c69f49"
          }
        ],
        "updatedAt": "2026-10-09T10:42:53.406705714Z"
      }
    },
    "facts": [
      {
        "a": "Model API",
        "b": "Model API",
        "name": "Kind"
      },
      {
        "a": "Deep Infra Inc.",
        "b": "Prism Technologies Inc",
        "name": "Vendor"
      },
      {
        "a": "https://api.deepinfra.com/v1/openai",
        "b": "https://api.prisminference.com/v1",
        "name": "Hosted endpoint"
      },
      {
        "a": "HTTP",
        "b": "HTTP",
        "name": "Transports"
      },
      {
        "a": "API key",
        "b": "API key",
        "name": "Auth"
      },
      {
        "a": "Pay per use",
        "b": "Pay per use",
        "name": "Pricing"
      },
      {
        "a": "no",
        "b": "no",
        "name": "x402"
      },
      {
        "a": "Proprietary service under the DeepInfra Terms of Service. The Python and Node SDKs and the docs repository are MIT",
        "b": "Proprietary service under Prism's terms of service. The OpenAPI file declares `LicenseRef-Proprietary`. The Hermes provider plugin repository carries no licence file",
        "name": "Licence"
      },
      {
        "a": "no",
        "b": "no",
        "name": "Read-only variant documented"
      },
      {
        "a": "yes",
        "b": "yes",
        "name": "llms.txt"
      },
      {
        "a": "2026-10-07",
        "b": "2026-10-06",
        "name": "Last release"
      },
      {
        "a": "2026-08-17",
        "b": "2026-09-09",
        "name": "Terms last updated"
      },
      {
        "a": "2026-08-15",
        "b": "2026-09-09",
        "name": "Privacy policy last updated"
      },
      {
        "a": "not found in the text",
        "b": "not found in the text",
        "name": "Customer content may train models"
      },
      {
        "a": "not found in the text",
        "b": "yes",
        "name": "Terms restrict automated access"
      },
      {
        "a": "yes",
        "b": "not found in the text",
        "name": "Terms restrict benchmarking"
      },
      {
        "a": "not found in the text",
        "b": "not found in the text",
        "name": "Terms or service can change without notice"
      },
      {
        "a": "yes",
        "b": "yes",
        "name": "Arbitration or class-action waiver"
      },
      {
        "a": "21 stars, 1.2k npm/wk, 50 PyPI/wk",
        "b": "none",
        "name": "Popularity"
      }
    ],
    "faq": [
      {
        "answer": "DeepInfra scores 63 (B) on agent readiness against Prism Inference's 60.1 (C), and leads in 4 of 7 scored categories. Prism Inference leads on schema \u0026 documentation and payments \u0026 pricing.",
        "question": "Which is better for AI agents, DeepInfra or Prism Inference?"
      },
      {
        "answer": "Both need an API key.",
        "question": "Do DeepInfra and Prism Inference need an API key?"
      },
      {
        "answer": "Yes. DeepInfra has a hosted endpoint at https://api.deepinfra.com/v1/openai and Prism Inference at https://api.prisminference.com/v1.",
        "question": "Can an agent call DeepInfra and Prism Inference without installing anything?"
      }
    ],
    "goodFor": [
      {
        "aheadOn": [
          "Reliability, 70 against 65",
          "Agent ergonomics, 77 against 68",
          "Maintenance \u0026 community, 65 against 49",
          "Transparency \u0026 trust, 67 against 61"
        ],
        "also": null,
        "goodFor": "Agents that want many open-weight models, embeddings, image and speech behind one OpenAI-style key at low per-token prices, with spend-capped tokens.",
        "slug": "deepinfra",
        "watchFor": "A deprecated model gets at least one week's notice, and requests are then forwarded to a replacement model under the old id"
      },
      {
        "aheadOn": [
          "Schema \u0026 documentation, 82 against 69",
          "Payments \u0026 pricing, 30 against 20"
        ],
        "also": null,
        "goodFor": "Coding agents that want DeepSeek-V4.1-Flash at a low input price, with no retention, through whichever of the three wire formats the harness already speaks.",
        "slug": "prism-inference",
        "watchFor": "Two models. The docs mark Gemma 4 31B as request access per organisation, while llms.txt and the keyless catalogue list it as available"
      }
    ],
    "job": {
      "capability": "inference.llm",
      "name": "LLM inference"
    },
    "others": [
      {
        "json": "https://www.anchorterminal.com/compare/anthropic-api-vs-deepinfra.json",
        "title": "Claude API vs DeepInfra",
        "url": "https://www.anchorterminal.com/compare/anthropic-api-vs-deepinfra"
      },
      {
        "json": "https://www.anchorterminal.com/compare/anthropic-api-vs-prism-inference.json",
        "title": "Claude API vs Prism Inference",
        "url": "https://www.anchorterminal.com/compare/anthropic-api-vs-prism-inference"
      },
      {
        "json": "https://www.anchorterminal.com/compare/antseed-vs-deepinfra.json",
        "title": "Antseed vs DeepInfra",
        "url": "https://www.anchorterminal.com/compare/antseed-vs-deepinfra"
      },
      {
        "json": "https://www.anchorterminal.com/compare/antseed-vs-prism-inference.json",
        "title": "Antseed vs Prism Inference",
        "url": "https://www.anchorterminal.com/compare/antseed-vs-prism-inference"
      },
      {
        "json": "https://www.anchorterminal.com/compare/blockrun-ai-vs-deepinfra.json",
        "title": "BlockRun.AI vs DeepInfra",
        "url": "https://www.anchorterminal.com/compare/blockrun-ai-vs-deepinfra"
      },
      {
        "json": "https://www.anchorterminal.com/compare/blockrun-ai-vs-prism-inference.json",
        "title": "BlockRun.AI vs Prism Inference",
        "url": "https://www.anchorterminal.com/compare/blockrun-ai-vs-prism-inference"
      },
      {
        "json": "https://www.anchorterminal.com/compare/deepinfra-vs-deepseek-api.json",
        "title": "DeepInfra vs DeepSeek API",
        "url": "https://www.anchorterminal.com/compare/deepinfra-vs-deepseek-api"
      },
      {
        "json": "https://www.anchorterminal.com/compare/deepinfra-vs-gemini-api.json",
        "title": "DeepInfra vs Gemini Developer API",
        "url": "https://www.anchorterminal.com/compare/deepinfra-vs-gemini-api"
      },
      {
        "json": "https://www.anchorterminal.com/compare/deepinfra-vs-groq.json",
        "title": "DeepInfra vs GroqCloud",
        "url": "https://www.anchorterminal.com/compare/deepinfra-vs-groq"
      },
      {
        "json": "https://www.anchorterminal.com/compare/deepinfra-vs-mistral-api.json",
        "title": "DeepInfra vs Mistral AI API",
        "url": "https://www.anchorterminal.com/compare/deepinfra-vs-mistral-api"
      },
      {
        "json": "https://www.anchorterminal.com/compare/deepinfra-vs-openai-api.json",
        "title": "DeepInfra vs OpenAI API",
        "url": "https://www.anchorterminal.com/compare/deepinfra-vs-openai-api"
      },
      {
        "json": "https://www.anchorterminal.com/compare/deepinfra-vs-openrouter.json",
        "title": "DeepInfra vs OpenRouter",
        "url": "https://www.anchorterminal.com/compare/deepinfra-vs-openrouter"
      },
      {
        "json": "https://www.anchorterminal.com/compare/deepinfra-vs-sambanova.json",
        "title": "DeepInfra vs SambaCloud",
        "url": "https://www.anchorterminal.com/compare/deepinfra-vs-sambanova"
      },
      {
        "json": "https://www.anchorterminal.com/compare/deepseek-api-vs-prism-inference.json",
        "title": "DeepSeek API vs Prism Inference",
        "url": "https://www.anchorterminal.com/compare/deepseek-api-vs-prism-inference"
      },
      {
        "json": "https://www.anchorterminal.com/compare/gemini-api-vs-prism-inference.json",
        "title": "Gemini Developer API vs Prism Inference",
        "url": "https://www.anchorterminal.com/compare/gemini-api-vs-prism-inference"
      },
      {
        "json": "https://www.anchorterminal.com/compare/groq-vs-prism-inference.json",
        "title": "GroqCloud vs Prism Inference",
        "url": "https://www.anchorterminal.com/compare/groq-vs-prism-inference"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mistral-api-vs-prism-inference.json",
        "title": "Mistral AI API vs Prism Inference",
        "url": "https://www.anchorterminal.com/compare/mistral-api-vs-prism-inference"
      },
      {
        "json": "https://www.anchorterminal.com/compare/openai-api-vs-prism-inference.json",
        "title": "OpenAI API vs Prism Inference",
        "url": "https://www.anchorterminal.com/compare/openai-api-vs-prism-inference"
      },
      {
        "json": "https://www.anchorterminal.com/compare/openrouter-vs-prism-inference.json",
        "title": "OpenRouter vs Prism Inference",
        "url": "https://www.anchorterminal.com/compare/openrouter-vs-prism-inference"
      },
      {
        "json": "https://www.anchorterminal.com/compare/prism-inference-vs-sambanova.json",
        "title": "Prism Inference vs SambaCloud",
        "url": "https://www.anchorterminal.com/compare/prism-inference-vs-sambanova"
      }
    ],
    "scores": [
      {
        "by": 5,
        "deepinfra": 70,
        "edge": "deepinfra",
        "key": "reliability",
        "name": "Reliability",
        "prism-inference": 65,
        "weight": 16
      },
      {
        "key": "performance",
        "name": "Performance",
        "pending": true,
        "weight": 10
      },
      {
        "by": 13,
        "deepinfra": 69,
        "edge": "prism-inference",
        "key": "schema",
        "name": "Schema \u0026 documentation",
        "prism-inference": 82,
        "weight": 13
      },
      {
        "by": 9,
        "deepinfra": 77,
        "edge": "deepinfra",
        "key": "ergonomics",
        "name": "Agent ergonomics",
        "prism-inference": 68,
        "weight": 13
      },
      {
        "by": 1,
        "deepinfra": 64,
        "edge": "prism-inference",
        "key": "security",
        "name": "Security \u0026 auth",
        "prism-inference": 65,
        "weight": 14
      },
      {
        "by": 10,
        "deepinfra": 20,
        "edge": "prism-inference",
        "key": "payments",
        "name": "Payments \u0026 pricing",
        "prism-inference": 30,
        "weight": 10
      },
      {
        "key": "tasks",
        "name": "Task success",
        "pending": true,
        "weight": 10
      },
      {
        "by": 16,
        "deepinfra": 65,
        "edge": "deepinfra",
        "key": "maintenance",
        "name": "Maintenance \u0026 community",
        "prism-inference": 49,
        "weight": 7
      },
      {
        "by": 6,
        "deepinfra": 67,
        "edge": "deepinfra",
        "key": "transparency",
        "name": "Transparency \u0026 trust",
        "prism-inference": 61,
        "weight": 7
      }
    ],
    "summary": "DeepInfra scores 63 (B) on agent readiness against Prism Inference's 60.1 (C), and leads in 4 of 7 scored categories. Prism Inference leads on schema \u0026 documentation and payments \u0026 pricing. Both do llm inference.",
    "verdicts": {
      "deepinfra": "The model list, context sizes and per-token prices are readable without a key, and keys can carry an IP allowlist, a monthly spending cap and model-limited JWTs. Deprecated models get one week's notice and are then redirected to another model, there is no changelog or SLA, and an account needs a card or prepayment before any call.",
      "prism-inference": "Three wire formats, a public OpenAPI 3.1 file, per-token prices in a keyless catalogue and zero data retention by default on every tier. The service launched on 24 September 2026 with two models, one of them by request, from a two-person company. No rate-limit numbers, SLA document, free tier or deprecation policy was found."
    }
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/compare/deepinfra-vs-prism-inference",
    "json": "https://www.anchorterminal.com/compare/deepinfra-vs-prism-inference.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/compare/deepinfra-vs-prism-inference.md",
    "slim": "https://www.anchorterminal.com/compare/deepinfra-vs-prism-inference.min.md"
  },
  "markdown": "DeepInfra scores 63 (B) on agent readiness against Prism Inference's 60.1 (C), and leads in 4 of 7 scored categories. Prism Inference leads on schema \u0026 documentation and payments \u0026 pricing. Both do llm inference.\n\n- DeepInfra: grade B, 63/100, rank #371 of 842. Markdown https://www.anchorterminal.com/tools/deepinfra.md · JSON https://www.anchorterminal.com/api/v1/tools/deepinfra.json\n- Prism Inference: grade C, 60.1/100, rank #479 of 842. Markdown https://www.anchorterminal.com/tools/prism-inference.md · JSON https://www.anchorterminal.com/api/v1/tools/prism-inference.json\n\n## Which one, for what\n\n### DeepInfra (B)\n\nGood for: Agents that want many open-weight models, embeddings, image and speech behind one OpenAI-style key at low per-token prices, with spend-capped tokens.\n\nAhead on:\n- Reliability, 70 against 65\n- Agent ergonomics, 77 against 68\n- Maintenance \u0026 community, 65 against 49\n- Transparency \u0026 trust, 67 against 61\n\nWatch for: A deprecated model gets at least one week's notice, and requests are then forwarded to a replacement model under the old id\n\n### Prism Inference (C)\n\nGood for: Coding agents that want DeepSeek-V4.1-Flash at a low input price, with no retention, through whichever of the three wire formats the harness already speaks.\n\nAhead on:\n- Schema \u0026 documentation, 82 against 69\n- Payments \u0026 pricing, 30 against 20\n\nWatch for: Two models. The docs mark Gemma 4 31B as request access per organisation, while llms.txt and the keyless catalogue list it as available\n\n\n## Score by category\n\n| Category | Weight | DeepInfra | Prism Inference | Edge |\n| --- | --- | --- | --- | --- |\n| Reliability | 16% (20 this run) | 70 | 65 | DeepInfra +5 |\n| Performance | 10%, pending | pending | pending | not scored in this run |\n| Schema \u0026 documentation | 13% (16.2 this run) | 69 | 82 | Prism Inference +13 |\n| Agent ergonomics | 13% (16.2 this run) | 77 | 68 | DeepInfra +9 |\n| Security \u0026 auth | 14% (17.5 this run) | 64 | 65 | Prism Inference +1 |\n| Payments \u0026 pricing | 10% (12.5 this run) | 20 | 30 | Prism Inference +10 |\n| Task success | 10%, pending | pending | pending | not scored in this run |\n| Maintenance \u0026 community | 7% (8.8 this run) | 65 | 49 | DeepInfra +16 |\n| Transparency \u0026 trust | 7% (8.8 this run) | 67 | 61 | DeepInfra +6 |\n| Negative events | ≤15 | 0 | -2 | |\n| **Total** | | **63 · B** | **60.1 · C** | |\n\n## Facts side by side\n\n| Fact | DeepInfra | Prism Inference |\n| --- | --- | --- |\n| Kind | Model API | Model API |\n| Vendor | Deep Infra Inc. | Prism Technologies Inc |\n| Hosted endpoint | `https://api.deepinfra.com/v1/openai` | `https://api.prisminference.com/v1` |\n| Transports | HTTP | HTTP |\n| Auth | API key | API key |\n| Pricing | Pay per use | Pay per use |\n| x402 | no | no |\n| Licence | Proprietary service under the DeepInfra Terms of Service. The Python and Node SDKs and the docs repository are MIT | Proprietary service under Prism's terms of service. The OpenAPI file declares `LicenseRef-Proprietary`. The Hermes provider plugin repository carries no licence file |\n| Read-only variant documented | no | no |\n| llms.txt | yes | yes |\n| Last release | 2026-10-07 | 2026-10-06 |\n| Terms last updated | 2026-08-17 | 2026-09-09 |\n| Privacy policy last updated | 2026-08-15 | 2026-09-09 |\n| Customer content may train models | not found in the text | not found in the text |\n| Terms restrict automated access | not found in the text | yes |\n| Terms restrict benchmarking | yes | not found in the text |\n| Terms or service can change without notice | not found in the text | not found in the text |\n| Arbitration or class-action waiver | yes | yes |\n| Popularity | 21 stars, 1.2k npm/wk, 50 PyPI/wk | none |\n\n## Verdicts\n\n**DeepInfra.** The model list, context sizes and per-token prices are readable without a key, and keys can carry an IP allowlist, a monthly spending cap and model-limited JWTs. Deprecated models get one week's notice and are then redirected to another model, there is no changelog or SLA, and an account needs a card or prepayment before any call.\n\n**Prism Inference.** Three wire formats, a public OpenAPI 3.1 file, per-token prices in a keyless catalogue and zero data retention by default on every tier. The service launched on 24 September 2026 with two models, one of them by request, from a two-person company. No rate-limit numbers, SLA document, free tier or deprecation policy was found.\n\n## Before you call either\n\n### DeepInfra\n\n1. Call `GET https://api.deepinfra.com/v1/openai/models` at start-up for ids, context sizes and prices. No key is needed\n2. Check the `model` field of each response. After a deprecation date, requests to the old id are served by a replacement model\n3. Ask the account owner for a scoped JWT limited to the models and spend the task needs, not the full API key\n4. Stay under 200 concurrent requests per model. On 429 `engine_overloaded`, retry after a delay, or send `models` with up to four fallbacks\n5. Never inspect a JWT with `GET /v1/scoped-jwt?jwtoken=`, which puts the token in the URL. Keep credentials in the `Authorization` header\n\n### Prism Inference\n\n1. Call `GET https://api.prisminference.com/v1/models` at start-up, with no key, and use only ids it returns. Expect 403 on `gemma-4-31b` without organisation access\n2. Use base URL `https://api.prisminference.com/v1` for OpenAI clients and `https://api.prisminference.com` with no `/v1` for Anthropic clients\n3. Read `error.retryable` before retrying, and wait for `Retry-After` on 429, which covers both key limits and model capacity\n4. Send `reasoning_effort: \"none\"` or `low` when latency matters. Reasoning is on by default and its tokens are billed as output\n5. Keep conversation state yourself and send `store: false` on Responses. `previous_response_id`, stored responses and hosted tools aren't supported\n\n## Questions\n\n### Which is better for AI agents, DeepInfra or Prism Inference?\n\nDeepInfra scores 63 (B) on agent readiness against Prism Inference's 60.1 (C), and leads in 4 of 7 scored categories. Prism Inference leads on schema \u0026 documentation and payments \u0026 pricing.\n\n### Do DeepInfra and Prism Inference need an API key?\n\nBoth need an API key.\n\n### Can an agent call DeepInfra and Prism Inference without installing anything?\n\nYes. DeepInfra has a hosted endpoint at https://api.deepinfra.com/v1/openai and Prism Inference at https://api.prisminference.com/v1.\n\n\n## For agents\n\n- This comparison as JSON: https://www.anchorterminal.com/compare/deepinfra-vs-prism-inference.json, and with the fewest tokens: https://www.anchorterminal.com/compare/deepinfra-vs-prism-inference.min.md\n- Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {\"a\": \"deepinfra\", \"b\": \"prism-inference\"}`. From a terminal: `anchor compare deepinfra prism-inference`\n- Each listing in full: https://www.anchorterminal.com/api/v1/tools/deepinfra.json and https://www.anchorterminal.com/api/v1/tools/prism-inference.json\n\n## Other comparisons with DeepInfra or Prism Inference\n\n- [Claude API vs DeepInfra](https://www.anchorterminal.com/compare/anthropic-api-vs-deepinfra.md)\n- [Claude API vs Prism Inference](https://www.anchorterminal.com/compare/anthropic-api-vs-prism-inference.md)\n- [Antseed vs DeepInfra](https://www.anchorterminal.com/compare/antseed-vs-deepinfra.md)\n- [Antseed vs Prism Inference](https://www.anchorterminal.com/compare/antseed-vs-prism-inference.md)\n- [BlockRun.AI vs DeepInfra](https://www.anchorterminal.com/compare/blockrun-ai-vs-deepinfra.md)\n- [BlockRun.AI vs Prism Inference](https://www.anchorterminal.com/compare/blockrun-ai-vs-prism-inference.md)\n- [DeepInfra vs DeepSeek API](https://www.anchorterminal.com/compare/deepinfra-vs-deepseek-api.md)\n- [DeepInfra vs Gemini Developer API](https://www.anchorterminal.com/compare/deepinfra-vs-gemini-api.md)\n- [DeepInfra vs GroqCloud](https://www.anchorterminal.com/compare/deepinfra-vs-groq.md)\n- [DeepInfra vs Mistral AI API](https://www.anchorterminal.com/compare/deepinfra-vs-mistral-api.md)\n- [DeepInfra vs OpenAI API](https://www.anchorterminal.com/compare/deepinfra-vs-openai-api.md)\n- [DeepInfra vs OpenRouter](https://www.anchorterminal.com/compare/deepinfra-vs-openrouter.md)\n- [DeepInfra vs SambaCloud](https://www.anchorterminal.com/compare/deepinfra-vs-sambanova.md)\n- [DeepSeek API vs Prism Inference](https://www.anchorterminal.com/compare/deepseek-api-vs-prism-inference.md)\n- [Gemini Developer API vs Prism Inference](https://www.anchorterminal.com/compare/gemini-api-vs-prism-inference.md)\n- [GroqCloud vs Prism Inference](https://www.anchorterminal.com/compare/groq-vs-prism-inference.md)\n- [Mistral AI API vs Prism Inference](https://www.anchorterminal.com/compare/mistral-api-vs-prism-inference.md)\n- [OpenAI API vs Prism Inference](https://www.anchorterminal.com/compare/openai-api-vs-prism-inference.md)\n- [OpenRouter vs Prism Inference](https://www.anchorterminal.com/compare/openrouter-vs-prism-inference.md)\n- [Prism Inference vs SambaCloud](https://www.anchorterminal.com/compare/prism-inference-vs-sambanova.md)\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-09",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "Compare",
        "url": "https://www.anchorterminal.com/compare/"
      },
      {
        "name": "DeepInfra vs Prism Inference",
        "url": ""
      }
    ],
    "description": "DeepInfra scores 63 (B) on agent readiness against Prism Inference's 60.1 (C), and leads in 4 of 7 scored categories. Prism Inference leads on schema \u0026 documentation and payments \u0026 pricing. Both do llm inference. Category scores, facts, verdicts and agent notes side by side.",
    "facts": [
      "DeepInfra B 63",
      "Prism Inference C 60.1",
      "scores"
    ],
    "h1": "DeepInfra vs Prism Inference",
    "image": "https://www.anchorterminal.com/assets/og/compare-deepinfra-vs-prism-inference.png",
    "path": "/compare/deepinfra-vs-prism-inference",
    "published": "2026-10-01",
    "section": "tools",
    "title": "DeepInfra vs Prism Inference for AI agents, B 63 vs C 60.1",
    "toc": null,
    "updated": "2026-10-09",
    "url": "https://www.anchorterminal.com/compare/deepinfra-vs-prism-inference"
  },
  "tokens": {
    "markdown": 2450,
    "slim": 680
  },
  "version": 1
}
