# Prism Inference (slim) > Prism is a hosted inference API from Prism Technologies Inc for open-weight models, aimed at coding agents. It accepts OpenAI Chat Completions, OpenAI Responses and Anthropic Messages requests at api.prisminference.com. It launched on 24 September 2026. - Full: https://www.anchorterminal.com/tools/prism-inference.md (~8,050 tokens) · this version ~1,830 tokens · JSON https://www.anchorterminal.com/tools/prism-inference.json · canonical https://www.anchorterminal.com/tools/prism-inference - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-08 **C · 60.1/100 · rank #363 of 629 · #8 in Model APIs & inference · not agent-ready · confidence medium** Assessment: Three wire formats, a public OpenAPI 3.1 file, per-token prices in a keyless catalogue and zero data retention by default on every tier. The service launched on 24 September 2026 with two models, one of them by request, from a two-person company. No rate-limit numbers, SLA document, free tier or deprecation policy was found. ## Facts - Kind: Model API · vendor: Prism Technologies Inc · category: Model APIs & inference · legal entity: Prism Technologies Inc · provenance 79/100 - Endpoint: `https://api.prisminference.com/v1` (HTTP) - Auth: API key · pricing: Pay per use · x402: no · licence: Proprietary service under Prism's terms of service. The OpenAPI file declares `LicenseRef-Proprietary`. The Hermes provider plugin repository carries no licence file - Probe metrics: not measured yet (probes haven't run) - Age: Launched 24 September 2026. Domain registered 9 September 2026. OpenAPI file at version 0.1.0. One changelog entry, 6 October 2026 - Endpoints: POST /v1/chat/completions, /v1/responses, /v1/messages and /v1/messages/count_tokens. GET /v1/models, /v1/models/{model} and /health - Models: `deepseek-v4.1-flash` (text and image input, 1M context, 384,000 output tokens) and `gemma-4-31b` (text and image, 32,768 context, 8,192 output, BF16, request access per the docs) - Free tier: None found. Prepaid credit, 402 when the workspace has none - Trains on API data: No, per the privacy policy, the terms and the docs - Data retention: Zero retention of inputs and outputs by default on every tier. Service metadata (key id, model, token counts, status, latency, source IP) kept for as long as reasonably needed, with no period stated - Data location: Prism is based in the United States and may process data there and in other countries. No region list or sub-processor list found - Rate limits: Per key, no numbers published. 429 with `Retry-After` in seconds, also used when a model is at capacity. Requests are capped at 4 MiB - Errors: OpenAI-style envelope with `code`, `retryable`, `fix` and `docs_url`. 13 status codes documented with retry advice - Structured output: JSON mode and strict JSON Schema on Chat Completions, `text.format` on Responses - Reasoning: On by default. `reasoning_effort` none, low, medium or high, and reasoning tokens are billed as output - Speed: 550 tokens a second on DeepSeek-V4.1-Flash per the vendor's home page (547 in the launch post). Not measured by us - SDKs: None of its own. OpenAI and Anthropic client libraries, plus provider plugins for Hermes and OpenClaw and setup guides for Claude Code, Codex, Cursor and OpenCode - SLA: None published. The terms say no guaranteed uptime without a separate written agreement, and dedicated deployments list a negotiated latency SLA - Status: status.prisminference.com on incident.io, one component per model, no incidents posted - Company: Prism Technologies Inc, San Francisco, Y Combinator Spring 2025, team of 2 per the YC directory - Models ($/1M in/out): `deepseek-v4.1-flash` $0.09/$1.20; `gemma-4-31b` $0.30/$0.40 - Scores: Reliability 65, Performance pending, Schema & documentation 82, Agent ergonomics 68, Security & auth 65, Payments & pricing 30, Task success pending, Maintenance & community 49, Transparency & trust 61 · negative events -2 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, A public OpenAPI 3.1 file covering ten paths, including chat completions, responses, messages and models (25). · Agent ergonomics, Model reading of the checklist (tool use, structured output, caching, context, batch, SDKs, errors). · Security & auth, Model reading. · Payments & pricing, No machine payment protocol (0). · Maintenance & community, Model reading. · Transparency & trust, Closed service with clear terms that name Prism Technologies Inc (15). - Sources: 31, open questions: 9, both in the full twin - Capabilities: inference.fast, inference.open-weights, inference.llm - JSON: https://www.anchorterminal.com/api/v1/tools/prism-inference.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/prism-inference.svg` or a link to https://www.anchorterminal.com/tools/prism-inference from a page on prisminference.com or one of its subdomains, or the README of github.com/prismhq/hermes-prism-provider, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Call `GET https://api.prisminference.com/v1/models` at start-up, with no key, and use only ids it returns. Expect 403 on `gemma-4-31b` without organisation access 2. Use base URL `https://api.prisminference.com/v1` for OpenAI clients and `https://api.prisminference.com` with no `/v1` for Anthropic clients 3. Read `error.retryable` before retrying, and wait for `Retry-After` on 429, which covers both key limits and model capacity 4. Send `reasoning_effort: "none"` or `low` when latency matters. Reasoning is on by default and its tokens are billed as output 5. Keep conversation state yourself and send `store: false` on Responses. `previous_response_id`, stored responses and hosted tools aren't supported ## Connect ```bash pip install openai # or: npm install openai, base URL https://api.prisminference.com/v1. Anthropic SDKs use https://api.prisminference.com with no /v1 ``` ```bash curl "https://api.prisminference.com/v1/chat/completions" \ -H "Authorization: Bearer $PRISM_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"deepseek-v4.1-flash","messages":[{"role":"user","content":"Return pong."}]}' ``` ```bash export ANTHROPIC_BASE_URL=https://api.prisminference.com export ANTHROPIC_AUTH_TOKEN=$PRISM_API_KEY export ANTHROPIC_MODEL=deepseek-v4.1-flash export ANTHROPIC_SMALL_FAST_MODEL=gemma-4-31b ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | GroqCloud | BB | 75.6 | inference.llm, inference.fast, inference.open-weights | https://www.anchorterminal.com/tools/groq.min.md | | Mistral AI API | BB | 71.2 | inference.llm, inference.open-weights | https://www.anchorterminal.com/tools/mistral-api.min.md | | Docker Model Runner | C | 57.1 | inference.open-weights, inference.llm | https://www.anchorterminal.com/tools/docker-model-runner.min.md | | Ollama | C | 56.3 | inference.open-weights, inference.llm | https://www.anchorterminal.com/tools/ollama.min.md | | DeepSeek API | D | 46.8 | inference.llm, inference.open-weights | https://www.anchorterminal.com/tools/deepseek-api.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)