# SambaCloud (slim) > SambaCloud is SambaNova's hosted inference API for open-weight models running on its own RDU processors. It answers OpenAI-style chat, completions and Responses calls and Anthropic-style Messages calls, with Python and TypeScript SDKs. - Full: https://www.anchorterminal.com/tools/sambanova.md (~8,000 tokens) · this version ~1,980 tokens · JSON https://www.anchorterminal.com/tools/sambanova.json · canonical https://www.anchorterminal.com/tools/sambanova - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **B · 66.4/100 · rank #269 of 842 · #7 in Model APIs & inference · not agent-ready · confidence medium** Assessment: Per-token prices and the model list are readable without a key at `/v1/models`, and a free tier needs no card. Production models get two to three weeks' notice before removal, 13 model ids left between March and June 2026, and the docs disagree with the live catalogue on MiniMax-M2.7. ## Facts - Kind: Model API · vendor: SambaNova Systems, Inc. · category: Model APIs & inference · legal entity: SambaNova Systems, Inc. · provenance 81/100 - Endpoint: `https://api.sambanova.ai/v1` (HTTP) - Auth: API key · pricing: Freemium · x402: no · licence: Proprietary service under the SambaCloud Terms of Service. The SDKs and the OpenAPI document are Apache-2.0 - Probe metrics: not measured yet (probes haven't run) - Endpoints: `/chat/completions`, `/completions`, `/responses`, `/messages`, `/messages/count_tokens`, `/embeddings`, `/audio/transcriptions`, `/audio/translations`, `/models` and `/models/{model_id}` under https://api.sambanova.ai/v1, per the OpenAPI document (version 1.2.0) - Models on 8 October 2026: Six ids from `/v1/models` and the pricing page. gpt-oss-120b, DeepSeek-V3.1, Meta-Llama-3.3-70B-Instruct (production), and DeepSeek-V3.2, gemma-4-31B-it, MiniMax-M3 (preview). No embedding or audio model is listed - Free tier: No payment method. 20 requests a minute, 20 requests a day and 200,000 tokens a day per model. MiniMax models are not in the free tier table - Developer tier: Applies once a payment method is linked. 60 requests a minute and 12,000 a day on most models, 240 and 48,000 on Llama 3.3 70B, and 20M tokens a day across all models - Prompt caching: Automatic on MiniMax-M3 and MiniMax-M2.7 for prefixes of 4,096 to 192,000 tokens. MiniMax-M3 cached input is $0.06 per 1M tokens against $0.60. Cache is per serving node, evicted least recently used, and not guaranteed - Structured output: `response_format` takes `json_object` and `json_schema`. Schema matching is best effort, and `strict: true` is accepted with no effect - Responses API: Stateless. `previous_response_id` and built-in or server-executed tools are not supported - Errors: OpenAI-style `error` object with `message`, `type`, `param` and `code`, plus `request_id`. 13 documented codes, among them `context_length_exceeded`, `model_deprecated` (410), `insufficient_quota` and `queue_full` (429) and `maintenance` (503) - Credentials: Bearer API key, or the same key in `x-api-key` on the Messages routes. Up to 25 keys per user, each shown once at creation - Deprecation notice: Production models at least two to three weeks, preview models at short notice, by email and on the deprecations page - Data use in the terms: SambaNova may access, store and process Customer Content solely to supply the service or as law requires. It may use query logs and Service Usage Data, defined as excluding Customer Content, to improve its products - Certifications: SOC 2 Type 2 and ISO 27001, per the trust centre's page description. The reports and the rest of the page need JavaScript and were not read - SDKs: Python `sambanova` 1.13.0 and TypeScript `sambanova` 1.10.0, both published 17 September 2026, Apache-2.0, generated by Stainless. 429 and server errors are retried twice with backoff by default - Other routes: SaaS listing on AWS Marketplace with PrivateLink, and SambaStack for self-managed or hosted private deployments under separate terms - Models ($/1M in/out): `gpt-oss-120b` $0.22/$0.59; `DeepSeek-V3.1` $3/$4.50; `Meta-Llama-3.3-70B-Instruct` $0.60/$1.20 - 2026-03-09 Shutdown: `Whisper-Large-v3` and `ALLaM-7B` removed - 2026-03-20 Shutdown: `DeepSeek-R1-Distill-Llama-70B` removed - 2026-04-06 Shutdown: `DeepSeek-V3.1-Terminus`, `Qwen3-235B-A22B-Instruct-2507`, `Qwen3-32B` and `E5-Mistral-7B-Instruct` removed - 2026-04-14 Shutdown: `DeepSeek-V3-0324`, `DeepSeek-R1-0528` and `Meta-Llama-3.1-8B-Instruct` removed - 2026-05-18 Shutdown: `MiniMax-M2.5` removed from SambaCloud - 2026-06-09 Shutdown: `gemma-3-12b-it` and `Llama-4-Maverick-17B-128E-Instruct` removed - Scores: Reliability 85, Performance pending, Schema & documentation 81, Agent ergonomics 67, Security & auth 58, Payments & pricing 40, Task success pending, Maintenance & community 77, Transparency & trust 63 · negative events -2 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, Public OpenAPI 3.1.1 document, version 1.2.0, with 10 operations and 169 schemas, last updated 17 September 2026 (25). · Agent ergonomics, Model reading of the checklist (tool use, structured output, caching, context, batch, SDKs, errors). · Security & auth, Model reading. · Payments & pricing, No machine payment protocol (0). · Maintenance & community, Model reading. · Transparency & trust, Closed service with clear terms that name the entity and the governing law, and Apache-2.0 SDKs and OpenAPI document (15). - Sources: 28, open questions: 7, both in the full twin - Capabilities: inference.llm, inference.open-weights, inference.fast - JSON: https://www.anchorterminal.com/api/v1/tools/sambanova.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/sambanova.svg` or a link to https://www.anchorterminal.com/tools/sambanova from a page on sambanova.ai or one of its subdomains, or the README of github.com/sambanova/sambanova-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Call `GET https://api.sambanova.ai/v1/models` at start-up and use only ids it returns. The docs name models the endpoint no longer lists 2. Read `x-ratelimit-remaining-requests` and `x-ratelimit-remaining-requests-day` on every response. The free tier allows 20 requests a day per model 3. Treat 429 `queue_full` and 503 `maintenance` as retryable after a delay, and 410 `model_deprecated` as a signal to change model 4. Validate JSON output yourself. Schema enforcement is best effort and `strict: true` changes nothing 5. Check `max_completion_tokens` per model. DeepSeek-V3.1 caps output at 7,168 tokens and Llama 3.3 70B at 3,072 ## Connect ```bash pip install sambanova # or: npm install sambanova ``` ```bash curl https://api.sambanova.ai/v1/chat/completions \ -H "Authorization: Bearer $SAMBANOVA_API_KEY" -H "Content-Type: application/json" \ -d '{"model":"gpt-oss-120b","messages":[{"role":"user","content":"hello"}]}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | GroqCloud | BB | 75.6 | inference.llm, inference.fast, inference.open-weights | https://www.anchorterminal.com/tools/groq.min.md | | Prism Inference | C | 60.1 | inference.fast, inference.open-weights, inference.llm | https://www.anchorterminal.com/tools/prism-inference.min.md | | Mistral AI API | BB | 71.2 | inference.llm, inference.open-weights | https://www.anchorterminal.com/tools/mistral-api.min.md | | DeepInfra | B | 63 | inference.llm, inference.open-weights | https://www.anchorterminal.com/tools/deepinfra.min.md | | Docker Model Runner | C | 57.1 | inference.open-weights, inference.llm | https://www.anchorterminal.com/tools/docker-model-runner.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)