# SiliconFlow (slim) > SiliconFlow is a hosted inference API for open-weight models, covering chat, vision, embeddings, reranking, image, video and speech. It answers OpenAI-style and Anthropic-style calls at api.siliconflow.com with a Bearer key. - Full: https://www.anchorterminal.com/tools/siliconflow.md (~9,150 tokens) · this version ~2,130 tokens · JSON https://www.anchorterminal.com/tools/siliconflow.json · canonical https://www.anchorterminal.com/tools/siliconflow - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-10 **D · 46.7/100 · rank #848 of 950 · #17 in Model APIs & inference · not agent-ready · confidence medium** Assessment: Per-token prices for every listed model are public, and one key reaches chat, embeddings, reranking, image, video and speech through OpenAI-style and Anthropic-style routes. No status page, SLA, security page or official SDK was found, release notes stop at 11 June 2026, and several model removals are dated the same day as their notice. ## Facts - Kind: Model API · vendor: SiliconFlow Labs Pte. Ltd. · category: Model APIs & inference · legal entity: SiliconFlow Labs Pte. Ltd. · provenance 68/100 - Endpoint: `https://api.siliconflow.com/v1` (HTTP) - Auth: API key · pricing: Pay per use · x402: no · licence: Proprietary service under the SiliconFlow Terms of Use - Probe metrics: not measured yet (probes haven't run) - Endpoints: 22 operations under https://api.siliconflow.com/v1 in the OpenAPI file. `/chat/completions`, `/completions`, `/messages` (Anthropic-style), `/embeddings`, `/rerank`, `/images/generations`, `/video/submit` and `/video/status`, `/audio/speech`, `/audio/transcriptions`, voice upload and list, `/files`, `/batches`, `/models`, `/user/info` and `/systemone` (alpha) - Models on 9 October 2026: The pricing page lists chat models from DeepSeek, Qwen, Z.ai, Moonshot AI, MiniMax, Tencent, Google (Gemma) and OpenAI (gpt-oss), FLUX and Z-Image for images, Wan2.2 for video and three text-to-speech models. A full count was not taken, because the tables load more rows on request - Rate limits: Per account and per model, set by monthly spend tier. L0 1,000 requests and 40,000 tokens a minute, L1 1,200 and 60,000, L2 2,000 and 80,000, L3 4,000 and 160,000, L4 8,000 and 500,000, L5 10,000 and 2,000,000. DeepSeek-R1 and DeepSeek-V3 also have 30 requests an hour. Per-model figures are in the console - Errors: JSON body with a numeric `code`, a `message` and `data`, for example `20012` for an unknown model. 400, 401, 403, 429, 500, 503 and 504 are described. The 429 message names the limit reached, such as TPM. No `Retry-After` header is documented - Tool calling: OpenAI-style `tools` with up to 128 functions. Each model page says whether Tools are supported. No `tool_choice` field is in the OpenAPI file's chat schema - Structured output: `response_format` of `json_object`. The DeepSeek-V4.1-Flash model page says JSON Mode is supported and Structured Outputs are not - Prompt caching: Cached input prices are published per model, such as $0.003 per 1M tokens on DeepSeek-V4.1-Flash. No caching guide is in the docs index - Batch: `/files` and `/batches` in the OpenAPI file, for `/v1/chat/completions` only. The field description gives a maximum window of 24 hours and a minimum of 336 hours - Credentials: API keys created in the console at cloud.siliconflow.com/account/ak, sent as a Bearer token, or in `x-api-key` on `/messages`. No scopes, expiry or per-key limits are documented. Sign-up is through a Google or GitHub account - Deprecation notices: Dated entries in the release notes. Ten model removal notices between 31 December 2025 and 11 June 2026, six of them dated the day of removal. No stated notice period - Data use in the terms: Clause 7.4 says SiliconFlow will not access Interaction Data, has no obligation to store it, will not use or disclose it without authorisation, and uses it only to supply the service. No retention period and no statement on model training were found - Generated media: Image and video URLs are valid for one hour, per the API reference - Contract: SiliconFlow Labs Pte. Ltd., Singapore law, SIAC arbitration in Singapore. Fees are not refunded once due, and a fee dispute must be raised within seven days - Support: help@siliconflow.com and contact@siliconflow.com, and a Discord server - Models ($/1M in/out): `deepseek-ai/DeepSeek-V4.1-Flash` $0.15/$0.60 - Prices: Z-Image-Turbo, one image $0.005 per image; FLUX.1-dev, one image $0.014 per image; FLUX 1.1 [pro], one image $0.04 per image; Wan2.2-T2V-A14B, one video $0.29 per call - 2026-04-29 Shutdown: `Qwen/QwQ-32B`, `IndexTeam/IndexTTS-2` and six other models removed - 2026-05-15 Shutdown: `moonshotai/Kimi-K2-Instruct`, `zai-org/GLM-4.6` and five other models removed - 2026-06-08 Shutdown: `nex-agi/DeepSeek-V3.1-Nex-N1` removed - 2026-06-11 Shutdown: `zai-org/GLM-4.7` removed, and traffic for `zai-org/GLM-5` and `moonshotai/Kimi-K2.5` routed to GLM-5.1 and Kimi-K2.6 - Scores: Reliability 33, Performance pending, Schema & documentation 65, Agent ergonomics 60, Security & auth 39, Payments & pricing 35, Task success pending, Maintenance & community 47, Transparency & trust 51 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, Model reading. · Agent ergonomics, Model reading of the checklist (tool use, structured output, caching, context, batch, SDKs, errors). · Security & auth, Model reading. · Payments & pricing, No machine payment protocol (0). · Maintenance & community, Model reading. · Transparency & trust, Closed service. - Sources: 25, open questions: 14, both in the full twin - Capabilities: inference.llm, inference.open-weights, embed.text, rerank, image.generate, video.generate, speech.tts, speech.stt - JSON: https://www.anchorterminal.com/api/v1/tools/siliconflow.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/siliconflow.svg` or a link to https://www.anchorterminal.com/tools/siliconflow from a page on siliconflow.com or one of its subdomains, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Set the OpenAI client's base URL to `https://api.siliconflow.com/v1`, or the Anthropic client's to `https://api.siliconflow.com/`, and send the key as a Bearer token 2. Take model ids from the model pages or `GET /v1/models`, not from the OpenAPI enum or the function calling guide, which list removed models 3. Check the `model` field of each response. On 11 June 2026 traffic for GLM-5 and Kimi-K2.5 was routed to successor models 4. Use `response_format` of `json_object` only where the model page says JSON Mode is supported, and keep `max_tokens` about 10,000 below the context length 5. Set the Claude Code environment variables by hand. The automated route pipes a script from an Amazon S3 bucket into bash ## Connect ```bash pip install --upgrade openai ``` ```bash curl --request POST \ --url https://api.siliconflow.com/v1/chat/completions \ --header 'authorization: Bearer $SILICONFLOW_API_KEY' \ --header 'content-type: application/json' \ --data '{"model":"deepseek-ai/DeepSeek-V4.1-Flash","messages":[{"role":"user","content":"Hello"}],"max_tokens":128}' ``` ```bash export ANTHROPIC_BASE_URL="https://api.siliconflow.com/" export ANTHROPIC_MODEL="your-preferred-model" export ANTHROPIC_API_KEY="sk-your-api-key-here" ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Cloudflare Workers AI | B | 68.3 | inference.llm, inference.open-weights, embed.text, rerank, image.generate, speech.stt, speech.tts | https://www.anchorterminal.com/tools/cloudflare-workers-ai.min.md | | LocalAI | B | 68 | inference.open-weights, embed.text, rerank, speech.stt, speech.tts, image.generate, video.generate | https://www.anchorterminal.com/tools/localai.min.md | | DeepInfra | B | 63 | inference.llm, inference.open-weights, embed.text, rerank, image.generate, speech.stt, speech.tts | https://www.anchorterminal.com/tools/deepinfra.min.md | | Novita AI | D | 53.5 | inference.llm, inference.open-weights, embed.text, rerank, image.generate, video.generate, speech.tts | https://www.anchorterminal.com/tools/novita-ai.min.md | | Lemonade | B | 63.8 | inference.open-weights, embed.text, rerank, speech.stt, speech.tts, image.generate | https://www.anchorterminal.com/tools/lemonade.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)