# Fix list: SiliconFlow From Anchor Terminal's listing at https://www.anchorterminal.com/tools/siliconflow, the October 2026 research run, assessed 9 October 2026. Grade D, 46.7 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on SiliconFlow: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Reliability, 33 out of 100, up to 13.4 more on the total Why it scored 33: Hosted reading. No status page is linked from the site, the docs index or any page read (0). With no page there is no incident record to read (5 of 30). Rate limits are published with numbers by spend tier, from 1,000 requests and 40,000 tokens a minute at L0 to 10,000 and 2,000,000 at L5, on a page served only in Chinese, with per-model figures left to the console (12 of 15). The 429 body is documented and names the limit reached, and the docs say to wait and retry. No `Retry-After` header, backoff figures or idempotency guidance were found (6 of 15). No SLA found. Clause 4.2 of the terms disclaims any uptime or availability commitment (0). The serverless API is generally available, and `/systemone` is marked alpha (10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 2. Security & auth, 39 out of 100, up to 10.7 more on the total Why it scored 39: Model reading. Bearer keys created in the console. No scopes, expiry, per-key limits or rotation are documented, and deletion of a key was not confirmed because the console was not read (15 of 30). The terms say Interaction Data is used only to supply the service and is not used or disclosed without authorisation. No statement on model training was found (12 of 20). The terms say there is no obligation to store Interaction Data. No retention period or zero-retention option is stated, and clause 5.11 allows communications to be recorded and passed to the authorities (6 of 15). The release note of 26 March 2026 describes usage and cost breakdowns and invoice export, and the pricing FAQ says monthly spending limits can be set. No audit log was found (6 of 15). No security.txt (404), disclosure policy, bug bounty or named certification was found on the pages read. The product introduction says the service complies with industry standards and names none (0 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 3. Payments & pricing, 35 out of 100, up to 8.1 more on the total Why it scored 35: No machine payment protocol (0). Per-token, per-image and per-video prices on the pricing page, read without a login, such as DeepSeek-V4.1-Flash at $0.15 in and $0.60 out per 1M tokens (20). $1 in free credits for a new account. Whether a card is needed was not established, so this line takes 15 of 20. A person signs up in a browser with a Google or GitHub account, and no key API was found (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 4. Agent ergonomics, 60 out of 100, up to 6.5 more on the total Why it scored 60: Model reading of the checklist (tool use, structured output, caching, context, batch, SDKs, errors). OpenAI-style `tools` with up to 128 functions, and each model page says whether Tools are supported. No `tool_choice` field is in the chat schema and the guide is out of date (12 of 20). `response_format` of `json_object`. The DeepSeek-V4.1-Flash page says Structured Outputs are not supported (7 of 15). Cached input prices are published per model. No caching guide was found (8 of 15). Context listed as 1049K tokens on eleven chat models on the home page, with output up to 393K (14 of 15). `/files` and `/batches` are in the OpenAPI file for chat completions only, with no guide in the English index, no batch price and a window description that contradicts itself (5 of 10). No official SDK. The OpenAI and Anthropic clients work with a changed base URL (6 of 10). Errors carry an HTTP status, a numeric code and a message. No list of codes was found (8 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 5. Schema & documentation, 65 out of 100, up to 5.7 more on the total Why it scored 65: Model reading. Public OpenAPI 3.0 file with 22 operations and 102 schemas, linked from `llms.txt` under the Chinese path only. Its chat model enum of 71 ids includes removed models and omits every model added since July 2026 (19 of 25). `llms.txt` and a Markdown copy of each docs page (10). Guides for function calling, JSON mode, reasoning, prefix and FIM completion. The function calling guide lists only models the release notes show as removed, and its examples use DeepSeek-V2.5 (10 of 20). Typed parameters with 104 enums and required fields. `response_format` is typed as a bare string with an example (10 of 15). Curl and Python examples, an error page with seven HTTP codes and one body code, and 400, 401, 404, 429, 503 and 504 responses in the file (9 of 15). The file's version reads 1.0.0. The release notes are dated, stop at 11 June 2026 and since March 2026 record only model removals (7 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 6. Maintenance & community, 47 out of 100, up to 4.6 more on the total Why it scored 47: Model reading. Hy4-preview is listed as released on 14 September 2026 and the blog's newest posts are dated 28 September 2026 (30). No notice period is stated. Six removal notices from 31 December 2025 to 11 June 2026 are dated the day of removal, and two in March 2026 gave six and seven days (2 of 12). The release notes record about 60 model ids removed in ten notices over the last twelve months (2 of 8). The release notes stop at 11 June 2026 while nine chat models were added afterwards. Support is by email and Discord (6 of 15). Replies in the support channels were not examined (2 of 10). No official SDK. The `langchain-siliconflow` package in the vendor's GitHub organisation was last pushed on 12 December 2025 (3 of 15). The API description and two guides are out of date (2 of 10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## 7. Transparency & trust, 51 out of 100, up to 4.3 more on the total Made of editorial 34, provenance 68. Why it scored 51: Closed service. The Terms of Use name the entity, Singapore law and SIAC arbitration. They also grant a licence "for in Singapore" and forbid use for any commercial purposes, which sits oddly with a paid API (10 of 30). The terms and the privacy policy agree that Interaction Data is processed on the customer's instructions with no obligation to store it. No retention period or DPA was found, and the privacy policy keeps drafting brackets around several values (12 of 30). Dated removal notices in the release notes, with no written policy or notice period (8 of 20). Stripe is named as payment processor and other providers only by type. Data may be handled within or outside Singapore, with no locations given (4 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Domain age: siliconflow.com, registered 2019-02-14 (7 years) (11 of 15) - Terms of service: read, states 6 of the 7 things a reader expects, and has 3 clauses that cost points (3.1 of 10) - Privacy policy: read, states 7 of the 8 things a reader expects (9.3 of 10) - Status page: not found (0 of 10) - security.txt: not found (0 of 10) ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - unchecked: the console at cloud.siliconflow.com, so key deletion, per-key settings, spending limits, per-model rate limits, usage logs and whether a card is needed for the free credit were not confirmed - unchecked: no request was sent to api.siliconflow.com, so the live model list, the 429 body and headers, and the routing of retired model ids rest on the docs - unchecked: whether a status page exists. None is linked from the site or the docs pages read, and no address was guessed - unchecked: the rendered API reference pages. The OpenAPI file was read in their place - unchecked: the Chinese-language docs beyond the rate limit page, the blog posts themselves, and replies in the Discord server - unchecked: model ids for GLM-5.3 and Kimi-K3 follow the vendor's naming pattern and were not read from their model pages - The release notes do not say when each entry was published. Six entries carry a label equal to the removal date. Whether customers were told earlier by email was not established, so no deduction is taken - The Terms of Use forbid use for any commercial purposes, scraping tools and benchmarking (clauses 3.4 p, i and l). Recorded as a fact with no deduction. It matters before any probe is run - The lead was right about the interface and the host. The fine-tuning product is console-only and was removed from batch 4. It is not graded here - The pricing page's audio table is headed per 1M UTF-8 bytes while its footnote says per minute and per 1,000 characters - `lastRelease` is the release date the site gives for Hy4-preview, the newest model, since the release notes stop at 11 June 2026 - One slip. The RDAP lookup at rdap.org redirected to rdap.verisign.com before that host's robots.txt was read. rdap.org's own robots.txt request answered 400, and api.github.com's answered 404 - No SLA, DPA, sub-processor list, security page or official SDK was found. Whether they exist on request is not established - Reserved GPUs and fine-tuning share the account and terms and were not graded. The products page marks dedicated endpoints as coming soon ## Weaknesses - No status page, incident history or SLA was found, and the terms disclaim any uptime or availability commitment - Release notes stop at 11 June 2026. Six notices from December 2025 to June 2026 give a removal date equal to the notice date - The Terms of Use of 9 September 2026 forbid use for any commercial purposes, scraping tools and benchmarking. This matters before any probe is run - The OpenAPI model list and the function calling guide name models already removed, and omit every model added since July 2026 - No security.txt, disclosure policy, certification, DPA or sub-processor list was found, and no official SDK is published ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Set the OpenAI client's base URL to `https://api.siliconflow.com/v1`, or the Anthropic client's to `https://api.siliconflow.com/`, and send the key as a Bearer token - Take model ids from the model pages or `GET /v1/models`, not from the OpenAPI enum or the function calling guide, which list removed models - Check the `model` field of each response. On 11 June 2026 traffic for GLM-5 and Kimi-K2.5 was routed to successor models - Use `response_format` of `json_object` only where the model page says JSON Mode is supported, and keep `max_tokens` about 10,000 below the context length - Set the Claude Code environment variables by hand. The automated route pipes a script from an Amazon S3 bucket into bash ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.