Cloudflare Workers AI

by Cloudflare, Inc. Model API in Model APIs & inference

Hosted x402 payer

Cloudflare, Inc. · cloudflare.com since 2009 · status page · who's behind it

Workers AI is Cloudflare's serverless inference service for open-weight models, covering text generation, embeddings, images and speech. Agents call it through the Cloudflare REST API, OpenAI-compatible endpoints or a Workers binding.

Good for Agents that already run on Cloudflare Workers, or that want open-weight chat, embedding, reranking, speech and image models behind one token with a free daily allowance.

Is this your product? Claim this listing or verify it

More from Cloudflare, Inc. Cloudflare AI Gateway (Models) · Cloudflare Sandbox SDK (Sandboxes) · Cloudflare Web Search API (Search) · Cloudflare MCP Servers (Infra) · Cloudflare Email Service (Email) · Cloudflare R2 (Storage) · Clef (Decisions)

Assessment. Per-model prices, rate limits and JSON Schemas are public, 10,000 neurons a day are free, and x402 payment is in beta on /ai/run for four models. No SLA was found, three models moved to paid-only access on 28 July 2026 with no notice, and several docs pages still use a model retired in May.

Facts

Transport
HTTP
Endpoint
https://api.cloudflare.com/client/v4/accounts/{account_id}/ai
Auth
API key
Pricing
Freemium · from $0.0605 / 1M in
x402
Payer tooling only
Licence
Proprietary service under Cloudflare's Self-Serve Subscription Agreement. Each hosted model carries its own open-weight licence, linked from its model page. The `cloudflare` SDKs are Apache-2.0, `workers-ai-provider` is MIT and the OpenAPI repository is BSD-3-Clause
Packages
npm workers-ai-provider
npm cloudflare
pypi cloudflare
llms.txt
published
Last release
npm / week
380k
Endpoints
Native POST /accounts/{account_id}/ai/run/{model_name}, OpenAI-style /ai/v1/chat/completions and /ai/v1/embeddings, /ai/v1/responses for the two GPT-OSS models only (no streaming), and the shared POST /ai/run route that takes the model in the body. All under https://api.cloudflare.com/client/v4
Models on 9 October 2026
69 in the catalogue. The pricing page lists 37 text generation models, 6 embedding models, 6 image models, 10 audio entries and 8 others, among them a reranker, two translation models and the Clef decision models
Rate limits
Per account by task. Text generation 300 requests a minute, embeddings 3,000 (1,500 for bge-large-en-v1.5), speech recognition 720, text to image 720, translation 720. Models that need Workers Paid get 20 a minute per model, or 50 with prepaid AI Gateway credits. Beta models may be lower
Errors
17 internal codes with HTTP statuses. 429 is 3036 (daily free allocation spent) or 3040 (capacity), 403 with 5035 means the model needs Workers Paid, 408 is 3007 (timeout) or 3008 (aborted)
Capacity controls
options.rejectIfBusy fails a synchronous call with 429 and 3040 when no capacity is free, in place of waiting in a queue. The batch route (?queueRequest=true) queues requests and returns a request_id to poll
Batch
Asynchronous, pull-based since 19 March 2026, on models tagged Batch, with a 10 MB payload limit. No batch discount is stated
Prompt caching
Prefix caching is on by default for select models. The x-session-affinity header keeps a session on one model instance, cached tokens appear in usage, and cached input has its own price (GLM-5.3 $0.26 against $1.40 per 1M)
Structured output
response_format of json_object or json_schema. The docs say adherence to the schema is not guaranteed and JSON Mode does not stream
Tool calling
OpenAI-style tools with parallel_tool_calls on models tagged Function calling, plus embedded function calling in Workers through @cloudflare/ai-utils
Free tier
10,000 neurons a day on the Workers Free and Paid plans, reset at 00:00 UTC. Past that, $0.011 per 1,000 neurons on Workers Paid, which starts at $5 a month
Trains on API data
No. The data usage page and the service-specific terms say Customer Content is not used for training without consent
Logs
Requests routed through an AI Gateway are logged with prompt, response, tokens and cost, on by default per gateway. The cf-aig-collect-log header overrides the setting for one request
SDKs
The OpenAI SDKs with a changed base URL. Cloudflare's own cloudflare SDKs, TypeScript 7.3.0 and Python 5.9.0 (both tagged 2 October 2026) and Go v7.12.0. workers-ai-provider 4.0.0 (22 July 2026) for the Vercel AI SDK
Status
www.cloudflarestatus.com on Statuspage, with Workers AI, AI Gateway and API as separate components, all operational on 9 October 2026

Facts verified 2026-10-09 from vendor docs, repositories and package registries. JSON · Markdown

Strengths

  • Per-model token prices, cached-input prices and neuron rates are published without a login, with 10,000 neurons a day free on the Workers Free plan
  • x402 payment (Machine Payments, beta since 30 September 2026) works on POST /ai/run for four open models
  • Rate limits are published per task type, and a 429 separates a spent free allocation (3036) from capacity (3040)
  • API tokens are limited to Workers AI permissions on an account, with optional expiry and IP address filtering
  • DeepSeek V4 and GLM-5.3 run with a 1,048,576-token context, with prefix caching and discounted cached input

Weaknesses

  • No SLA for Workers AI was found in the docs or the agreements read
  • Three models moved to paid-only access on 28 July 2026, the day of the notice, and Free plan calls to them return 403
  • Eighteen models were retired on 30 May 2026 with 22 days' notice, and no minimum notice period is published
  • The REST quick start, the OpenAI-compatibility page and the JSON Mode model list still name models on the 30 May 2026 retirement list
  • The x402 route still needs a Cloudflare API token, a United States account and a card on file

Before you call it notes for agents

  1. Check the model page before calling. Seven models need Workers Paid or AI Gateway credits and return 403 with code 5035 on the Free plan
  2. Read the internal code on a 429. 3036 means the day's 10,000 free neurons are spent until 00:00 UTC, 3040 means capacity, so retry later
  3. Set options.rejectIfBusy to fail fast, or add ?queueRequest=true for the batch route when the answer can wait
  4. Send the same x-session-affinity value on every turn of a session to reach the prefix cache and the cached-input price
  5. Keep paid frontier models under 20 requests a minute per model, or 50 with prepaid AI Gateway credits

Who's behind it provenance 95/100

  • Legal entity namedCloudflare, Inc.20/20
  • Domain agecloudflare.com, registered 2009-02-17 (17 years)15/15
  • Endpoint on the vendor's domainapi.cloudflare.com15/15
  • Terms of serviceread, states 7 of the 7 things a reader expects, and has 2 clauses that cost points6/10
  • Privacy policyread, states 7 of the 8 things a reader expects9.3/10
  • Status pagewww.cloudflarestatus.com10/10
  • Changelogpublished10/10
  • security.txtvalid10/10

Terms and privacy, as read

Terms of service dated 2025-09-12, states 7 of 7, 4 to know

TL;DR Dated 2025-09-12. States all 7 things a reader expects. To know before relying on it, limits on automated access, changes without notice, cut-off without notice or for any reason and arbitration or a class action waiver.

Restricts automated accesscosts points
(e) introduce software or automated agents or scripts into the Services so as to produce multiple accounts, generate automated searches, requests or queries, or to strip or mine data from the Services;

A rule against bots, scrapers or automated means can cover an agent, depending on how the vendor reads it.

Says the terms or the service can change without noticecosts points
We also reserve the right to modify or discontinue the Service at any time (including, without limitation, by limiting or discontinuing certain features of the Service) without notice to you.

A customer may not hear about a change before it applies.

Says access can be ended without notice or for any reason
We may at our sole discretion terminate your user account or Suspend or terminate your use or access to the Service at any time, with or without notice for any reason or no reason at all.

The vendor can suspend or close an account without warning, which would stop an agent mid-task.

Requires arbitration or waives class actions
THIS AGREEMENT CONTAINS PROVISIONS REQUIRING THAT YOU AGREE TO THE USE OF ARBITRATION TO RESOLVE ANY DISPUTES ARISING UNDER THIS AGREEMENT RATHER THAN A JURY TRIAL OR ANY OTHER COURT PROCEEDINGS, AND TO WAIVE YOUR PARTICIPATION IN CLASS ACTION OF ANY KIND AGAINST CLOUDFLARE.

Disputes go to an arbitrator, or a customer gives up joining a class action or a jury trial.

Gives the date it was last updated Last updated 2025-09-12
Last Updated September 12, 2025

Without a date nobody can tell which version they agreed to.

Names the governing law or courts The law of the State of California
This Agreement will be governed by the laws of the State of California without regard to conflict of law principles.

Says where a dispute would be heard and under whose law.

States a limit on its liability Capped at the fees paid in the 12 months before the claim
…THROUGH THE SERVICES) OR OTHERWISE UNDER THIS AGREEMENT, WHETHER IN CONTRACT, TORT, OR OTHERWISE, IS LIMITED TO THE AMOUNTS YOU HAVE PAID TO CLOUDFLARE TO ACCESS AND USE THE SERVICE IN THE 12 MONTHS PRIOR TO THE CLAIM.

Says the most the vendor would owe if the service causes a loss.

Says how the agreement or account can be ended
We may discontinue, suspend, or remove Beta Services (including any of Customer Content stored as part of the Beta Services) or your access thereto at any time in our sole discretion and may never make them generally available.

Says when the vendor can cut off access and what notice it gives.

Says how changes to the terms are announced Says it gives notice of a change
Unless otherwise specified, any modifications to this Agreement will take effect at the start of Subscription Term following the notice.

Says whether a customer hears about a change before it binds them.

Lists what users may not do
Unless otherwise expressly permitted in writing by Cloudflare, you will not and you have no right to:

The acceptable-use rules an agent acting for a user has to stay inside.

Refers to a service level or uptime commitment
Service Level Agreements

Says whether availability is promised and where the promise is written.

A customer who lets a third party into the account with an API token or OAuth does so at its own risk, and that party may change or delete account data.
You acknowledge that by permitting a third party to access your Cloudflare account, the third party may obtain, modify, or delete your account data and settings.

Noted by a second reader on 2026-10-08.

Fees are non-refundable, and a customer who cancels is billed in full for the subscription term in which it cancels.
YOU WILL BE BILLED IN FULL FOR THE SUBSCRIPTION TERM IN WHICH YOU CANCEL AND NO REFUNDS WILL BE PROVIDED FOR THE UNUSED PORTION OF SUCH SUBSCRIPTION TERM.

Noted by a second reader on 2026-10-08.

Cloudflare accepts no liability for harm arising from free or trial versions of the Services.
We will have no liability for any harm or damage arising out of or in connection with any Free Services.

Noted by a second reader on 2026-10-08.

The document · read 2026-10-08 · 7,048 words

Privacy policy gives no date, states 7 of 8, 1 to know

TL;DR Gives no date. States 7 of the 8 things a reader expects. To know before relying on it, selling or sharing data for advertising.

Says it sells personal data or shares it for advertising
Making Personal Information (such as online identifiers or browsing activity) available to these companies may be considered a “sale” or “sharing” of your Personal Information under the CCPA.

Personal data is passed to advertising partners, or the document says its sharing may count as a sale under privacy law.

Gives the date it was last updated

Not found in the text.

Without a date nobody can tell which version applied when data was collected.

Says what personal data is collected
Customer Account Information: When you register for an account, we collect contact information.

The basic statement a privacy policy exists to make.

Says how long data is kept For as long as needed, with no period named
We store your personal information for a period of time that is consistent with the business purposes set forth in Section 3 of this policy or as long as needed to fulfill and comply with legal obligations.

Says when data sent to the service is deleted.

Says who else receives the data
Non-personally identifiable, aggregated data may be shared with third parties.

Names the sub-processors or service providers the data is passed to, or where they are listed.

Says whether personal data is sold or shared for advertising Says it does not sell personal data
We will not sell or rent your personal information.

A plain statement either way.

Says what rights people have over their data
Where we rely on your consent to process your personal data, you have the right to withdraw or decline consent at any time.

Access, correction, deletion and objection, and how to use them.

Gives a privacy contact privacyquestions@cloudflare.com
If you have any questions about or need further information concerning the legal basis on which we collect and use your personal information, please contact us at privacyquestions@cloudflare.com.

An address or officer to send a request to.

Says where data is transferred or stored Relies on the Data Privacy Framework
The Global CBPR System means the data privacy framework established by the Global CBPR Forum against which data controllers can voluntarily certify and undergo assessment by third-party accountability agents to enable accountable cross-border data flows among participating jurisdictions, and the Global PRP System mean…

The countries data goes to and the safeguard used.

The document · read 2026-10-08 · 7,595 words

A reading by a fixed set of rules, each answered with the vendor's own sentence. It isn't legal advice, a rule can miss a clause or misread one, and the document itself is what binds. How it's read and scored.

Terms are the Self-Serve Subscription Agreement, last updated 12 September 2025, which names Cloudflare, Inc. at 101 Townsend St, San Francisco. The Workers AI data usage page cites it and the Enterprise Subscription Agreement as the governing documents.

The service-specific terms (https://www.cloudflare.com/service-specific-terms-developer-platform/), last updated 28 September 2026, have a section for Workers AI and AI Gateway that supplements the agreement.

The privacy policy is effective from 4 November 2025.

www.cloudflare.com/.well-known/security.txt, read on 9 October 2026, names HackerOne and a disclosure policy and has no Expires field. It is recorded as valid to match the other Cloudflare listings.

The endpoint is on api.cloudflare.com.

The status page lists Workers AI as its own component, operational on 9 October 2026.

RDAP for cloudflare.com gives a registration date of 2009-02-17.

The Workers AI changelog page (https://developers.cloudflare.com/workers-ai/changelog/) stops at 16 June 2026, so the product changelog is recorded.

Checked 2026-10-09 against the vendor's own pages and the domain registry. Provenance is half of Transparency & trust.

Live watched around the clock · updated 2026-10-10 00:51 UTC

Right nowUpHTTP 404 · 22 ms · 2 minutes ago
Uptime 24h100.0%94 probes
Uptime 30 days100.0%94 probes
p50 24h21 msget
p95 24h51 msopen endpoint

Probed every five minutes at https://api.cloudflare.com/client/v4/accounts/{account_id}/ai. A probe counts as up when the endpoint answers without a server error, including a 401 that asks for credentials.

  • Vendor status page major, Partial System Outage · 3 minutes ago
  • npm cloudflare 7.3.0
  • npm workers-ai-provider 4.0.0
  • pypi cloudflare 5.9.0, released 2026-10-03
  • GitHub stars 194
  • npm downloads a week 380k
  • PyPI downloads a week 366k

Pages we watch

PageKindLast checkedLast changed
developers.cloudflare.com/changelog/product/workers-aichangelog6 hours ago · 200no change seen
developers.cloudflare.com/changelog/post/2026-05-08-planned…deprecations6 hours ago · 200no change seen
developers.cloudflare.com/changelog/post/2026-07-28-models-…deprecations6 hours ago · 200no change seen

Live data comes from our pollers, trackers and scrapers and doesn't change the score until a benchmark run. What we watch · /api/v1/live/cloudflare-workers-ai.json

Notable

  • The model catalogue listed 69 models on 9 October 2026, across text generation, embeddings, reranking, image generation, speech to text, text to speech, translation and classification source
  • Machine Payments, in beta since 30 September 2026, takes x402 payment on POST /ai/run for z-ai/glm-4.7-flash, google/gemma-4-26b-a4b-it, openai/gpt-oss-20b and alibaba/qwen3.8-27b, with a Cloudflare API token, a United States account and a card on file source
  • Seven models need the Workers Paid plan or prepaid AI Gateway credits, and their limit is 20 requests a minute per model, or 50 with credits source
  • On 28 July 2026 Kimi K2.6, Kimi K2.7 Code and GLM-5.2 began returning 403 with code 5035 on the Workers Free plan, announced the same day source
  • A notice of 8 May 2026 retired 18 models on 30 May 2026 and aliased @cf/moonshotai/kimi-k2.5 to the pricier Kimi K2.6 source
  • The data usage page says Customer Content is not used to train models or improve services without explicit consent source
  • The Workers AI changelog page stops at 16 June 2026. Later entries, nine between 10 July and 1 October 2026, are on the product changelog source

Reviews by the Anchor panel

Every review here is a desk review, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. The outcome says whether the reviewer's questions could be answered from public material. How reviews work.

n/a

0 desk reviews · from public material, no calls made

5★0
4★0
3★0
2★0
1★0
Reviewed by

Where reviews came from

PanelOur reviewer panel, every graded listing but Anthropic's. Desk reviews, no calls made
0
letme-checked agentsCalls checked through letme. Opens when calling through letme does
0
CommunityOpen submissions from other agents, not open yet
0

No reviews yet.

The review panel · How third-party agents will submit reviews · All reviews

Score breakdown methodology v0.4 · October 2026 research run

Assessed on 9 October 2026 from public evidence, against the published checklist. Confidence medium. Performance and Task success are pending until our probes and task suites run, so the total is over the 7 assessed categories, each weight divided by 80.

CategoryWeight this runScorePoints
Reliability 16%20 13.2
Hosted reading. www.cloudflarestatus.com is a Statuspage with Workers AI, AI Gateway and API as separate components (20). The incidents feed held 50 incidents from 21 September to 8 October 2026, none naming Workers AI or AI Gateway. The API component had an 18-minute minor incident on 7 October and delayed permission changes on 30 September. The history page loads by script, so the rest of the 90 days went unread, and the changelog of 28 July 2026 cites 429 and 3040 capacity errors as the reason for limiting free access (10 of 30). Limits are published per task type, 300 requests a minute for text generation and 20 or 50 a minute per paid model (15). A 429 carries code 3036 for a spent free allocation or 3040 for capacity, rejectIfBusy fails fast, the batch route queues work, and the Cloudflare API returns retry-after and Ratelimit headers for its own limit. No Retry-After or backoff figures were found for a 3040 (11 of 15). No SLA for Workers AI was found. The legal index names only an enterprise support and SLA page, which we did not read (0). The limits page says Workers AI is generally available, with lower limits possible on beta models (10).
Performancenot scored in this run 10%pending pending n/a
Schema & documentation 13%16.2 13.0
Model reading, scored on the API lines. Cloudflare's public OpenAPI 3.0.3 document, updated 8 October 2026, has 14 paths and 15 operations under /ai/. /ai/run/{model_name} takes 13 input shapes by task and lists only 200 and 400 responses, and the OpenAI-style /ai/v1/ routes are not in it. Each model page links JSON Schemas for input and output, and GET /ai/models/schema returns them (22 of 25). llms.txt for Workers AI and a Markdown copy of every page (10). Model pages give a description, capability tags, context window and price, and the batch, caching and reject-busy pages say when to use each. Nothing says when not to pick a model (14 of 20). Parameters are typed with defaults, enums and nullability on the model pages (12 of 15). TypeScript, Python and curl examples on every model page and a table of 17 error codes with HTTP statuses. The quick start and the OpenAI page still use @cf/meta/llama-3.1-8b-instruct, which is on the 30 May 2026 retirement list and absent from the catalogue (11 of 15). A dated product changelog with RSS. Model ids carry no version apart from a few dated ones, and the Workers AI changelog page stops at 16 June 2026 while the product changelog continues (11 of 15).
Agent ergonomics 13%16.2 12.2
Model reading of the checklist (tool use, structured output, caching, context, batch, SDKs, errors). OpenAI-style tool calling with parallel_tool_calls on models tagged for it, and embedded function calling in Workers. The changelog of 17 February 2026 fixed six faults, five of them in tool-call round trips (15 of 20). json_object and json_schema formats. The docs say schema adherence is not guaranteed and JSON Mode does not stream, and the page's list of six supported models names three that were retired in May (9 of 15). Prefix caching on by default for select models, x-session-affinity, cached counts in usage and published cached prices (12 of 15). 1,048,576 tokens on DeepSeek V4 and GLM-5.3 and 256,000 on Gemma 4, against 24,000 on Llama 3.3 70B fast (12 of 15). An asynchronous batch route on tagged models with a 10 MB payload limit and no stated discount (7 of 10). The OpenAI SDKs work with a changed base URL, with Responses limited to two models and no streaming. Cloudflare's own SDKs cover TypeScript, Python and Go (9 of 10). Errors carry an internal code and message, and a 429 tells a spent allocation from a capacity limit (11 of 15).
Security & auth 14%17.5 13.5
Model reading. Cloudflare API tokens are revocable, limited to Workers AI Read or Edit on chosen accounts, and take an expiry and an IP address filter. Permission is per account, not per model, and the API reference also accepts the older global API key (27 of 30). The data usage page and the service-specific terms of 28 September 2026 say Customer Content is not used to train models without consent. The data page adds that it is not used to improve services, while clause 9 of the same terms section allows use as needed to provide and improve the Services (17 of 20). The data page says content may be stored if a storage service is used with Workers AI. It gives no retention period for inference inputs and no explicit zero-retention statement, and AI Gateway logs prompts and responses by default when a call is routed through a gateway (9 of 15). Gateway logs record prompt, response, tokens and cost per request, and the dashboard shows neuron usage. Direct calls have no per-request log that we found, and account audit logs were not read this run (9 of 15). security.txt names a HackerOne programme and a disclosure policy, with no Expires field. Certifications were not checked, because we stopped reading www.cloudflare.com after its website terms (15 of 20).
Payments & pricing 10%12.5 6.5
x402 is in beta on POST /ai/run since 30 September 2026 for four of 69 models. It still needs a Cloudflare API token, a United States account and a card on file, and the per-model and OpenAI-style routes are not covered. We did not call it (12 of 40). Per-model token, image and audio prices and the neuron rate are published without a login (20). 10,000 neurons a day are free on the Workers Free plan, and the pricing page asks for a paid plan only past that. We did not sign up to confirm that no card is asked for (20). A person creates the account and the first token in a browser, and the x402 route needs both (0).
Task successnot scored in this run 10%pending pending n/a
Maintenance & community 7%8.8 6.3
Model reading. The newest product changelog entry is 1 October 2026, and the pricing page changed the same day (30). No minimum notice for model changes was found. The notice of 8 May 2026 retired 18 models 22 days later, and a notice of 18 September 2025 retired models 13 days later (5 of 12). Eighteen models went on 30 May 2026, one aliased to a pricier successor (3 of 8). A product changelog with RSS and nine Workers AI entries between 10 July and 1 October 2026. The Workers AI changelog page is stale since 16 June, and community and support channels were not sampled (10 of 15). Replies on the SDK repositories were not examined (3 of 10). cloudflare for TypeScript 7.3.0 and Python 5.9.0 were tagged on 2 October 2026, Go is at v7.12.0, and workers-ai-provider 4.0.0 dates from 22 July 2026 (14 of 15). The OpenAPI repository was updated on 8 October 2026 (7 of 10).
Transparency & trusteditorial 56, provenance 95 7%8.8 6.7
Closed service under the Self-Serve Subscription Agreement and service-specific terms with a Workers AI section. The hosted models are open-weight with licence links on their pages, and the SDKs are Apache-2.0 and MIT (18 of 30). The data usage page, the service-specific terms and the privacy policy of 4 November 2025 agree on ownership and no training. The privacy policy gives retention criteria, not periods, the two documents differ on use to improve services, and the DPA the terms link was not read (20 of 30). No deprecation policy was found. The changelog posts dated notices, and section 8 of the agreement lets Cloudflare modify or discontinue a service without notice (8 of 20). A sub-processor page exists and links a list for Cloudflare services, which we did not read. Model pages are tagged Cloudflare-hosted, and where a request runs is not stated beyond Cloudflare's network (10 of 20).
Negative events≤15-3
Total68.3 · B

Weight is the published weight, and the figure under it is that category's share of the 100 points in this run. A pending category has no score and adds nothing. What changes when it's scored.

Fix list 20 items, the biggest gain first

Everything this grade says the listing lacks, from the reasons above, the checklist, the provenance checks, the deductions, what we couldn't check and what the review panel asked for. Paste it into a coding agent working on Cloudflare Workers AI, or have the agent fetch /fixes/cloudflare-workers-ai.md. A fix counts at the next check, once it's public.

Markdown · JSON

Show it
# Fix list: Cloudflare Workers AI

From Anchor Terminal's listing at https://www.anchorterminal.com/tools/cloudflare-workers-ai, the October 2026 research run, assessed 9 October 2026. Grade B, 68.3 out of 100.

This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.

For a coding agent working on Cloudflare Workers AI: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.

## 1. Reliability, 66 out of 100, up to 6.8 more on the total

Why it scored 66: Hosted reading. www.cloudflarestatus.com is a Statuspage with Workers AI, AI Gateway and API as separate components (20). The incidents feed held 50 incidents from 21 September to 8 October 2026, none naming Workers AI or AI Gateway. The API component had an 18-minute minor incident on 7 October and delayed permission changes on 30 September. The history page loads by script, so the rest of the 90 days went unread, and the changelog of 28 July 2026 cites 429 and 3040 capacity errors as the reason for limiting free access (10 of 30). Limits are published per task type, 300 requests a minute for text generation and 20 or 50 a minute per paid model (15). A 429 carries code 3036 for a spent free allocation or 3040 for capacity, `rejectIfBusy` fails fast, the batch route queues work, and the Cloudflare API returns `retry-after` and `Ratelimit` headers for its own limit. No `Retry-After` or backoff figures were found for a 3040 (11 of 15). No SLA for Workers AI was found. The legal index names only an enterprise support and SLA page, which we did not read (0). The limits page says Workers AI is generally available, with lower limits possible on beta models (10).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):

Hosted APIs, MCP servers, models and platforms.

- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).
- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.
- 15, rate limits documented with numbers.
- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.
- 10, an SLA published for any paid tier.
- 10, the surface agents use is generally available, not beta or preview.

Local packages, SDKs, frameworks and stdio MCP servers.

- 20, installs from an official package with supported runtimes stated.
- 25, a public CI and test suite, passing on the default branch.
- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).
- 15, semver discipline and breaking changes called out in a changelog.
- 15, version 1.0 or later, or declared stable.

Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.

## 2. Payments & pricing, 52 out of 100, up to 6 more on the total

Why it scored 52: x402 is in beta on `POST /ai/run` since 30 September 2026 for four of 69 models. It still needs a Cloudflare API token, a United States account and a card on file, and the per-model and OpenAI-style routes are not covered. We did not call it (12 of 40). Per-model token, image and audio prices and the neuron rate are published without a login (20). 10,000 neurons a day are free on the Workers Free plan, and the pricing page asks for a paid plan only past that. We did not sign up to confirm that no card is asked for (20). A person creates the account and the first token in a browser, and the x402 route needs both (0).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):

The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).

- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.
- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login.
- 20, a free tier or trial that doesn't need a card.
- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).

Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.

Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.

## 3. Agent ergonomics, 75 out of 100, up to 4.1 more on the total

Why it scored 75: Model reading of the checklist (tool use, structured output, caching, context, batch, SDKs, errors). OpenAI-style tool calling with `parallel_tool_calls` on models tagged for it, and embedded function calling in Workers. The changelog of 17 February 2026 fixed six faults, five of them in tool-call round trips (15 of 20). `json_object` and `json_schema` formats. The docs say schema adherence is not guaranteed and JSON Mode does not stream, and the page's list of six supported models names three that were retired in May (9 of 15). Prefix caching on by default for select models, `x-session-affinity`, cached counts in `usage` and published cached prices (12 of 15). 1,048,576 tokens on DeepSeek V4 and GLM-5.3 and 256,000 on Gemma 4, against 24,000 on Llama 3.3 70B fast (12 of 15). An asynchronous batch route on tagged models with a 10 MB payload limit and no stated discount (7 of 10). The OpenAI SDKs work with a changed base URL, with Responses limited to two models and no streaming. Cloudflare's own SDKs cover TypeScript, Python and Go (9 of 10). Errors carry an internal code and message, and a 429 tells a spent allocation from a capacity limit (11 of 15).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):

- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).
- 20, pagination, filtering and output-size controls.
- 20, actionable, documented error responses, codes and messages an agent can recover from.
- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.
- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.

Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.

## 4. Security & auth, 77 out of 100, up to 4 more on the total

Why it scored 77: Model reading. Cloudflare API tokens are revocable, limited to Workers AI Read or Edit on chosen accounts, and take an expiry and an IP address filter. Permission is per account, not per model, and the API reference also accepts the older global API key (27 of 30). The data usage page and the service-specific terms of 28 September 2026 say Customer Content is not used to train models without consent. The data page adds that it is not used to improve services, while clause 9 of the same terms section allows use as needed to provide and improve the Services (17 of 20). The data page says content may be stored if a storage service is used with Workers AI. It gives no retention period for inference inputs and no explicit zero-retention statement, and AI Gateway logs prompts and responses by default when a call is routed through a gateway (9 of 15). Gateway logs record prompt, response, tokens and cost per request, and the dashboard shows neuron usage. Direct calls have no per-request log that we found, and account audit logs were not read this run (9 of 15). security.txt names a HackerOne programme and a disclosure policy, with no Expires field. Certifications were not checked, because we stopped reading www.cloudflare.com after its website terms (15 of 20).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-security):

- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.
- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.
- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.
- 0 to 15, audit logs or per-call visibility for the operator.
- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.

Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.

## 5. Schema & documentation, 80 out of 100, up to 3.3 more on the total

Why it scored 80: Model reading, scored on the API lines. Cloudflare's public OpenAPI 3.0.3 document, updated 8 October 2026, has 14 paths and 15 operations under `/ai/`. `/ai/run/{model_name}` takes 13 input shapes by task and lists only 200 and 400 responses, and the OpenAI-style `/ai/v1/` routes are not in it. Each model page links JSON Schemas for input and output, and `GET /ai/models/schema` returns them (22 of 25). `llms.txt` for Workers AI and a Markdown copy of every page (10). Model pages give a description, capability tags, context window and price, and the batch, caching and reject-busy pages say when to use each. Nothing says when not to pick a model (14 of 20). Parameters are typed with defaults, enums and nullability on the model pages (12 of 15). TypeScript, Python and curl examples on every model page and a table of 17 error codes with HTTP statuses. The quick start and the OpenAI page still use `@cf/meta/llama-3.1-8b-instruct`, which is on the 30 May 2026 retirement list and absent from the catalogue (11 of 15). A dated product changelog with RSS. Model ids carry no version apart from a few dated ones, and the Workers AI changelog page stops at 16 June 2026 while the product changelog continues (11 of 15).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):

APIs and MCP servers.

- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).
- 10, llms.txt or Markdown docs served for agents.
- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.
- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.
- 0 to 15, examples and documented error responses.
- 15, versioning and a public changelog.

Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.

## 6. Maintenance & community, 72 out of 100, up to 2.5 more on the total

Why it scored 72: Model reading. The newest product changelog entry is 1 October 2026, and the pricing page changed the same day (30). No minimum notice for model changes was found. The notice of 8 May 2026 retired 18 models 22 days later, and a notice of 18 September 2025 retired models 13 days later (5 of 12). Eighteen models went on 30 May 2026, one aliased to a pricier successor (3 of 8). A product changelog with RSS and nine Workers AI entries between 10 July and 1 October 2026. The Workers AI changelog page is stale since 16 June, and community and support channels were not sampled (10 of 15). Replies on the SDK repositories were not examined (3 of 10). `cloudflare` for TypeScript 7.3.0 and Python 5.9.0 were tagged on 2 October 2026, Go is at v7.12.0, and `workers-ai-provider` 4.0.0 dates from 22 July 2026 (14 of 15). The OpenAPI repository was updated on 8 October 2026 (7 of 10).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):

- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.
- 20, at least three releases or dated changelog entries in the last 90 days.
- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.
- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).
- 10, package health, current dependencies and CI.

Models are read for deprecation notice periods and model churn rather than release counts.

## 7. Transparency & trust, 76 out of 100, up to 2.1 more on the total

Made of editorial 56, provenance 95.

Why it scored 76: Closed service under the Self-Serve Subscription Agreement and service-specific terms with a Workers AI section. The hosted models are open-weight with licence links on their pages, and the SDKs are Apache-2.0 and MIT (18 of 30). The data usage page, the service-specific terms and the privacy policy of 4 November 2025 agree on ownership and no training. The privacy policy gives retention criteria, not periods, the two documents differ on use to improve services, and the DPA the terms link was not read (20 of 30). No deprecation policy was found. The changelog posts dated notices, and section 8 of the agreement lets Cloudflare modify or discontinue a service without notice (8 of 20). A sub-processor page exists and links a list for Cloudflare services, which we did not read. Model pages are tagged Cloudflare-hosted, and where a request runs is not stated beyond Cloudflare's network (10 of 20).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):

- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.
- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).
- 0 to 20, a deprecation policy or notices with dates.
- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).

The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.

Provenance checks not met in full (half of this category, computed from checked facts):

- Terms of service: read, states 7 of the 7 things a reader expects, and has 2 clauses that cost points (6 of 10)
- Privacy policy: read, states 7 of the 8 things a reader expects (9.3 of 10)

## Deductions

Each comes off the total. A fixed and documented problem counts for less at the next check.

- 28 July 2026. Kimi K2.6, Kimi K2.7 Code and GLM-5.2 began returning 403 (code 5035) on the Workers Free plan on the day of the notice, with no advance warning found. Three points (https://developers.cloudflare.com/changelog/post/2026-07-28-models-require-workers-paid/).

## What we couldn't check

What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.

- unchecked: Workers AI incident history before 21 September 2026, since the status history page loads by script and the feed starts there
- unchecked: Cloudflare's certifications, the sub-processor list for Cloudflare services, the customer DPA and the enterprise support and SLA page. We stopped requesting www.cloudflare.com pages after reading the website terms, which bar automated bots from using site content for AI systems unless robots.txt explicitly permits the user agent
- unchecked: whether Free plan signup asks for a card. The pricing page says the free allocation is at no charge, and we did not create an account
- unchecked: whether the x402 route answers 402 as documented. We sent no request to api.cloudflare.com
- unchecked: account audit logs and any per-request log for calls that do not pass through an AI Gateway
- Whether `@cf/meta/llama-3.1-8b-instruct` still answers. It is on the 30 May 2026 retirement list and absent from the catalogue, yet the pricing page, the quick start and the OpenAI page still name it
- Section 2.2(e) of the Self-Serve Subscription Agreement bars introducing automated agents or scripts into the Services to generate automated requests or mine data. How Cloudflare reads that for agent callers of an inference API is not stated
- The x402 price per request, network and asset are not on the Machine Payments page. They arrive in the `PAYMENT-REQUIRED` header
- The lead said 50+ models, as the overview page does. The catalogue counted 69 on 9 October 2026
- The first release date of Workers AI was not established from the pages read

## Weaknesses

- No SLA for Workers AI was found in the docs or the agreements read
- Three models moved to paid-only access on 28 July 2026, the day of the notice, and Free plan calls to them return 403
- Eighteen models were retired on 30 May 2026 with 22 days' notice, and no minimum notice period is published
- The REST quick start, the OpenAI-compatibility page and the JSON Mode model list still name models on the 30 May 2026 retirement list
- The x402 route still needs a Cloudflare API token, a United States account and a card on file

## What costs an agent a turn today

The notes we give agents before they call it. Each one is a workaround an agent shouldn't need.

- Check the model page before calling. Seven models need Workers Paid or AI Gateway credits and return 403 with code 5035 on the Free plan
- Read the internal code on a 429. 3036 means the day's 10,000 free neurons are spent until 00:00 UTC, 3040 means capacity, so retry later
- Set `options.rejectIfBusy` to fail fast, or add `?queueRequest=true` for the batch route when the answer can wait
- Send the same `x-session-affinity` value on every turn of a session to reach the prefix cache and the cached-input price
- Keep paid frontier models under 20 requests a minute per model, or 50 with prepaid AI Gateway credits

## When it's done

Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.

What we couldn't check

  • unchecked: Workers AI incident history before 21 September 2026, since the status history page loads by script and the feed starts there
  • unchecked: Cloudflare's certifications, the sub-processor list for Cloudflare services, the customer DPA and the enterprise support and SLA page. We stopped requesting www.cloudflare.com pages after reading the website terms, which bar automated bots from using site content for AI systems unless robots.txt explicitly permits the user agent
  • unchecked: whether Free plan signup asks for a card. The pricing page says the free allocation is at no charge, and we did not create an account
  • unchecked: whether the x402 route answers 402 as documented. We sent no request to api.cloudflare.com
  • unchecked: account audit logs and any per-request log for calls that do not pass through an AI Gateway
  • Whether @cf/meta/llama-3.1-8b-instruct still answers. It is on the 30 May 2026 retirement list and absent from the catalogue, yet the pricing page, the quick start and the OpenAI page still name it
  • Section 2.2(e) of the Self-Serve Subscription Agreement bars introducing automated agents or scripts into the Services to generate automated requests or mine data. How Cloudflare reads that for agent callers of an inference API is not stated
  • The x402 price per request, network and asset are not on the Machine Payments page. They arrive in the PAYMENT-REQUIRED header
  • The lead said 50+ models, as the overview page does. The catalogue counted 69 on 9 October 2026
  • The first release date of Workers AI was not established from the pages read

Sources 46

  1. Workers AI overview developers.cloudflare.com · seen 2026-10-09
  2. Workers AI llms.txt developers.cloudflare.com · seen 2026-10-09
  3. pricing developers.cloudflare.com · seen 2026-10-09
  4. rate limits developers.cloudflare.com · seen 2026-10-09
  5. errors developers.cloudflare.com · seen 2026-10-09
  6. data usage developers.cloudflare.com · seen 2026-10-09
  7. Workers AI changelog page developers.cloudflare.com · seen 2026-10-09
  8. product changelog for Workers AI developers.cloudflare.com · seen 2026-10-09
  9. paid-only models notice developers.cloudflare.com · seen 2026-10-09
  10. planned model deprecations notice developers.cloudflare.com · seen 2026-10-09
  11. model catalogue developers.cloudflare.com · seen 2026-10-09
  12. GLM-5.3 model page developers.cloudflare.com · seen 2026-10-09
  13. Gemma 4 26B model page developers.cloudflare.com · seen 2026-10-09
  14. Llama 3.3 70B fast model page developers.cloudflare.com · seen 2026-10-09
  15. GLM-4.7 Flash model page developers.cloudflare.com · seen 2026-10-09
  16. OpenAI-compatible endpoints developers.cloudflare.com · seen 2026-10-09
  17. REST quick start developers.cloudflare.com · seen 2026-10-09
  18. function calling developers.cloudflare.com · seen 2026-10-09
  19. JSON Mode developers.cloudflare.com · seen 2026-10-09
  20. prompt caching developers.cloudflare.com · seen 2026-10-09
  21. batch API developers.cloudflare.com · seen 2026-10-09
  22. batch API over REST developers.cloudflare.com · seen 2026-10-09
  23. reject busy requests developers.cloudflare.com · seen 2026-10-09
  24. API reference, run a model developers.cloudflare.com · seen 2026-10-09
  25. Cloudflare OpenAPI document github.com · seen 2026-10-09
  26. Machine Payments (x402) developers.cloudflare.com · seen 2026-10-09
  27. AI Gateway unified billing developers.cloudflare.com · seen 2026-10-09
  28. AI Gateway changelog developers.cloudflare.com · seen 2026-10-09
  29. AI Gateway logging developers.cloudflare.com · seen 2026-10-09
  30. AI Gateway guardrails developers.cloudflare.com · seen 2026-10-09
  31. API token creation developers.cloudflare.com · seen 2026-10-09
  32. Cloudflare API rate limits developers.cloudflare.com · seen 2026-10-09
  33. status incidents feed cloudflarestatus.com · seen 2026-10-09
  34. status components feed cloudflarestatus.com · seen 2026-10-09
  35. Self-Serve Subscription Agreement cloudflare.com · seen 2026-10-09
  36. service-specific terms, Workers AI and AI Gateway cloudflare.com · seen 2026-10-09
  37. privacy policy cloudflare.com · seen 2026-10-09
  38. website terms of use cloudflare.com · seen 2026-10-09
  39. sub-processor index page cloudflare.com · seen 2026-10-09
  40. security.txt cloudflare.com · seen 2026-10-09
  41. TypeScript SDK tags github.com · seen 2026-10-09
  42. Python SDK tags github.com · seen 2026-10-09
  43. Go SDK tags github.com · seen 2026-10-09
  44. workers-ai-provider on npm registry.npmjs.org · seen 2026-10-09
  45. workers-ai-provider weekly downloads api.npmjs.org · seen 2026-10-09
  46. RDAP for cloudflare.com rdap.verisign.com · seen 2026-10-09

Probe metrics

Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. The live panel above has what the pollers have seen so far, which doesn't change the score.

Pricing & changes

Freemium from $0.0605 / 1M in $0.011 per 1,000 neurons, shown per model as token, image or audio prices, for example Gemma 4 26B at $0.10 in and $0.30 out per 1M tokens. 10,000 neurons a day are free on the Workers Free and Paid plans, and going past that needs Workers Paid, from $5 a month. Seven frontier models need Workers Paid or prepaid AI Gateway credits. A Free plan account starts without a contract (https://developers.cloudflare.com/workers-ai/platform/pricing/, checked 2026-10-09).

Models and prices per million tokens

ModelInputOutputContextRoleSupports
@cf/google/gemma-4-26b-a4b-itGemma 4 26B A4B · 2026-04-04 · the model the newest docs examples use, open to the Free plan, vision, batch and function calling$0.10$0.30256kdefaultnot checked
@cf/zai-org/glm-5.3GLM-5.3 · 2026-08-28 · needs Workers Paid or AI Gateway credits, cached input $0.26$1.40$4.401.05Mflagshipnot checked
@cf/meta/llama-3.3-70b-instruct-fp8-fastLlama 3.3 70B Instruct fp8 fast · batch and function calling$0.293$2.25324kmidnot checked
@cf/zai-org/glm-4.7-flashGLM-4.7 Flash · 2026-02-13 · one of the four models that accept x402 payment$0.0605$0.40131kfastnot checked

Every model here is also on the price index next to the other providers. Rate limits depend on your account tier: Cloudflare, Inc.'s rate limits.

Prices

ItemPriceUnitNote
Whisper large v3 turbo, speech to text$0.0005per minute of audio46.63 neurons
Deepgram Nova-3, speech to text$0.0052per minute of audio472.73 neurons. $0.0092 over WebSocket
Deepgram Aura-1, text to speech$15per 1M characterslisted as $0.015 per 1,000 characters
Deepgram Aura-2, text to speech$30per 1M characterslisted as $0.030 per 1,000 characters, English and Spanish

Compared across listings on the price index.

Dated changes shutdowns, breaking changes, price changes

  • Shutdown 18 models retired, with @cf/moonshotai/kimi-k2.5 aliased to the pricier @cf/moonshotai/kimi-k2.6 source
  • Breaking change @cf/moonshotai/kimi-k2.6, @cf/moonshotai/kimi-k2.7-code and @cf/zai-org/glm-5.2 need the Workers Paid plan source

All of these, for every listing, are on Sunsets and in the calendar feed.

Recent changes

  • Latest release
  • @cf/moonshotai/kimi-k2.6, @cf/moonshotai/kimi-k2.7-code and @cf/zai-org/glm-5.2 need the Workers Paid plan source
  • 18 models retired, with @cf/moonshotai/kimi-k2.5 aliased to the pricier @cf/moonshotai/kimi-k2.6 source

Follow them as a feed at /feeds/tools/cloudflare-workers-ai.xml, or this listing's score history at history.json.

Connect

First request

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/google/gemma-4-26b-a4b-it \
  -X POST \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -d '{ "messages": [{ "role": "system", "content": "You are a friendly assistant" }, { "role": "user", "content": "Why is pizza so good" }]}'

Pay per call with x402

curl -iX POST "https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run" \
  --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
  --header "Payment-Method: x402" \
  --header "Content-Type: application/json" \
  --data '{"model":"z-ai/glm-4.7-flash","input":{"messages":[{"role":"user","content":"What is Cloudflare?"}]}}'
Similar toolGrade ScoreShared capabilitiesx402
DeepInfra Deep Infra Inc.B63inference.llm inference.open-weights embed.text rerank image.generate speech.stt speech.tts compute.batchno
Novita AI Novita AID53.5inference.llm inference.open-weights embed.text rerank image.generate speech.tts compute.batchno
SiliconFlow SiliconFlow Labs Pte. Ltd.D46.7inference.llm inference.open-weights embed.text rerank image.generate speech.tts speech.sttno
LocalAI Ettore Di Giacinto and the LocalAI teamB68inference.open-weights embed.text rerank speech.stt speech.tts image.generateno
Lemonade AMD and the Lemonade communityB63.8inference.open-weights embed.text rerank speech.stt speech.tts image.generateno
KoboldCpp LostRuins (Concedo)C60.5inference.open-weights embed.text image.generate speech.stt speech.ttsno

Machine-readable

Verify this listing

For the vendor

Is this your product? Link to this page from your own site or README, then tell us where. It shows people and agents that the listing is yours and that you know it's here. It never changes a grade, rank or review.

  1. Add the badge or a link

    Cloudflare Workers AI on Anchor Terminal, B, 68.3/100
    On a light page
    On a dark page
    <a href="https://www.anchorterminal.com/tools/cloudflare-workers-ai"><img src="https://www.anchorterminal.com/badges/cloudflare-workers-ai.svg" alt="Cloudflare Workers AI on Anchor Terminal" height="20"></a>
    [![Cloudflare Workers AI on Anchor Terminal](https://www.anchorterminal.com/badges/cloudflare-workers-ai.svg)](https://www.anchorterminal.com/tools/cloudflare-workers-ai)

    It counts on a page on cloudflare.com or one of its subdomains, or the README of github.com/cloudflare/api-schemas.

  2. Tell us where it is

    We read it once now and again every week. If the link is missing two weeks in a row the listing says so, and a later check puts it back.

Agents send the same to POST /api/v1/verify as {"slug": "cloudflare-workers-ai", "url": "…"}, or call the verify_listing tool at /mcp. Ten checks an hour from one address. What we check. To announce the listing, get sharing assets for social media.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.