# Cloudflare Workers AI (slim) > Workers AI is Cloudflare's serverless inference service for open-weight models, covering text generation, embeddings, images and speech. Agents call it through the Cloudflare REST API, OpenAI-compatible endpoints or a Workers binding. - Full: https://www.anchorterminal.com/tools/cloudflare-workers-ai.md (~10,550 tokens) · this version ~2,280 tokens · JSON https://www.anchorterminal.com/tools/cloudflare-workers-ai.json · canonical https://www.anchorterminal.com/tools/cloudflare-workers-ai - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-10 **B · 68.3/100 · rank #224 of 950 · #7 in Model APIs & inference · not agent-ready · confidence medium** Assessment: Per-model prices, rate limits and JSON Schemas are public, 10,000 neurons a day are free, and x402 payment is in beta on `/ai/run` for four models. No SLA was found, three models moved to paid-only access on 28 July 2026 with no notice, and several docs pages still use a model retired in May. ## Facts - Kind: Model API · vendor: Cloudflare, Inc. · category: Model APIs & inference · legal entity: Cloudflare, Inc. · provenance 95/100 - Endpoint: `https://api.cloudflare.com/client/v4/accounts/{account_id}/ai` (HTTP) - Auth: API key · pricing: Freemium · x402: payer tooling only · licence: Proprietary service under Cloudflare's Self-Serve Subscription Agreement. Each hosted model carries its own open-weight licence, linked from its model page. The `cloudflare` SDKs are Apache-2.0, `workers-ai-provider` is MIT and the OpenAPI repository is BSD-3-Clause - Probe metrics: not measured yet (probes haven't run) - Endpoints: Native `POST /accounts/{account_id}/ai/run/{model_name}`, OpenAI-style `/ai/v1/chat/completions` and `/ai/v1/embeddings`, `/ai/v1/responses` for the two GPT-OSS models only (no streaming), and the shared `POST /ai/run` route that takes the model in the body. All under https://api.cloudflare.com/client/v4 - Models on 9 October 2026: 69 in the catalogue. The pricing page lists 37 text generation models, 6 embedding models, 6 image models, 10 audio entries and 8 others, among them a reranker, two translation models and the Clef decision models - Rate limits: Per account by task. Text generation 300 requests a minute, embeddings 3,000 (1,500 for `bge-large-en-v1.5`), speech recognition 720, text to image 720, translation 720. Models that need Workers Paid get 20 a minute per model, or 50 with prepaid AI Gateway credits. Beta models may be lower - Errors: 17 internal codes with HTTP statuses. 429 is 3036 (daily free allocation spent) or 3040 (capacity), 403 with 5035 means the model needs Workers Paid, 408 is 3007 (timeout) or 3008 (aborted) - Capacity controls: `options.rejectIfBusy` fails a synchronous call with 429 and 3040 when no capacity is free, in place of waiting in a queue. The batch route (`?queueRequest=true`) queues requests and returns a `request_id` to poll - Batch: Asynchronous, pull-based since 19 March 2026, on models tagged Batch, with a 10 MB payload limit. No batch discount is stated - Prompt caching: Prefix caching is on by default for select models. The `x-session-affinity` header keeps a session on one model instance, cached tokens appear in `usage`, and cached input has its own price (GLM-5.3 $0.26 against $1.40 per 1M) - Structured output: `response_format` of `json_object` or `json_schema`. The docs say adherence to the schema is not guaranteed and JSON Mode does not stream - Tool calling: OpenAI-style `tools` with `parallel_tool_calls` on models tagged Function calling, plus embedded function calling in Workers through `@cloudflare/ai-utils` - Free tier: 10,000 neurons a day on the Workers Free and Paid plans, reset at 00:00 UTC. Past that, $0.011 per 1,000 neurons on Workers Paid, which starts at $5 a month - Trains on API data: No. The data usage page and the service-specific terms say Customer Content is not used for training without consent - Logs: Requests routed through an AI Gateway are logged with prompt, response, tokens and cost, on by default per gateway. The `cf-aig-collect-log` header overrides the setting for one request - SDKs: The OpenAI SDKs with a changed base URL. Cloudflare's own `cloudflare` SDKs, TypeScript 7.3.0 and Python 5.9.0 (both tagged 2 October 2026) and Go v7.12.0. `workers-ai-provider` 4.0.0 (22 July 2026) for the Vercel AI SDK - Status: www.cloudflarestatus.com on Statuspage, with Workers AI, AI Gateway and API as separate components, all operational on 9 October 2026 - Models ($/1M in/out): `@cf/google/gemma-4-26b-a4b-it` $0.10/$0.30; `@cf/zai-org/glm-5.3` $1.40/$4.40; `@cf/meta/llama-3.3-70b-instruct-fp8-fast` $0.293/$2.253; `@cf/zai-org/glm-4.7-flash` $0.0605/$0.40 - Prices: Whisper large v3 turbo, speech to text $0.0005 per minute of audio; Deepgram Nova-3, speech to text $0.0052 per minute of audio; Deepgram Aura-1, text to speech $15 per 1M characters; Deepgram Aura-2, text to speech $30 per 1M characters - 2026-05-30 Shutdown: 18 models retired, with `@cf/moonshotai/kimi-k2.5` aliased to the pricier `@cf/moonshotai/kimi-k2.6` - 2026-07-28 Breaking change: `@cf/moonshotai/kimi-k2.6`, `@cf/moonshotai/kimi-k2.7-code` and `@cf/zai-org/glm-5.2` need the Workers Paid plan - Scores: Reliability 66, Performance pending, Schema & documentation 80, Agent ergonomics 75, Security & auth 77, Payments & pricing 52, Task success pending, Maintenance & community 72, Transparency & trust 76 · negative events -3 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, Model reading, scored on the API lines. · Agent ergonomics, Model reading of the checklist (tool use, structured output, caching, context, batch, SDKs, errors). · Security & auth, Model reading. · Payments & pricing, x402 is in beta on `POST /ai/run` since 30 September 2026 for four of 69 models. · Maintenance & community, Model reading. · Transparency & trust, Closed service under the Self-Serve Subscription Agreement and service-specific terms with a Workers AI section. - Sources: 46, open questions: 10, both in the full twin - Capabilities: inference.llm, inference.open-weights, embed.text, rerank, image.generate, speech.stt, speech.tts, compute.batch - JSON: https://www.anchorterminal.com/api/v1/tools/cloudflare-workers-ai.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/cloudflare-workers-ai.svg` or a link to https://www.anchorterminal.com/tools/cloudflare-workers-ai from a page on cloudflare.com or one of its subdomains, or the README of github.com/cloudflare/api-schemas, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Check the model page before calling. Seven models need Workers Paid or AI Gateway credits and return 403 with code 5035 on the Free plan 2. Read the internal code on a 429. 3036 means the day's 10,000 free neurons are spent until 00:00 UTC, 3040 means capacity, so retry later 3. Set `options.rejectIfBusy` to fail fast, or add `?queueRequest=true` for the batch route when the answer can wait 4. Send the same `x-session-affinity` value on every turn of a session to reach the prefix cache and the cached-input price 5. Keep paid frontier models under 20 requests a minute per model, or 50 with prepaid AI Gateway credits ## Connect ```bash curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/google/gemma-4-26b-a4b-it \ -X POST \ -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \ -d '{ "messages": [{ "role": "system", "content": "You are a friendly assistant" }, { "role": "user", "content": "Why is pizza so good" }]}' ``` x402: ```bash curl -iX POST "https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run" \ --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \ --header "Payment-Method: x402" \ --header "Content-Type: application/json" \ --data '{"model":"z-ai/glm-4.7-flash","input":{"messages":[{"role":"user","content":"What is Cloudflare?"}]}}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | DeepInfra | B | 63 | inference.llm, inference.open-weights, embed.text, rerank, image.generate, speech.stt, speech.tts, compute.batch | https://www.anchorterminal.com/tools/deepinfra.min.md | | Novita AI | D | 53.5 | inference.llm, inference.open-weights, embed.text, rerank, image.generate, speech.tts, compute.batch | https://www.anchorterminal.com/tools/novita-ai.min.md | | SiliconFlow | D | 46.7 | inference.llm, inference.open-weights, embed.text, rerank, image.generate, speech.tts, speech.stt | https://www.anchorterminal.com/tools/siliconflow.min.md | | LocalAI | B | 68 | inference.open-weights, embed.text, rerank, speech.stt, speech.tts, image.generate | https://www.anchorterminal.com/tools/localai.min.md | | Lemonade | B | 63.8 | inference.open-weights, embed.text, rerank, speech.stt, speech.tts, image.generate | https://www.anchorterminal.com/tools/lemonade.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)