# Cloudflare AI Gateway (slim) > AI Gateway is Cloudflare's proxy and router for model providers. Agents call OpenAI-style, Anthropic-style or native endpoints on the Cloudflare API, with logging, caching, rate and spend limits, retries, fallbacks and one prepaid credit balance. - Full: https://www.anchorterminal.com/tools/cloudflare-ai-gateway.md (~11,200 tokens) · this version ~2,480 tokens · JSON https://www.anchorterminal.com/tools/cloudflare-ai-gateway.json · canonical https://www.anchorterminal.com/tools/cloudflare-ai-gateway - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-10 **B · 66.8/100 · rank #275 of 950 · #8 in Model APIs & inference · not agent-ready · confidence medium** Assessment: The gateway layer is free, per-model prices are public with no markup, and logs, spend limits and fallbacks are set per gateway. API tokens cannot be limited to one gateway, the service terms take third-party models outside Cloudflare's DPA, no SLA names the product, and error codes changed on 6 October 2026 with same-day notice. ## Facts - Kind: Model router · vendor: Cloudflare, Inc. · category: Model APIs & inference · legal entity: Cloudflare, Inc. · provenance 95/100 - Endpoint: `https://api.cloudflare.com/client/v4/accounts/{account_id}/ai` (HTTP) - Auth: API key · pricing: Freemium · x402: payer tooling only · licence: Proprietary service under Cloudflare's Self-Serve Subscription Agreement and its service-specific terms. Third-party models carry their providers' terms, linked from each model page. `ai-gateway-provider` is MIT and Cloudflare's OpenAPI repository is BSD-3-Clause - Probe metrics: not measured yet (probes haven't run) - Endpoints: `POST /ai/run` (all models and modalities, model in the body), `/ai/v1/chat/completions` (OpenAI), `/ai/v1/responses` (OpenAI Responses) and `/ai/v1/messages` (Anthropic), all under https://api.cloudflare.com/client/v4/accounts/{account_id}. Provider-native endpoints at https://gateway.ai.cloudflare.com/v1/{account_id}/{gateway_id}/{provider} keep each provider's own schema - Management API: 36 paths and 67 operations under `/accounts/{account_id}/ai-gateway` in Cloudflare's OpenAPI 3.0.3 document, covering gateways, logs, dynamic routes, provider keys, custom providers, custom domains and billing (credit balance, top-up, usage history). 13 operations for datasets, evaluations and the older spending limit are marked deprecated - Models on 9 October 2026: 241 in the shared catalogue. Each model page gives the context window, request formats, price and whether Zero Data Retention applies. Claude Sonnet 5 is listed at $2.00 in and $10.00 out per 1M tokens, GPT-4.1 mini at $0.40 and $1.60 - Credentials for the upstream provider: Resolved in order. A provider key on the request, then a key stored on the gateway under the `default` alias (BYOK, held in Secrets Store), then Unified Billing with Cloudflare-managed credentials. `byok_only` or `cf-aig-no-wholesale: true` stops the last step and returns 400 - Rate limits: Unified Billing requests are limited to 200 per 60 seconds per gateway and return 429 past that. BYOK requests are not subject to that limit. The Cloudflare API allows 1,200 requests per five minutes per token. A gateway's own rate limit is set by its owner, fixed or sliding window - Spend limits: Up to 20 rules a gateway, each a dollar budget over a fixed or sliding window, scoped by model, provider or metadata key. Over budget returns 429. Enforcement is eventually consistent, so a burst can pass the limit briefly - Retries and fallbacks: `cf-aig-max-attempts` (up to 5), `cf-aig-retry-delay` (up to 60,000 ms) and `cf-aig-backoff` (constant, linear, exponential) per request, or a retry policy on the gateway. Dynamic routes add conditions, percentage splits, rate and budget nodes and model fallbacks, and accept the Chat Completions shape only - Auto Router: `cloudflare/auto` in place of a model name picks a model by task and cost, for Chat Completions and Responses. Response headers `cf-aig-routed-model` and `cf-aig-routing-reason` name the choice - Caching: Off by default. Identical text and image requests only, TTL up to one month, request size up to 25 MB, with `cf-aig-cache-ttl`, `cf-aig-cache-key` and `cf-aig-skip-cache` per request. Streaming responses are not cached by default - Background requests: `options.background` with a `webhookUrl` on `/ai/run` returns at once and posts the result to the webhook. Delivery is best-effort and not retried - Logs: On by default per gateway. A log holds prompt, response, provider, status, tokens, cost, duration and user agent. `cf-aig-collect-log` and `cf-aig-collect-log-payload` override per request. Logpush export needs Workers Paid, with 4 jobs an account - Guardrails and DLP: Guardrails score prompts and responses with `@cf/meta/llama-guard-3-8b` and flag or block by category, billed as Workers AI inference. DLP scanning is free, with two predefined profiles for accounts without a Zero Trust subscription - Zero Data Retention: Routes Unified Billing traffic to provider endpoints that do not retain prompts or responses, on models the catalogue marks for it. It does not apply to BYOK and does not switch off AI Gateway's own logs - Trains on API data: The service-specific terms say Cloudflare does not use Customer Content to train generative AI tools unless otherwise agreed. Third-party model providers act under their own terms - SDKs: The OpenAI and Anthropic SDKs with a changed base URL. Cloudflare's `cloudflare` SDKs, TypeScript v7.3.0 (tagged 2 October 2026), Python v5.9.0 and Go v7.12.0. `ai-gateway-provider` 4.0.1 (11 September 2026) for the Vercel AI SDK - Limits: 10 gateways an account on the free plan and 20 on paid, 5 custom metadata entries a request, 64 characters in a gateway name (https://developers.cloudflare.com/ai-gateway/reference/limits/) - Status: www.cloudflarestatus.com on Statuspage, with AI Gateway, Workers AI and API as separate components, all operational on 9 October 2026 - Prices: Unified Billing credit purchase fee 5% percentage fee; Logpush requests past 10 million a month $0.0001 per 1,000 requests - Scores: Reliability 66, Performance pending, Schema & documentation 75, Agent ergonomics 73, Security & auth 75, Payments & pricing 52, Task success pending, Maintenance & community 75, Transparency & trust 73 · negative events -3 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, Router, scored on the API lines. · Agent ergonomics, Router, read on the model lines (tool use, structured output, caching, context, batch, SDKs, errors). · Security & auth, Router, read on the model lines. · Payments & pricing, x402 is in beta on `POST /ai/run` since 30 September 2026 for four open models of 241 in the catalogue. · Maintenance & community, Router, read on the model lines. · Transparency & trust, Closed service under the Self-Serve Subscription Agreement and service-specific terms with a section for Workers AI and AI Gateway. - Sources: 52, open questions: 16, both in the full twin - Capabilities: inference.llm, inference.router, obs.gateway, guard.moderation, guard.pii - JSON: https://www.anchorterminal.com/api/v1/tools/cloudflare-ai-gateway.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/cloudflare-ai-gateway.svg` or a link to https://www.anchorterminal.com/tools/cloudflare-ai-gateway from a page on cloudflare.com or one of its subdomains, or the README of github.com/cloudflare/ai, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Use the REST API at `api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1` with a token holding Workers AI Read. A token with only an AI Gateway permission returns 401 with code 10000 2. Send `cf-aig-gateway-id` on every Workers AI (`@cf/`) call and on dynamic routes. Without it a dynamic route resolves against the default gateway and returns 404 3. Set `cf-aig-collect-log: false` or `cf-aig-collect-log-payload: false` on sensitive requests. Logging of prompts and responses is on by default 4. Treat a 429 as a rate limit or a spent budget and a 401 with code 2009 as a rejected provider key. Do not retry a 2009 5. Set `byok_only` on the gateway or send `cf-aig-no-wholesale: true` to stop a request without a provider key falling through to paid Unified Billing credits ## Connect ```bash curl -X POST "https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/v1/chat/completions" \ --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \ --header "Content-Type: application/json" \ --data '{ "model": "openai/gpt-4.1-mini", "messages": [{"role": "user", "content": "What is Cloudflare?"}] }' ``` x402: ```bash curl -iX POST "https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run" \ --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \ --header "Payment-Method: x402" \ --header "Content-Type: application/json" \ --data '{"model":"z-ai/glm-4.7-flash","input":{"messages":[{"role":"user","content":"What is Cloudflare?"}]}}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Google Cloud Model Armor | BB | 77.9 | guard.pii, guard.moderation | https://www.anchorterminal.com/tools/google-model-armor.min.md | | Amazon Bedrock Guardrails | BB | 74.8 | guard.pii, guard.moderation | https://www.anchorterminal.com/tools/amazon-bedrock-guardrails.min.md | | OpenAI Guardrails | B | 69.5 | guard.pii, guard.moderation | https://www.anchorterminal.com/tools/openai-guardrails.min.md | | OpenRouter | B | 68.5 | inference.llm, inference.router | https://www.anchorterminal.com/tools/openrouter.min.md | | NVIDIA NeMo Guardrails | B | 68.4 | guard.pii, guard.moderation | https://www.anchorterminal.com/tools/nemo-guardrails.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)