Head to head · LLM inference · October 2026 research run

Cloudflare Workers AI vs DeepInfra

Cloudflare Workers AI scores 68.3 (B) on agent readiness against DeepInfra's 63 (B), and leads in 5 of 7 scored categories. Both do llm inference.

Best model APIs and inference for AI agents · All 136 models comparisons

Which one, for what

Cloudflare Workers AI B

Good for Agents that already run on Cloudflare Workers, or that want open-weight chat, embedding, reranking, speech and image models behind one token with a free daily allowance.

Ahead on

  • Schema & documentation, 80 against 69
  • Security & auth, 77 against 64
  • Payments & pricing, 52 against 20
  • Maintenance & community, 72 against 65
  • Transparency & trust, 76 against 67

Watch for

No SLA for Workers AI was found in the docs or the agreements read

DeepInfra B

Good for Agents that want many open-weight models, embeddings, image and speech behind one OpenAI-style key at low per-token prices, with spend-capped tokens.

Also in its favour

  • No incidents deducted, where Cloudflare Workers AI loses 3 points for them

Watch for

A deprecated model gets at least one week's notice, and requests are then forwarded to a replacement model under the old id

Score by category

CategoryWeight this runCloudflare Workers AIDeepInfraEdge
Reliability16%206670DeepInfra +4
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28069Cloudflare Workers AI +11
Agent ergonomics13%16.27577DeepInfra +2
Security & auth14%17.57764Cloudflare Workers AI +13
Payments & pricing10%12.55220Cloudflare Workers AI +32
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.87265Cloudflare Workers AI +7
Transparency & trust7%8.87667Cloudflare Workers AI +9
Negative events≤15-30
Total68.3 · B63 · B

Facts side by side

FactCloudflare Workers AIDeepInfra
KindModel APIModel API
VendorCloudflare, Inc.Deep Infra Inc.
Hosted endpointhttps://api.cloudflare.com/client/v4/accounts/{account_id}/aihttps://api.deepinfra.com/v1/openai
TransportsHTTPHTTP
AuthAPI keyAPI key
PricingFreemiumPay per use
x402payer tooling onlyno
LicenceProprietary service under Cloudflare's Self-Serve Subscription Agreement. Each hosted model carries its own open-weight licence, linked from its model page. The cloudflare SDKs are Apache-2.0, workers-ai-provider is MIT and the OpenAPI repository is BSD-3-ClauseProprietary service under the DeepInfra Terms of Service. The Python and Node SDKs and the docs repository are MIT
Read-only variant documentednono
llms.txtyesyes
Last release2026-10-012026-10-07
Terms last updated2025-09-122026-08-17
Privacy policy last updatedno date given2026-08-15
Customer content may train modelsnot found in the textnot found in the text
Terms restrict automated accessyesnot found in the text
Terms restrict benchmarkingnot found in the textyes
Terms or service can change without noticeyesnot found in the text
Arbitration or class-action waiveryesyes
Popularity380k npm/wk21 stars, 1.2k npm/wk, 50 PyPI/wk

Verdicts

Cloudflare Workers AI

Per-model prices, rate limits and JSON Schemas are public, 10,000 neurons a day are free, and x402 payment is in beta on /ai/run for four models. No SLA was found, three models moved to paid-only access on 28 July 2026 with no notice, and several docs pages still use a model retired in May.

DeepInfra

The model list, context sizes and per-token prices are readable without a key, and keys can carry an IP allowlist, a monthly spending cap and model-limited JWTs. Deprecated models get one week's notice and are then redirected to another model, there is no changelog or SLA, and an account needs a card or prepayment before any call.

Before you call either

Cloudflare Workers AI

  1. Check the model page before calling. Seven models need Workers Paid or AI Gateway credits and return 403 with code 5035 on the Free plan
  2. Read the internal code on a 429. 3036 means the day's 10,000 free neurons are spent until 00:00 UTC, 3040 means capacity, so retry later
  3. Set options.rejectIfBusy to fail fast, or add ?queueRequest=true for the batch route when the answer can wait
  4. Send the same x-session-affinity value on every turn of a session to reach the prefix cache and the cached-input price
  5. Keep paid frontier models under 20 requests a minute per model, or 50 with prepaid AI Gateway credits

DeepInfra

  1. Call GET https://api.deepinfra.com/v1/openai/models at start-up for ids, context sizes and prices. No key is needed
  2. Check the model field of each response. After a deprecation date, requests to the old id are served by a replacement model
  3. Ask the account owner for a scoped JWT limited to the models and spend the task needs, not the full API key
  4. Stay under 200 concurrent requests per model. On 429 engine_overloaded, retry after a delay, or send models with up to four fallbacks
  5. Never inspect a JWT with GET /v1/scoped-jwt?jwtoken=, which puts the token in the URL. Keep credentials in the Authorization header

Questions

Which is better for AI agents, Cloudflare Workers AI or DeepInfra?

Cloudflare Workers AI scores 68.3 (B) on agent readiness against DeepInfra's 63 (B), and leads in 5 of 7 scored categories.

Do Cloudflare Workers AI and DeepInfra need an API key?

Both need an API key.

Can an agent call Cloudflare Workers AI and DeepInfra without installing anything?

Yes. Cloudflare Workers AI has a hosted endpoint at https://api.cloudflare.com/client/v4/accounts/{account_id}/ai and DeepInfra at https://api.deepinfra.com/v1/openai.

Other comparisons with Cloudflare Workers AI or DeepInfra

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.