Head to head · LLM inference · October 2026 research run

Cloudflare Workers AI vs Prism Inference

Cloudflare Workers AI scores 68.3 (B) on agent readiness against Prism Inference's 60.1 (C), and leads in 6 of 7 scored categories. Both do llm inference.

Best model APIs and inference for AI agents · All 136 models comparisons

Which one, for what

Cloudflare Workers AI B

Good for Agents that already run on Cloudflare Workers, or that want open-weight chat, embedding, reranking, speech and image models behind one token with a free daily allowance.

Ahead on

  • Agent ergonomics, 75 against 68
  • Security & auth, 77 against 65
  • Payments & pricing, 52 against 30
  • Maintenance & community, 72 against 49
  • Transparency & trust, 76 against 61

Watch for

No SLA for Workers AI was found in the docs or the agreements read

Prism Inference C

Good for Coding agents that want DeepSeek-V4.1-Flash at a low input price, with no retention, through whichever of the three wire formats the harness already speaks.

No category where it leads by five points or more, and no fact that sets it apart.

Watch for

Two models. The docs mark Gemma 4 31B as request access per organisation, while llms.txt and the keyless catalogue list it as available

Score by category

CategoryWeight this runCloudflare Workers AIPrism InferenceEdge
Reliability16%206665Cloudflare Workers AI +1
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28082Prism Inference +2
Agent ergonomics13%16.27568Cloudflare Workers AI +7
Security & auth14%17.57765Cloudflare Workers AI +12
Payments & pricing10%12.55230Cloudflare Workers AI +22
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.87249Cloudflare Workers AI +23
Transparency & trust7%8.87661Cloudflare Workers AI +15
Negative events≤15-3-2
Total68.3 · B60.1 · C

Facts side by side

FactCloudflare Workers AIPrism Inference
KindModel APIModel API
VendorCloudflare, Inc.Prism Technologies Inc
Hosted endpointhttps://api.cloudflare.com/client/v4/accounts/{account_id}/aihttps://api.prisminference.com/v1
TransportsHTTPHTTP
AuthAPI keyAPI key
PricingFreemiumPay per use
x402payer tooling onlyno
LicenceProprietary service under Cloudflare's Self-Serve Subscription Agreement. Each hosted model carries its own open-weight licence, linked from its model page. The cloudflare SDKs are Apache-2.0, workers-ai-provider is MIT and the OpenAPI repository is BSD-3-ClauseProprietary service under Prism's terms of service. The OpenAPI file declares LicenseRef-Proprietary. The Hermes provider plugin repository carries no licence file
Read-only variant documentednono
llms.txtyesyes
Last release2026-10-012026-10-06
Terms last updated2025-09-122026-09-09
Privacy policy last updatedno date given2026-09-09
Customer content may train modelsnot found in the textnot found in the text
Terms restrict automated accessyesyes
Terms restrict benchmarkingnot found in the textnot found in the text
Terms or service can change without noticeyesnot found in the text
Arbitration or class-action waiveryesyes
Popularity380k npm/wknone

Verdicts

Cloudflare Workers AI

Per-model prices, rate limits and JSON Schemas are public, 10,000 neurons a day are free, and x402 payment is in beta on /ai/run for four models. No SLA was found, three models moved to paid-only access on 28 July 2026 with no notice, and several docs pages still use a model retired in May.

Prism Inference

Three wire formats, a public OpenAPI 3.1 file, per-token prices in a keyless catalogue and zero data retention by default on every tier. The service launched on 24 September 2026 with two models, one of them by request, from a two-person company. No rate-limit numbers, SLA document, free tier or deprecation policy was found.

Before you call either

Cloudflare Workers AI

  1. Check the model page before calling. Seven models need Workers Paid or AI Gateway credits and return 403 with code 5035 on the Free plan
  2. Read the internal code on a 429. 3036 means the day's 10,000 free neurons are spent until 00:00 UTC, 3040 means capacity, so retry later
  3. Set options.rejectIfBusy to fail fast, or add ?queueRequest=true for the batch route when the answer can wait
  4. Send the same x-session-affinity value on every turn of a session to reach the prefix cache and the cached-input price
  5. Keep paid frontier models under 20 requests a minute per model, or 50 with prepaid AI Gateway credits

Prism Inference

  1. Call GET https://api.prisminference.com/v1/models at start-up, with no key, and use only ids it returns. Expect 403 on gemma-4-31b without organisation access
  2. Use base URL https://api.prisminference.com/v1 for OpenAI clients and https://api.prisminference.com with no /v1 for Anthropic clients
  3. Read error.retryable before retrying, and wait for Retry-After on 429, which covers both key limits and model capacity
  4. Send reasoning_effort: "none" or low when latency matters. Reasoning is on by default and its tokens are billed as output
  5. Keep conversation state yourself and send store: false on Responses. previous_response_id, stored responses and hosted tools aren't supported

Questions

Which is better for AI agents, Cloudflare Workers AI or Prism Inference?

Cloudflare Workers AI scores 68.3 (B) on agent readiness against Prism Inference's 60.1 (C), and leads in 6 of 7 scored categories.

Do Cloudflare Workers AI and Prism Inference need an API key?

Both need an API key.

Can an agent call Cloudflare Workers AI and Prism Inference without installing anything?

Yes. Cloudflare Workers AI has a hosted endpoint at https://api.cloudflare.com/client/v4/accounts/{account_id}/ai and Prism Inference at https://api.prisminference.com/v1.

Other comparisons with Cloudflare Workers AI or Prism Inference

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.