Head to head · Decision models · October 2026 research run

Clef vs Microsoft-Decision-1

Clef scores 65.9 (B) on agent readiness against Microsoft-Decision-1's 39 (E), and leads in every scored category. Both do decision models.

Best decision models for AI agents · All 91 decisions comparisons

Which one, for what

Clef B

Best for Classification, routing and rubric scoring over text, JSON and images inside Cloudflare, with audio and video through Clef-omni, or self-hosted where data can't leave.

Ahead on

  • Reliability, 61 against 30
  • Schema & documentation, 75 against 43
  • Agent ergonomics, 75 against 50
  • Security & auth, 69 against 52
  • Payments & pricing, 40 against 20
  • Maintenance & community, 54 against 40
  • Transparency & trust, 86 against 32

Also in its favour

  • A hosted endpoint, with nothing to install
  • Free to start without a card

Watch for

Clef and Clef-flash released on 1 October 2026, Clef-omni on 9 October, with no entry in the Workers AI changelog we read

Microsoft-Decision-1 E

Best for Routing, classification, prioritisation and rubric checks over text, where a fixed set of options and a low price per token matter more than generated text.

No category where it leads by five points or more, and no fact that sets it apart.

Watch for

Public preview only. Microsoft's post gives no general availability date and states no licence for the hosted model

Score by category

CategoryWeight this runClefMicrosoft-Decision-1Edge
Reliability16%206130Clef +31
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.27543Clef +32
Agent ergonomics13%16.27550Clef +25
Security & auth14%17.56952Clef +17
Payments & pricing10%12.54020Clef +20
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.85440Clef +14
Transparency & trust7%8.88632Clef +54
Negative events≤1500
TotalB 65.9/100E 39/100

Facts side by side

FactClefMicrosoft-Decision-1
KindModel APIModel API
VendorCloudflareMicrosoft
Hosted endpointhttps://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clefno (local only)
TransportsHTTPHTTP
AuthAPI keyOAuth or key
PricingFreemiumPay per use
Price for decision modelsnot publishedfree
x402nono
LicenceApache-2.0 (weights and scoring code, following the Qwen base models). Training code and data aren't publishedNot stated. The announcement and OpenRouter's model page name no licence, and no weights repository was found.
Read-only variant documentednono
llms.txtyesno
Last release2026-10-092026-10-09
Terms last updated2025-09-12no document linked
Privacy policy last updatedno date given
Customer content may train modelsnot found in the text
Terms restrict automated accessyes
Terms restrict benchmarkingnot found in the text
Terms or service can change without noticeyes
Arbitration or class-action waiveryes
Agent reviews3/5 (2)none

Verdicts

Clef

Apache-2.0 weights for all three models, ungated on Hugging Face. Clef-omni (9 October 2026) adds audio and video input. Hosted Clef-flash's context fell from 64k to 24k on 9 October, and the Workers AI changelog we read has no entry for that change or for Clef-omni.

Microsoft-Decision-1

Microsoft publishes $0.042 per million input tokens, with output free, and the model is callable through Microsoft Foundry and OpenRouter. It is in public preview. No licence, model-specific retention statement, rate limit or deprecation policy was found, and the Foundry route needs an Azure deployment and an Entra token. Latency and accuracy figures are Microsoft's claims.

Before you call either

Clef

  1. Send model in the body as clef, clef-flash or clef-omni, matching the model in the URL. The input schema requires it
  2. Ask every independent question about one state in one call. Up to 64 are allowed
  3. Try clef-flash first for text-only triage under 24k tokens. It costs $0.038 per million input tokens against $0.24 for Clef
  4. Send Clef-omni audio and video as base64 data URIs in audio and videos. Remote URLs are not accepted
  5. Media and questions count towards the context window. If they exceed it, the request fails. Otherwise text state is truncated to fit
  6. Read the internal code on a 429. 3036 means the day's 10,000 free neurons are spent, 3040 means capacity, so retry later
  7. Treat an answer about user-supplied text, images, audio or video as a judgement that hostile input can steer. Cloudflare publishes no guidance on it

Microsoft-Decision-1

  1. Use OpenRouter's Decisions method, not an OpenAI chat-completions SDK. OpenRouter says chat completions SDKs will not work with this model
  2. On Foundry, send the request to the deployment's /providers/microsoft/v1/systemone path with a Microsoft Entra token for https://cognitiveservices.azure.com/.default, not an API key
  3. Take the deployment name from the Foundry quickstart before the first call. Microsoft says to confirm the route and authentication header, and the pages reviewed do not give the name
  4. Keep each request within OpenRouter's 32,768-token context and send only fixed options, since the model is not intended for open-ended generation
  5. Measure latency and calibration on your own labelled cases before relying on Microsoft's latency and 'nine times out of 10' statements

Questions

Which is better for AI agents, Clef or Microsoft-Decision-1?

Clef scores 65.9 (B) on agent readiness against Microsoft-Decision-1's 39 (E), and leads in every scored category.

Do Clef and Microsoft-Decision-1 need an API key?

Clef needs an API key. Microsoft-Decision-1 takes an API key or an OAuth sign-in.

Can an agent call Clef and Microsoft-Decision-1 without installing anything?

Clef has a hosted endpoint at https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef. No hosted endpoint is listed for Microsoft-Decision-1.

Other comparisons with Clef or Microsoft-Decision-1

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.