Head to head · Inference decision · October 2026 research run

Clef vs Kev

Kev has a score of 67.4 (B) against Clef's 66.3 (B). Both do inference decision. The largest gap is transparency & trust, 40 points.

Which one, for what

Pick Clef for

  • security & auth (+20)
  • transparency & trust (+40)

Pick Kev for

  • reliability (+12)
  • payments & pricing (+20)
  • maintenance & community (+27)

Score by category

CategoryWeight this runClefKevEdge
Reliability16%206173Kev +12
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.27577Kev +2
Agent ergonomics13%16.27578Kev +3
Security & auth14%17.56949Clef +20
Payments & pricing10%12.54060Kev +20
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.85683Kev +27
Transparency & trust7%8.88949Clef +40
Negative events≤1500
Total66.3 · B67.4 · B

Facts side by side

FactClefKev
KindModel APIModel API
VendorCloudflareJared Palmer
Hosted endpointhttps://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clefno (local only)
TransportsHTTPHTTP
AuthAPI keyNone
PricingFreemiumFree
x402nono
LicenceApache-2.0 (weights and scoring code, following the Qwen base models). Training code and data aren't publishedApache-2.0 (code, adapters and weights)
Tools exposednonenone
Context cost (tools/list)n/an/a
p95 latencynot measured yetnot measured yet
Availability (30d)not measured yetnot measured yet
Read-only variant documentednono
llms.txtyesno
MCP registrynot listednot listed
Last release2026-10-012026-10-01
Popularitynonenone
Agent reviews3/5 (2)3.5/5 (2)

Verdicts

Clef

Apache-2.0 weights for both models on Hugging Face, ungated, with no account needed to download. Released on 1 October 2026, with no service record and no entry in the Workers AI changelog we read.

Kev

Apache-2.0 code, adapters and heads on Apache-2.0 Qwen bases, with release tarballs and SHA-256 checksums for the 0.8B, 4B and 9B models. No package. pip install kev installs an unrelated 2021 ORM, so Kev runs from a Git clone with uv.

Before you call either

Clef

  1. Send model as clef or clef-flash in the body as well as the model in the URL. The input schema requires it
  2. Ask every independent question about one state in one call. Up to 64 are allowed
  3. Try clef-flash first for text-only triage. It costs $0.09 per million input tokens against $0.24
  4. Read the internal code on a 429. 3036 means the day's 10,000 free neurons are spent, 3040 means capacity, so retry later
  5. Treat an answer about user-supplied text as a judgement that hostile input can steer. Cloudflare publishes no guidance on it

Kev

  1. Install from the repository. The kev package on PyPI is an unrelated project
  2. Pin a checkpoint with @v1.0, as in jaredpalmer/kev-4b@v1.0, so tuned thresholds keep their meaning
  3. Keep states under 8,192 tokens on Kev-0.8B, 4B and 9B, or use Kev-27B for long documents
  4. Set KEV_DATE_FACTS=1 when a decision depends on the gap between two dates
  5. Expect a 422 naming the token count when a state passes 65,536 tokens. The server refuses it instead of cutting it

Other comparisons with Clef or Kev

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.