# Decision models for AI agents > 4 decision models ranked by the Anchor benchmark. Leader Laya (B). Models that answer typed questions with probabilities instead of text, a yes or no, one label from a set or a level on a rubric, for routing, triage and checks inside agents and workflows. Compared on accuracy, calibration, context, latency, price and whether the weights are open. - Canonical: https://www.anchorterminal.com/categories/decision-models - Markdown: https://www.anchorterminal.com/categories/decision-models.md (~1,450 tokens) - Slim: https://www.anchorterminal.com/categories/decision-models.min.md (~330 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/categories/decision-models.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-04 Models that answer typed questions with probabilities instead of text, a yes or no, one label from a set or a level on a rubric, for routing, triage and checks inside agents and workflows. Compared on accuracy, calibration, context, latency, price and whether the weights are open. - Tools ranked: 4 · agent-ready (BB or better): 0 · accept x402: 0 · hosted endpoints: 2 · desk reviews by the panel: 8 - JSON: https://www.anchorterminal.com/api/v1/tools.json (list) · https://www.anchorterminal.com/api/v1/rankings.json (ranked) · https://www.anchorterminal.com/api/v1/x402.json (payable) · https://www.anchorterminal.com/api/v1/capabilities.json (by capability) - Grades run AA, A, BB, B, C, D, E, F · methodology: https://www.anchorterminal.com/benchmark/ - Capabilities in this category: inference.decision - https://letme.dev/inference.decision picks the top-graded tool in this list and says how to call it direct; calling through letme comes later (https://www.anchorterminal.com/letme/index.md) ## Ranking | # | Tool | Vendor | Kind | Category | Grade | Score | Confidence | x402 | Auth | Where | Reviews | Page | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 115 | Laya | Convai Innovations | Model API | Decisions | B | 69.2 | medium | no | None | local | 2.5/5 (2) | https://www.anchorterminal.com/tools/convai-laya.md | | 141 | Kev | Jared Palmer | Model API | Decisions | B | 67.4 | medium | no | None | local | 3.5/5 (2) | https://www.anchorterminal.com/tools/jaredpalmer-kev.md | | 159 | Clef | Cloudflare | Model API | Decisions | B | 66.3 | medium | no | API key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/cloudflare-clef.md | | 221 | Jev | TypeSafe AI | Model API | Decisions | B | 62.2 | medium | no | API key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/typesafe-jev.md | Scores are from public evidence against the published checklist (https://www.anchorterminal.com/benchmark/), with Performance and Task success pending. p95 latency and context cost come from our probes, which haven't run yet. ## Summaries ### 115. Laya, B (69.2) Open-source decision engine from Convai Innovations, released under Apache-2.0. Apache-2.0 code and weights, installed with `pip install laya`, with Python 3.10 to 3.13 tested in CI. Base checkpoints score 0.362 and 0.352 on the maintainers' typed-decisions benchmark against a 0.318 random baseline, so it needs fine-tuning. - Page: https://www.anchorterminal.com/tools/convai-laya · Markdown: https://www.anchorterminal.com/tools/convai-laya.md · JSON: https://www.anchorterminal.com/api/v1/tools/convai-laya.json - Capabilities: inference.decision ### 141. Kev, B (67.4) Kev is a family of four open-weight decision models by Jared Palmer, released together as Kev 1.0 on 1 October 2026 under Apache-2.0. Apache-2.0 code, adapters and heads on Apache-2.0 Qwen bases, with release tarballs and SHA-256 checksums for the 0.8B, 4B and 9B models. No package. `pip install kev` installs an unrelated 2021 ORM, so Kev runs from a Git clone with uv. - Page: https://www.anchorterminal.com/tools/jaredpalmer-kev · Markdown: https://www.anchorterminal.com/tools/jaredpalmer-kev.md · JSON: https://www.anchorterminal.com/api/v1/tools/jaredpalmer-kev.json - Capabilities: inference.decision ### 159. Clef, B (66.3) Clef and Clef-flash are Cloudflare's open-weight decision models, released on 1 October 2026 under Apache-2.0 and hosted on Workers AI. Apache-2.0 weights for both models on Hugging Face, ungated, with no account needed to download. Released on 1 October 2026, with no service record and no entry in the Workers AI changelog we read. - Page: https://www.anchorterminal.com/tools/cloudflare-clef · Markdown: https://www.anchorterminal.com/tools/cloudflare-clef.md · JSON: https://www.anchorterminal.com/api/v1/tools/cloudflare-clef.json - Capabilities: inference.decision · endpoint: `https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef` ### 221. Jev, B (62.2) Jev is TypeSafe AI's first System One model, a closed decision model behind an HTTP API. Typed answers with probabilities for `noul`, `choice` and `score` questions, many per call, with no text to parse. Early access behind a waitlist, with no free tier or free credits found. - Page: https://www.anchorterminal.com/tools/typesafe-jev · Markdown: https://www.anchorterminal.com/tools/typesafe-jev.md · JSON: https://www.anchorterminal.com/api/v1/tools/typesafe-jev.json - Capabilities: inference.decision · endpoint: `https://api.typesafe.ai/v1/systemone` ## How we test this category A fixed set of labelled decisions (yes or no, one label from a set, a level on a rubric) sent to every model in the same request shape, with the same states and questions. We check accuracy against the labels, calibration (expected calibration error, Brier score and how often an answer given 0.9 or more is wrong), whether answers move when the option order changes, latency and the cost per 1,000 decisions. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.