# Best decision models for AI agents (slim) > OpenAI Decisions API (BB), Decider (B) and Laya (B) lead the 12 ranked decision models. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Full: https://www.anchorterminal.com/best/decision-models/index.md (~6,100 tokens) · this version ~1,480 tokens · JSON https://www.anchorterminal.com/best/decision-models/index.json · canonical https://www.anchorterminal.com/best/decision-models/ - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 The 10 highest-scoring of 12 decision models on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 12 · agent-ready (BB or better): 1 · accept x402: 0 · hosted endpoints: 6 - Full ranked table: https://www.anchorterminal.com/categories/decision-models.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/decision-models/index.md (66) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [OpenAI Decisions API](https://www.anchorterminal.com/tools/openai-decisions-api.md) | BB | 71.5 | High-volume yes or no checks, routing among fixed options and rubric scoring over text and images, for teams already on an OpenAI key. | Pay per use | hosted | | 2 | [Decider](https://www.anchorterminal.com/tools/decider.md) | B | 69.5 | Local classification, routing, triage and checks where a team wants open weights in several sizes and a Jev-shaped route. | Free · OSS | local | | 3 | [Laya](https://www.anchorterminal.com/tools/convai-laya.md) | B | 69.2 | Fast, cheap classification and routing on short text in many languages, as a base to fine-tune on your own labels, and as an MCP or LangGraph routing step. | Free · OSS | local | | 4 | [Kev](https://www.anchorterminal.com/tools/jaredpalmer-kev.md) | B | 67.4 | Self-hosted classification, routing, triage and rubric scoring where a probability matters, especially for teams already calling Jev who want the same API on their own hardware. | Free · OSS | local | | 5 | [Vela 2.0](https://www.anchorterminal.com/tools/vela.md) | B | 66.5 | Self-hosted routing and guardrail checks in one call, where span offsets for personal data or unsupported claims matter. | Free · OSS | local | | 6 | [Clef](https://www.anchorterminal.com/tools/cloudflare-clef.md) | B | 66.1 | Classification, routing and rubric scoring over text, JSON and images inside Cloudflare, or self-hosted where data can't leave. | Freemium | hosted | | 7 | [pplx-decider](https://www.anchorterminal.com/tools/pplx-decider.md) | B | 62.9 | High-volume classification, routing and rubric scoring over text and images where a team wants a low hosted price and the option to run the same weights itself. | Pay per use | hosted | | 8 | [Jev](https://www.anchorterminal.com/tools/typesafe-jev.md) | B | 62.1 | High-volume yes or no answers, labelling, routing and rubric scoring where a probability is more useful than prose, such as ticket triage, invoice checks or picking a tool or skill from a list. | Pay per use | hosted | | 9 | [Strands Decider 2B](https://www.anchorterminal.com/tools/strands-decider.md) | C | 61.3 | Cheap, local classification, routing, triage and tool-call checks on short text inside Strands or other Python agents, and for teams who want to retrain a decision model from a published recipe. | Free · OSS | local | | 10 | [GLiClass](https://www.anchorterminal.com/tools/gliclass.md) | D | 49.9 | Topic, intent and sentiment routing over a known label set on the owner's own CPU or GPU, where many labels must be scored at once. | Free · OSS | local | ## Picks by need - Highest score overall: [OpenAI Decisions API](https://www.anchorterminal.com/tools/openai-decisions-api.md), BB, 71.5/100 on the benchmark. Also [Decider](https://www.anchorterminal.com/tools/decider.md), B, 69.5/100. - Reliability: [Decider](https://www.anchorterminal.com/tools/decider.md), 90/100 on reliability, against 53 for the overall leader. - Schema & documentation: [pplx-decider](https://www.anchorterminal.com/tools/pplx-decider.md), 90/100 on schema & documentation, against 89 for the overall leader. - Maintenance & community: [Decider](https://www.anchorterminal.com/tools/decider.md), 87/100 on maintenance & community, against 77 for the overall leader. - Transparency & trust: [Clef](https://www.anchorterminal.com/tools/cloudflare-clef.md), 86/100 on transparency & trust, against 83 for the overall leader. - Self-hosting under an open licence: [Decider](https://www.anchorterminal.com/tools/decider.md), self-hosted, Apache-2 licence. Also [Laya](https://www.anchorterminal.com/tools/convai-laya.md), self-hosted, Apache-2 licence. ## How to choose - Calibration of stated probabilities: Check calibration on labelled cases, since a high stated probability that often turns out wrong sends an agent down the wrong branch. - Stable answers to option order: Check whether answers stay the same when the order of the options changes, since a routing label that depends on position sends work to the wrong queue. - Accuracy on your own labels: Measure accuracy on labelled decisions drawn from your own workflow, since the answer labels an agent acts on are the ones that count. - Cost per 1,000 decisions: Compare the cost per 1,000 decisions for your own label mix, since output length and retries add to the list price. - How the benchmark tests this category: A fixed set of labelled decisions (yes or no, one label from a set, a level on a rubric) sent to every model in the same request shape, with the same states and questions. We check accuracy against the labels, calibration (expected calibration error, Brier score and how often an answer given 0.9 or more is wrong), whether answers move when the option order changes, latency and the cost per 1,000 decisions. Each listing's verdict, strengths and weaknesses: https://www.anchorterminal.com/best/decision-models/index.md