Best of · Models & inference
Best decision models for AI agents
The 10 highest-scoring of 12 decision models on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 12 ranked
- 1 agent-ready
- 6 hosted endpoints
- Updated 9 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | OpenAI Decisions API OpenAI |
BB 71.5 | High-volume yes or no checks, routing among fixed options and rubric scoring over text and images, for teams already on an OpenAI key. | Pay per use | hosted |
| 2 | Decider Mark Marosi (Mapika) |
B 69.5 | Local classification, routing, triage and checks where a team wants open weights in several sizes and a Jev-shaped route. | Free · OSS | local |
| 3 | Laya Convai Innovations |
B 69.2 | Fast, cheap classification and routing on short text in many languages, as a base to fine-tune on your own labels, and as an MCP or LangGraph routing step. | Free · OSS | local |
| 4 | Kev Jared Palmer |
B 67.4 | Self-hosted classification, routing, triage and rubric scoring where a probability matters, especially for teams already calling Jev who want the same API on their own hardware. | Free · OSS | local |
| 5 | Vela 2.0 vLLM Semantic Router project and KR Labs |
B 66.5 | Self-hosted routing and guardrail checks in one call, where span offsets for personal data or unsupported claims matter. | Free · OSS | local |
| 6 | Clef Cloudflare |
B 66.1 | Classification, routing and rubric scoring over text, JSON and images inside Cloudflare, or self-hosted where data can't leave. | Freemium | hosted |
| 7 | pplx-decider Perplexity |
B 62.9 | High-volume classification, routing and rubric scoring over text and images where a team wants a low hosted price and the option to run the same weights itself. | Pay per use | hosted |
| 8 | Jev TypeSafe AI |
B 62.1 | High-volume yes or no answers, labelling, routing and rubric scoring where a probability is more useful than prose, such as ticket triage, invoice checks or picking a tool or skill from a list. | Pay per use | hosted |
| 9 | Strands Decider 2B Amazon Web Services (Strands Agents) |
C 61.3 | Cheap, local classification, routing, triage and tool-call checks on short text inside Strands or other Python agents, and for teams who want to retrain a decision model from a published recipe. | Free · OSS | local |
| 10 | GLiClass Knowledgator |
D 49.9 | Topic, intent and sentiment routing over a known label set on the owner's own CPU or GPU, where many labels must be scored at once. | Free · OSS | local |
2 more are ranked in the full table.
How to choose
- Calibration of stated probabilitiesCheck calibration on labelled cases, since a high stated probability that often turns out wrong sends an agent down the wrong branch.
- Stable answers to option orderCheck whether answers stay the same when the order of the options changes, since a routing label that depends on position sends work to the wrong queue.
- Accuracy on your own labelsMeasure accuracy on labelled decisions drawn from your own workflow, since the answer labels an agent acts on are the ones that count.
- Cost per 1,000 decisionsCompare the cost per 1,000 decisions for your own label mix, since output length and retries add to the list price.
How the benchmark tests this category. A fixed set of labelled decisions (yes or no, one label from a set, a level on a rubric) sent to every model in the same request shape, with the same states and questions. We check accuracy against the labels, calibration (expected calibration error, Brier score and how often an answer given 0.9 or more is wrong), whether answers move when the option order changes, latency and the cost per 1,000 decisions.
Each one in detail
OpenAI Decisions API
BB 71.5/100The Decisions API is an OpenAI endpoint, in public beta since 6 October 2026, that reads text and images and returns typed answers with probabilities, a predicate, a choice from a fixed set or a rubric score.
Verdict Typed predicate, choice and score answers with probabilities, up to 200 questions a call, at $0.10 per million input tokens with no output charge. It is a public beta released on 6 October 2026, with one model alias, no dated snapshot, no Batch route and no SLA found.
Choose it for High-volume yes or no checks, routing among fixed options and rubric scoring over text and images, for teams already on an OpenAI key.
Strengths
- Three typed question kinds, predicate, choice and score, with per-option probabilities and up to 200 questions about one input in a call
- $0.10 per million input tokens on
gpt-6-luna, with no output, cache-read or cache-write charge - Text and up to 128 inline images in one request
Weaknesses
- Public beta released on 6 October 2026. The guide expects general availability in the coming weeks and gives no date
- One model alias,
gpt-6-luna, with no dated snapshot to pin - Images must be inline base64 data URLs. Hosted image URLs and
file_idinputs aren't accepted
Price Pay per useAuth API keyx402 nohosted
Decider
B 69.5/100Decider is a family of open-weight decision models by Mark Marosi (Mapika), from 0.8B to 35B parameters under Apache-2.0. It answers typed yes or no, choice and score questions with probabilities, and runs locally from the decider-ai Python package.
Verdict An Apache-2.0 decision model family with a dated changelog, passing CI, 21 package releases since 22 September 2026 and model cards that list measured regressions. One person maintains it, the local server has no authentication option, states over 32,768 tokens are cut without an error, and no security policy is published.
Choose it for Local classification, routing, triage and checks where a team wants open weights in several sizes and a Jev-shaped route.
Strengths
- Apache-2.0 code and weights, with the training code, data builders and per-version measurements in the repository
- Sizes from 0.8B to 35B parameters, with GGUF files for CPU and builds for CUDA, Apple silicon and vLLM
POST /v1/systemonefollows TypeSafe's wire format, and the README says TypeSafe's SDKs work withTYPESAFE_BASE_URLset to the local server
Weaknesses
- The local server has no authentication option. It binds to 127.0.0.1 since 1.7.1
- States over 32,768 tokens are truncated (
DECIDER_MAX_STATE_TOKENS) without an error - No SECURITY.md, disclosure policy or security contact was found in the repository
Price Free · OSSAuth Nonex402 nolocal
Laya
B 69.2/100Open-source decision engine from Convai Innovations, released under Apache-2.0.
Verdict Apache-2.0 code and weights, installed with pip install laya, with Python 3.10 to 3.13 tested in CI. Base checkpoints score 0.362 and 0.352 on the maintainers' typed-decisions benchmark against a 0.318 random baseline, so it needs fine-tuning.
Choose it for Fast, cheap classification and routing on short text in many languages, as a base to fine-tune on your own labels, and as an MCP or LangGraph routing step.
Strengths
- Apache-2.0 code and weights, installed with
pip install laya, with Python 3.10 to 3.13 tested in CI - 421M and 322M-parameter encoders that the README times at 32.8 to 39.5 ms for one question on a Tesla T4
- A Jev-compatible HTTP server with a batch route for up to 64 states, an 8-tool MCP server, and LangChain, LlamaIndex and CrewAI wrappers
Weaknesses
- Base checkpoints score 0.362 and 0.352 on the maintainers' typed-decisions benchmark against a 0.318 random baseline, so it needs fine-tuning
- Choice options share a 192 or 256-token budget, and the README reports 0.425 on Banking77's 77 labels
- 512 tokens of context on the English checkpoint and 1,024 by default on the others
Price Free · OSSAuth Nonex402 nolocal
Kev
B 67.4/100Kev is a family of four open-weight decision models by Jared Palmer, released together as Kev 1.0 on 1 October 2026 under Apache-2.0.
Verdict Apache-2.0 code, adapters and heads on Apache-2.0 Qwen bases, with release tarballs and SHA-256 checksums for the 0.8B, 4B and 9B models. No package. pip install kev installs an unrelated 2021 ORM, so Kev runs from a Git clone with uv.
Choose it for Self-hosted classification, routing, triage and rubric scoring where a probability matters, especially for teams already calling Jev who want the same API on their own hardware.
Strengths
- Apache-2.0 code, adapters and heads on Apache-2.0 Qwen bases, with release tarballs and SHA-256 checksums for the 0.8B, 4B and 9B models
- The same
/v1/systemonerequest and answer shapes as Jev, and the README says TypeSafe's Python SDK works against it unchanged - A fitted temperature per checkpoint, with Brier scores, calibration error and confident-error rates published for each model
Weaknesses
- No package.
pip install kevinstalls an unrelated 2021 ORM, so Kev runs from a Git clone with uv - Kev-0.8B, 4B and 9B are validated to 8,192 tokens of state, though the server accepts 65,536
- Jared Palmer wrote 312 of the 333 commits we cloned
Price Free · OSSAuth Nonex402 nolocal
Vela 2.0
B 66.5/100Vela 2.0 is a family of four open-weight decision models from the vLLM Semantic Router project and KR Labs, released on 6 October 2026 under Apache-2.0 for routing, safety checks, personal-data spans and hallucination checks.
Verdict One self-hosted call answers routing, prompt-attack, personal-data and unsupported-claim questions with probabilities and character offsets, under Apache-2.0 with SHA-256 manifests. The models are days old and carry no Hub version tags, and the three larger sizes keep 74 to 89 per cent of their Decision 2.0 bases on the Jev Decision Index by the authors' figures.
Choose it for Self-hosted routing and guardrail checks in one call, where span offsets for personal data or unsupported claims matter.
Strengths
- Five question types in one request (choice, noul, score, set and span), with span answers as labelled character offsets and a probability each
- Apache-2.0 weights, code and documentation, ungated on Hugging Face, with safetensors files and a SHA256SUMS manifest in the three decoder repositories
- Two serving routes. A bundled FastAPI server on
POST /v1/systemone, and the router's model runtime with an OpenAPI 3.0.3 contract and Prometheus metrics
Weaknesses
- No version tags on the four Hub repositories, and the 4B and 9B weights were replaced in place on 3 October 2026
- Loading with
transformersneedstrust_remote_code=True, which runs Python from the model repository - The router's
vllm-sr serve MODELengine mode is newer than the 0.4.0 stable release and needs the development channel
Price Free · OSSAuth Nonex402 nolocal
Clef
B 66.1/100Clef and Clef-flash are Cloudflare's open-weight decision models, released on 1 October 2026 under Apache-2.0 and hosted on Workers AI.
Verdict Apache-2.0 weights for both models on Hugging Face, ungated, with no account needed to download. Released on 1 October 2026, with no service record and no entry in the Workers AI changelog we read.
Choose it for Classification, routing and rubric scoring over text, JSON and images inside Cloudflare, or self-hosted where data can't leave.
Strengths
- Apache-2.0 weights for both models on Hugging Face, ungated, with no account needed to download
- $0.24 per million input tokens for Clef and $0.09 for Clef-flash, with 10,000 free neurons a day on Workers AI
- Typed answers with a probability for every allowed option, 1 to 64 questions a call, and up to 4 images
Weaknesses
- Released on 1 October 2026, with no service record and no entry in the Workers AI changelog we read
- No limitations, failure modes or prompt-injection guidance in the model page or the launch post
- No SLA found for Workers AI, and Clef shares the Text Generation default of 300 requests a minute
Price FreemiumAuth API keyx402 nohosted
pplx-decider
B 62.9/100pplx-decider is Perplexity's multimodal decision model, a fine-tune of Qwen3.8-27B. It answers yes or no, choice and score questions about text, JSON and images with probabilities, through Perplexity's hosted Decisions API or from open weights on Hugging Face.
Verdict The hosted Decisions API has an OpenAPI description, typed questions, documented limits and a published price of $0.02 per million input tokens, and the v1.1 weights are public under Apache-2.0. The API is days old, the official Python SDK has no method for it, and the governing terms and privacy notice could not be read on 9 October 2026.
Choose it for High-volume classification, routing and rubric scoring over text and images where a team wants a low hosted price and the option to run the same weights itself.
Strengths
- Weights for v1.1 and v1 are public on Hugging Face under Apache-2.0, with training code, config and a data manifest in the repository
- One request takes up to 128 questions about one
state, mixingnoul,choiceandscore, with under 262,144 input tokens - $0.02 per million input tokens with free output and no per-request fee, published without a login
Weaknesses
- The Decisions API was released in October 2026, so its incident record is days long, and Perplexity's FAQ says it gives no uptime guarantee
- The official Python SDK at 0.43.8 has no Decisions method. The guide uses a plain HTTP client
- An image over 2,048 tiles is not rejected. The request waits about a minute and returns 504
Price Pay per useAuth API keyx402 nohosted
Jev
B 62.1/100Jev is TypeSafe AI's first System One model, a closed decision model behind an HTTP API.
Verdict Typed answers with probabilities for noul, choice and score questions, many per call, with no text to parse. Early access behind a waitlist, with no free tier or free credits found.
Choose it for High-volume yes or no answers, labelling, routing and rubric scoring where a probability is more useful than prose, such as ticket triage, invoice checks or picking a tool or skill from a list.
Strengths
- Typed answers with probabilities for
noul,choiceandscorequestions, many per call, with no text to parse - $0.042 per million input tokens, with output tokens free
- A public OpenAPI 3.1 document, llms.txt with Markdown twins, and Python and TypeScript SDKs that retry 408, 429 and 5xx with backoff
Weaknesses
- Early access behind a waitlist, with no free tier or free credits found
- No SLA, and the customer agreement promises only commercially reasonable efforts to give notice of API changes
- Published limits of 40 requests a second can change without notice
Price Pay per useAuth API keyx402 nohosted
Strands Decider 2B
C 61.3/100Strands Decider 2B is an open-weight decision model from AWS's Strands Labs, released on 1 October 2026 under Apache-2.0. It answers typed yes or no, choice and score questions with probabilities, and runs locally from a Python package.
Verdict A 1.9B-parameter Apache-2.0 decision model that runs on a laptop GPU, an Apple silicon Mac or a CPU, with its training data, recipe and per-version results published. It's an experimental 0.1.0 release with a 4,096-token window that cuts long states by default, and its local server has no authentication.
Choose it for Cheap, local classification, routing, triage and tool-call checks on short text inside Strands or other Python agents, and for teams who want to retrain a decision model from a published recipe.
Strengths
- Apache-2.0 code and weights, with the training recipe, data inventory and evaluation logs published
- Runs on CUDA, Apple silicon (MPS or MLX) or CPU, with a v19 median of 115 ms a question on an RTX 3090 per the README
- Brier score and expected calibration error published for each released checkpoint
Weaknesses
- Version 0.1.0, described as experimental in its package metadata, with no changelog file
- A 4,096-token window, and by default an over-long state is cut to fit without an error
- The local server has no authentication option
Price Free · OSSAuth Nonex402 nolocal
GLiClass
D 49.9/100GLiClass is an open-source Python library and family of open-weight zero-shot text classifiers from Knowledgator. It scores every candidate label in one forward pass and runs locally through a pipeline or a Ray Serve endpoint.
Verdict An Apache-2.0 classifier that scores a whole label set in one encoder pass on the owner's hardware, with single-label, multi-label, hierarchical and few-shot modes. It returns label scores with no calibration claim, the bundled server has no authentication, and the last three test runs on the main branch, on 24 September 2026, failed.
Choose it for Topic, intent and sentiment routing over a known label set on the owner's own CPU or GPU, where many labels must be scored at once.
Strengths
- Apache-2.0 code and weights, with
train.py, the training datasets named on the model cards and an arXiv paper (2508.07662) - One forward pass scores every label. The base v3.0 card reports 51.6 examples a second averaged over 1 to 128 labels on an A6000 (vendor figures)
- Single-label (softmax) and multi-label (sigmoid) modes, hierarchical label sets, task prompts and few-shot examples in one pipeline call
Weaknesses
- No calibration evidence. The cards report F1 only, and the docs tell users to calibrate thresholds on their own traffic
- The bundled server has no authentication, and
python -m gliclass.servebinds to 0.0.0.0 by default - The last three runs of the Tests workflow on main, all on 24 September 2026, failed
Price Free · OSSAuth Nonex402 nolocal
Head to head
- Decider vs OpenAI Decisions API B 69.5 vs BB 71.5
- Laya vs OpenAI Decisions API B 69.2 vs BB 71.5
- Kev vs OpenAI Decisions API B 67.4 vs BB 71.5
- OpenAI Decisions API vs Vela 2.0 BB 71.5 vs B 66.5
- Laya vs Decider B 69.2 vs B 69.5
- Decider vs Kev B 69.5 vs B 67.4
- Decider vs Vela 2.0 B 69.5 vs B 66.5
- Laya vs Kev B 69.2 vs B 67.4
- Laya vs Vela 2.0 B 69.2 vs B 66.5
- Kev vs Vela 2.0 B 67.4 vs B 66.5
Questions
What are the highest-rated decision models for AI agents?
OpenAI Decisions API has the highest benchmark score of the 12 ranked decision models, 71.5 (BB). Decider is second with 69.5 (B).
How many decision models are agent-ready?
1 of the 12 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which decision models accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 66 head-to-head comparisons · Best tools in every category