Best of · Models & inference

Best decision models for AI agents

The 10 highest-scoring of 12 decision models on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.

  • 12 ranked
  • 1 agent-ready
  • 6 hosted endpoints
  • Updated 9 October 2026

Top three

Picks by need

Worked out from the scores, prices and facts, so they change when the research does.

Highest score overall

OpenAI Decisions API BB

BB, 71.5/100 on the benchmark.

Also Decider, B, 69.5/100.

Reliability

Decider B

90/100 on reliability, against 53 for the overall leader.

Schema & documentation

pplx-decider B

90/100 on schema & documentation, against 89 for the overall leader.

Maintenance & community

Decider B

87/100 on maintenance & community, against 77 for the overall leader.

Transparency & trust

Clef B

86/100 on transparency & trust, against 83 for the overall leader.

Self-hosting under an open licence

Decider B

self-hosted, Apache-2 licence.

Also Laya, self-hosted, Apache-2 licence.

The shortlist

#ToolGradeBest forPriceWhere
1 OpenAI Decisions API
OpenAI
BB 71.5 High-volume yes or no checks, routing among fixed options and rubric scoring over text and images, for teams already on an OpenAI key. Pay per use hosted
2 Decider
Mark Marosi (Mapika)
B 69.5 Local classification, routing, triage and checks where a team wants open weights in several sizes and a Jev-shaped route. Free · OSS local
3 Laya
Convai Innovations
B 69.2 Fast, cheap classification and routing on short text in many languages, as a base to fine-tune on your own labels, and as an MCP or LangGraph routing step. Free · OSS local
4 Kev
Jared Palmer
B 67.4 Self-hosted classification, routing, triage and rubric scoring where a probability matters, especially for teams already calling Jev who want the same API on their own hardware. Free · OSS local
5 Vela 2.0
vLLM Semantic Router project and KR Labs
B 66.5 Self-hosted routing and guardrail checks in one call, where span offsets for personal data or unsupported claims matter. Free · OSS local
6 Clef
Cloudflare
B 66.1 Classification, routing and rubric scoring over text, JSON and images inside Cloudflare, or self-hosted where data can't leave. Freemium hosted
7 pplx-decider
Perplexity
B 62.9 High-volume classification, routing and rubric scoring over text and images where a team wants a low hosted price and the option to run the same weights itself. Pay per use hosted
8 Jev
TypeSafe AI
B 62.1 High-volume yes or no answers, labelling, routing and rubric scoring where a probability is more useful than prose, such as ticket triage, invoice checks or picking a tool or skill from a list. Pay per use hosted
9 Strands Decider 2B
Amazon Web Services (Strands Agents)
C 61.3 Cheap, local classification, routing, triage and tool-call checks on short text inside Strands or other Python agents, and for teams who want to retrain a decision model from a published recipe. Free · OSS local
10 GLiClass
Knowledgator
D 49.9 Topic, intent and sentiment routing over a known label set on the owner's own CPU or GPU, where many labels must be scored at once. Free · OSS local

2 more are ranked in the full table.

How to choose

  1. Calibration of stated probabilitiesCheck calibration on labelled cases, since a high stated probability that often turns out wrong sends an agent down the wrong branch.
  2. Stable answers to option orderCheck whether answers stay the same when the order of the options changes, since a routing label that depends on position sends work to the wrong queue.
  3. Accuracy on your own labelsMeasure accuracy on labelled decisions drawn from your own workflow, since the answer labels an agent acts on are the ones that count.
  4. Cost per 1,000 decisionsCompare the cost per 1,000 decisions for your own label mix, since output length and retries add to the list price.

How the benchmark tests this category. A fixed set of labelled decisions (yes or no, one label from a set, a level on a rubric) sent to every model in the same request shape, with the same states and questions. We check accuracy against the labels, calibration (expected calibration error, Brier score and how often an answer given 0.9 or more is wrong), whether answers move when the option order changes, latency and the cost per 1,000 decisions.

Each one in detail

#1

OpenAI Decisions API

BB 71.5/100

The Decisions API is an OpenAI endpoint, in public beta since 6 October 2026, that reads text and images and returns typed answers with probabilities, a predicate, a choice from a fixed set or a rubric score.

Verdict Typed predicate, choice and score answers with probabilities, up to 200 questions a call, at $0.10 per million input tokens with no output charge. It is a public beta released on 6 October 2026, with one model alias, no dated snapshot, no Batch route and no SLA found.

Choose it for High-volume yes or no checks, routing among fixed options and rubric scoring over text and images, for teams already on an OpenAI key.

Strengths

  • Three typed question kinds, predicate, choice and score, with per-option probabilities and up to 200 questions about one input in a call
  • $0.10 per million input tokens on gpt-6-luna, with no output, cache-read or cache-write charge
  • Text and up to 128 inline images in one request

Weaknesses

  • Public beta released on 6 October 2026. The guide expects general availability in the coming weeks and gives no date
  • One model alias, gpt-6-luna, with no dated snapshot to pin
  • Images must be inline base64 data URLs. Hosted image URLs and file_id inputs aren't accepted

Price Pay per useAuth API keyx402 nohosted

Full assessment

#2

Decider

B 69.5/100

Decider is a family of open-weight decision models by Mark Marosi (Mapika), from 0.8B to 35B parameters under Apache-2.0. It answers typed yes or no, choice and score questions with probabilities, and runs locally from the decider-ai Python package.

Verdict An Apache-2.0 decision model family with a dated changelog, passing CI, 21 package releases since 22 September 2026 and model cards that list measured regressions. One person maintains it, the local server has no authentication option, states over 32,768 tokens are cut without an error, and no security policy is published.

Choose it for Local classification, routing, triage and checks where a team wants open weights in several sizes and a Jev-shaped route.

Strengths

  • Apache-2.0 code and weights, with the training code, data builders and per-version measurements in the repository
  • Sizes from 0.8B to 35B parameters, with GGUF files for CPU and builds for CUDA, Apple silicon and vLLM
  • POST /v1/systemone follows TypeSafe's wire format, and the README says TypeSafe's SDKs work with TYPESAFE_BASE_URL set to the local server

Weaknesses

  • The local server has no authentication option. It binds to 127.0.0.1 since 1.7.1
  • States over 32,768 tokens are truncated (DECIDER_MAX_STATE_TOKENS) without an error
  • No SECURITY.md, disclosure policy or security contact was found in the repository

Price Free · OSSAuth Nonex402 nolocal

Full assessment · Against #1, OpenAI Decisions API

#3

Laya

B 69.2/100

Open-source decision engine from Convai Innovations, released under Apache-2.0.

Verdict Apache-2.0 code and weights, installed with pip install laya, with Python 3.10 to 3.13 tested in CI. Base checkpoints score 0.362 and 0.352 on the maintainers' typed-decisions benchmark against a 0.318 random baseline, so it needs fine-tuning.

Choose it for Fast, cheap classification and routing on short text in many languages, as a base to fine-tune on your own labels, and as an MCP or LangGraph routing step.

Strengths

  • Apache-2.0 code and weights, installed with pip install laya, with Python 3.10 to 3.13 tested in CI
  • 421M and 322M-parameter encoders that the README times at 32.8 to 39.5 ms for one question on a Tesla T4
  • A Jev-compatible HTTP server with a batch route for up to 64 states, an 8-tool MCP server, and LangChain, LlamaIndex and CrewAI wrappers

Weaknesses

  • Base checkpoints score 0.362 and 0.352 on the maintainers' typed-decisions benchmark against a 0.318 random baseline, so it needs fine-tuning
  • Choice options share a 192 or 256-token budget, and the README reports 0.425 on Banking77's 77 labels
  • 512 tokens of context on the English checkpoint and 1,024 by default on the others

Price Free · OSSAuth Nonex402 nolocal

Full assessment · Against #1, OpenAI Decisions API

#4

Kev

B 67.4/100

Kev is a family of four open-weight decision models by Jared Palmer, released together as Kev 1.0 on 1 October 2026 under Apache-2.0.

Verdict Apache-2.0 code, adapters and heads on Apache-2.0 Qwen bases, with release tarballs and SHA-256 checksums for the 0.8B, 4B and 9B models. No package. pip install kev installs an unrelated 2021 ORM, so Kev runs from a Git clone with uv.

Choose it for Self-hosted classification, routing, triage and rubric scoring where a probability matters, especially for teams already calling Jev who want the same API on their own hardware.

Strengths

  • Apache-2.0 code, adapters and heads on Apache-2.0 Qwen bases, with release tarballs and SHA-256 checksums for the 0.8B, 4B and 9B models
  • The same /v1/systemone request and answer shapes as Jev, and the README says TypeSafe's Python SDK works against it unchanged
  • A fitted temperature per checkpoint, with Brier scores, calibration error and confident-error rates published for each model

Weaknesses

  • No package. pip install kev installs an unrelated 2021 ORM, so Kev runs from a Git clone with uv
  • Kev-0.8B, 4B and 9B are validated to 8,192 tokens of state, though the server accepts 65,536
  • Jared Palmer wrote 312 of the 333 commits we cloned

Price Free · OSSAuth Nonex402 nolocal

Full assessment · Against #1, OpenAI Decisions API

#5

Vela 2.0

B 66.5/100

Vela 2.0 is a family of four open-weight decision models from the vLLM Semantic Router project and KR Labs, released on 6 October 2026 under Apache-2.0 for routing, safety checks, personal-data spans and hallucination checks.

Verdict One self-hosted call answers routing, prompt-attack, personal-data and unsupported-claim questions with probabilities and character offsets, under Apache-2.0 with SHA-256 manifests. The models are days old and carry no Hub version tags, and the three larger sizes keep 74 to 89 per cent of their Decision 2.0 bases on the Jev Decision Index by the authors' figures.

Choose it for Self-hosted routing and guardrail checks in one call, where span offsets for personal data or unsupported claims matter.

Strengths

  • Five question types in one request (choice, noul, score, set and span), with span answers as labelled character offsets and a probability each
  • Apache-2.0 weights, code and documentation, ungated on Hugging Face, with safetensors files and a SHA256SUMS manifest in the three decoder repositories
  • Two serving routes. A bundled FastAPI server on POST /v1/systemone, and the router's model runtime with an OpenAPI 3.0.3 contract and Prometheus metrics

Weaknesses

  • No version tags on the four Hub repositories, and the 4B and 9B weights were replaced in place on 3 October 2026
  • Loading with transformers needs trust_remote_code=True, which runs Python from the model repository
  • The router's vllm-sr serve MODEL engine mode is newer than the 0.4.0 stable release and needs the development channel

Price Free · OSSAuth Nonex402 nolocal

Full assessment · Against #1, OpenAI Decisions API

#6

Clef

B 66.1/100

Clef and Clef-flash are Cloudflare's open-weight decision models, released on 1 October 2026 under Apache-2.0 and hosted on Workers AI.

Verdict Apache-2.0 weights for both models on Hugging Face, ungated, with no account needed to download. Released on 1 October 2026, with no service record and no entry in the Workers AI changelog we read.

Choose it for Classification, routing and rubric scoring over text, JSON and images inside Cloudflare, or self-hosted where data can't leave.

Strengths

  • Apache-2.0 weights for both models on Hugging Face, ungated, with no account needed to download
  • $0.24 per million input tokens for Clef and $0.09 for Clef-flash, with 10,000 free neurons a day on Workers AI
  • Typed answers with a probability for every allowed option, 1 to 64 questions a call, and up to 4 images

Weaknesses

  • Released on 1 October 2026, with no service record and no entry in the Workers AI changelog we read
  • No limitations, failure modes or prompt-injection guidance in the model page or the launch post
  • No SLA found for Workers AI, and Clef shares the Text Generation default of 300 requests a minute

Price FreemiumAuth API keyx402 nohosted

Full assessment · Against #1, OpenAI Decisions API

#7

pplx-decider

B 62.9/100

pplx-decider is Perplexity's multimodal decision model, a fine-tune of Qwen3.8-27B. It answers yes or no, choice and score questions about text, JSON and images with probabilities, through Perplexity's hosted Decisions API or from open weights on Hugging Face.

Verdict The hosted Decisions API has an OpenAPI description, typed questions, documented limits and a published price of $0.02 per million input tokens, and the v1.1 weights are public under Apache-2.0. The API is days old, the official Python SDK has no method for it, and the governing terms and privacy notice could not be read on 9 October 2026.

Choose it for High-volume classification, routing and rubric scoring over text and images where a team wants a low hosted price and the option to run the same weights itself.

Strengths

  • Weights for v1.1 and v1 are public on Hugging Face under Apache-2.0, with training code, config and a data manifest in the repository
  • One request takes up to 128 questions about one state, mixing noul, choice and score, with under 262,144 input tokens
  • $0.02 per million input tokens with free output and no per-request fee, published without a login

Weaknesses

  • The Decisions API was released in October 2026, so its incident record is days long, and Perplexity's FAQ says it gives no uptime guarantee
  • The official Python SDK at 0.43.8 has no Decisions method. The guide uses a plain HTTP client
  • An image over 2,048 tiles is not rejected. The request waits about a minute and returns 504

Price Pay per useAuth API keyx402 nohosted

Full assessment · Against #1, OpenAI Decisions API

#8

Jev

B 62.1/100

Jev is TypeSafe AI's first System One model, a closed decision model behind an HTTP API.

Verdict Typed answers with probabilities for noul, choice and score questions, many per call, with no text to parse. Early access behind a waitlist, with no free tier or free credits found.

Choose it for High-volume yes or no answers, labelling, routing and rubric scoring where a probability is more useful than prose, such as ticket triage, invoice checks or picking a tool or skill from a list.

Strengths

  • Typed answers with probabilities for noul, choice and score questions, many per call, with no text to parse
  • $0.042 per million input tokens, with output tokens free
  • A public OpenAPI 3.1 document, llms.txt with Markdown twins, and Python and TypeScript SDKs that retry 408, 429 and 5xx with backoff

Weaknesses

  • Early access behind a waitlist, with no free tier or free credits found
  • No SLA, and the customer agreement promises only commercially reasonable efforts to give notice of API changes
  • Published limits of 40 requests a second can change without notice

Price Pay per useAuth API keyx402 nohosted

Full assessment · Against #1, OpenAI Decisions API

#9

Strands Decider 2B

C 61.3/100

Strands Decider 2B is an open-weight decision model from AWS's Strands Labs, released on 1 October 2026 under Apache-2.0. It answers typed yes or no, choice and score questions with probabilities, and runs locally from a Python package.

Verdict A 1.9B-parameter Apache-2.0 decision model that runs on a laptop GPU, an Apple silicon Mac or a CPU, with its training data, recipe and per-version results published. It's an experimental 0.1.0 release with a 4,096-token window that cuts long states by default, and its local server has no authentication.

Choose it for Cheap, local classification, routing, triage and tool-call checks on short text inside Strands or other Python agents, and for teams who want to retrain a decision model from a published recipe.

Strengths

  • Apache-2.0 code and weights, with the training recipe, data inventory and evaluation logs published
  • Runs on CUDA, Apple silicon (MPS or MLX) or CPU, with a v19 median of 115 ms a question on an RTX 3090 per the README
  • Brier score and expected calibration error published for each released checkpoint

Weaknesses

  • Version 0.1.0, described as experimental in its package metadata, with no changelog file
  • A 4,096-token window, and by default an over-long state is cut to fit without an error
  • The local server has no authentication option

Price Free · OSSAuth Nonex402 nolocal

Full assessment · Against #1, OpenAI Decisions API

#10

GLiClass

D 49.9/100

GLiClass is an open-source Python library and family of open-weight zero-shot text classifiers from Knowledgator. It scores every candidate label in one forward pass and runs locally through a pipeline or a Ray Serve endpoint.

Verdict An Apache-2.0 classifier that scores a whole label set in one encoder pass on the owner's hardware, with single-label, multi-label, hierarchical and few-shot modes. It returns label scores with no calibration claim, the bundled server has no authentication, and the last three test runs on the main branch, on 24 September 2026, failed.

Choose it for Topic, intent and sentiment routing over a known label set on the owner's own CPU or GPU, where many labels must be scored at once.

Strengths

  • Apache-2.0 code and weights, with train.py, the training datasets named on the model cards and an arXiv paper (2508.07662)
  • One forward pass scores every label. The base v3.0 card reports 51.6 examples a second averaged over 1 to 128 labels on an A6000 (vendor figures)
  • Single-label (softmax) and multi-label (sigmoid) modes, hierarchical label sets, task prompts and few-shot examples in one pipeline call

Weaknesses

  • No calibration evidence. The cards report F1 only, and the docs tell users to calibrate thresholds on their own traffic
  • The bundled server has no authentication, and python -m gliclass.serve binds to 0.0.0.0 by default
  • The last three runs of the Tests workflow on main, all on 24 September 2026, failed

Price Free · OSSAuth Nonex402 nolocal

Full assessment · Against #1, OpenAI Decisions API

Head to head

All 66 comparisons in this category

Questions

What are the highest-rated decision models for AI agents?

OpenAI Decisions API has the highest benchmark score of the 12 ranked decision models, 71.5 (BB). Decider is second with 69.5 (B).

How many decision models are agent-ready?

1 of the 12 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.

Which decision models accept x402 payments?

None of the ranked listings here accepts x402 for its main call yet.

How is this list ranked?

By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.

How this list is made

The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.

Full ranked table · 66 head-to-head comparisons · Best tools in every category

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.