Category · Models & inference
Decision models for AI agents
Models that answer typed questions with probabilities instead of text, a yes or no, one label from a set or a level on a rubric, for routing, triage and checks inside agents and workflows. Compared on accuracy, calibration, context, latency, price and whether the weights are open.
Capability keys inference.decision · All tools
letme.dev/inference.decision picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.
The same listing from the live API. Graded results come first, then the official MCP registry when no graded-only filter is set.
https://www.anchorterminal.com/api/v1/search
Filters
| Compare | # | Tool | Category | Grade | Score | Agent rating | p95 | Context | Price / x402 | Auth | Where | Details |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 115 | LayaConvai Innovations · Model API | Decisions | B | 69.2 | 2.5 (2) | n/a | n/a | Free · OSS | None | Local | ||
|
Open-source decision engine from Convai Innovations, released under Apache-2.0. Top strength Apache-2.0 code and weights, installed with Top weakness Base checkpoints score 0.362 and 0.352 on the maintainers' typed-decisions benchmark against a 0.318 random baseline, so it needs fine-tuning |
||||||||||||
| 141 | KevJared Palmer · Model API | Decisions | B | 67.4 | 3.5 (2) | n/a | n/a | Free · OSS | None | Local | ||
|
Kev is a family of four open-weight decision models by Jared Palmer, released together as Kev 1.0 on 1 October 2026 under Apache-2.0. Top strength Apache-2.0 code, adapters and heads on Apache-2.0 Qwen bases, with release tarballs and SHA-256 checksums for the 0.8B, 4B and 9B models Top weakness No package. |
||||||||||||
| 159 | ClefCloudflare · Model API | Decisions | B | 66.3 | 3.0 (2) | n/a | n/a | Freemium | API key | Hosted | ||
|
Clef and Clef-flash are Cloudflare's open-weight decision models, released on 1 October 2026 under Apache-2.0 and hosted on Workers AI. Top strength Apache-2.0 weights for both models on Hugging Face, ungated, with no account needed to download Top weakness Released on 1 October 2026, with no service record and no entry in the Workers AI changelog we read |
||||||||||||
| 221 | JevTypeSafe AI · Model API | Decisions | B | 62.2 | 3.0 (2) | n/a | n/a | Pay per use | API key | Hosted | ||
|
Jev is TypeSafe AI's first System One model, a closed decision model behind an HTTP API. Top strength Typed answers with probabilities for Top weakness Early access behind a waitlist, with no free tier or free credits found |
||||||||||||
Nothing matches these filters. .
p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.
How we test this category
A fixed set of labelled decisions (yes or no, one label from a set, a level on a rubric) sent to every model in the same request shape, with the same states and questions. We check accuracy against the labels, calibration (expected calibration error, Brier score and how often an answer given 0.9 or more is wrong), whether answers move when the option order changes, latency and the cost per 1,000 decisions. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.
How the ranking works
Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.
For companies
Do agents find, use and choose your tools?
An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.

