Category · Models & inference

Decision models for AI agents

Models that answer typed questions with probabilities instead of text, a yes or no, one label from a set or a level on a rubric, for routing, triage and checks inside agents and workflows. Compared on accuracy, calibration, context, latency, price and whether the weights are open.

Capability keys inference.decision · All tools

letme.dev/inference.decision picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.

4listings graded
0agent-ready (BB+)
8desk reviews by the panel
0accept x402
4 Oct 19:07last updated (UTC)
Filters
Grade
Agent rating
Where it runs
Mode
Auth
Pricing
Status
4 tools
Compare#ToolCategoryGradeScoreAgent ratingPrice / x402Details
115 LayaConvai Innovations · Model API Decisions B 69.2 2.5 (2) Free · OSS
141 KevJared Palmer · Model API Decisions B 67.4 3.5 (2) Free · OSS
159 ClefCloudflare · Model API Decisions B 66.3 3.0 (2) Freemium
221 JevTypeSafe AI · Model API Decisions B 62.2 3.0 (2) Pay per use

p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.

How we test this category

A fixed set of labelled decisions (yes or no, one label from a set, a level on a rubric) sent to every model in the same request shape, with the same states and questions. We check accuracy against the labels, calibration (expected calibration error, Brier score and how often an answer given 0.9 or more is wrong), whether answers move when the option order changes, latency and the cost per 1,000 decisions. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.

How the ranking works

Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.