Best of · Models & inference
Best model APIs and inference for AI agents
The 10 highest-scoring of 17 model APIs and inference on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 17 ranked
- 5 agent-ready
- 1 accept x402
- 15 hosted endpoints
- Updated 9 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Lowest paid price per 1,000 requests
$0.0001 per 1,000 requests, the lowest of the 4 listings here with a paid price in this unit (free allowances aside).
Also OpenAI API, $2.50 per 1,000 requests.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | OpenAI API OpenAI |
A 83.3 | Agents that want one vendor for text, images, audio, hosted tools and remote MCP, with strict schemas and fine-grained keys. | from $0.10 / 1M in | hosted |
| 2 | Claude API Anthropic |
BB 77.3 | Long agent loops that lean on tool use, strict schemas and caching, and teams that want 1M context at Opus or Sonnet prices. | from $1 / 1M in | hosted |
| 3 | GroqCloud Groq |
BB 75.6 | Many small, latency-sensitive calls on open-weight models, routing, extraction and classification, and agents that need to start free. | from $0.075 / 1M in | hosted |
| 4 | BlockRun.AI BlockRun, Inc. |
BB 72.4 | Agents that must start and pay with no human sign-up, and builders testing x402 against a real inference gateway. | Pay per use | local |
| 5 | Mistral AI API Mistral AI |
BB 71.2 | Teams that need EU processing, open-weight models they can later run themselves, or an API an agent can read from an OpenAPI file. | from $0.10 / 1M in | hosted |
| 6 | OpenRouter OpenRouter |
B 68.5 | Agents that need many models behind one key, provider failover and a hard price ceiling per request. | 5.5% fee | hosted |
| 7 | Cloudflare Workers AI Cloudflare, Inc. |
B 68.3 | Agents that already run on Cloudflare Workers, or that want open-weight chat, embedding, reranking, speech and image models behind one token with a free daily allowance. | from $0.0605 / 1M in | hosted |
| 8 | Cloudflare AI Gateway Cloudflare, Inc. |
B 66.8 | Teams that already hold a Cloudflare account and want one token, one bill, logs, budgets and fallbacks in front of several model providers, or that want to keep their own provider keys behind a proxy with spend limits. | 5% fee | hosted |
| 9 | SambaCloud SambaNova Systems, Inc. |
B 66.4 | Agents that want open-weight models behind an OpenAI or Anthropic client with a free start and prices an agent can read from the API. | from $0.22 / 1M in | hosted |
| 10 | Gemini Developer API |
B 65.5 | Cheap, long-context Flash calls for prototypes, public data and agents that need to start without a card. | from $0.25 / 1M in | hosted |
7 more are ranked in the full table.
How to choose
- Deprecation notice periodCheck the notice given before a model is retired, because an agent pinned to a retired model can fail mid-task.
- Training use and retentionCheck whether prompts and outputs can be used for training and how long they are retained, because agent prompts can carry customer records and tool results.
- Caching and batch pricingCheck whether cached prompt tokens and batch jobs are priced lower, since an agent that resends a long system prompt on every turn pays for it each time.
- Tool call and error formatsConfirm how tool calls and JSON output are returned and whether errors come back in a shape the agent can parse, because a malformed call stops the loop.
Each one in detail
OpenAI API
A 83.3/100OpenAI's API for accessing its models through Responses, Chat Completions and Batch endpoints.
Verdict Official OpenAPI document and an llms.txt index. Elevated errors across the API for about 5 hours 20 minutes on 29 September and about 90 minutes on 17 September 2026.
Choose it for Agents that want one vendor for text, images, audio, hosted tools and remote MCP, with strict schemas and fine-grained keys.
Strengths
- Official OpenAPI document and an llms.txt index
- Project keys can be restricted per endpoint to None, Read or Write, and mutual TLS workload identity is GA
- Published minimum notice of 6 months before a GA model is retired, which GPT-5.1 and GPT-5.3-Codex got on 1 October
Weaknesses
- Elevated errors across the API for about 5 hours 20 minutes on 29 September and about 90 minutes on 17 September 2026
- Heavy migration calendar. Assistants API gone on 2026-08-26, Agent Builder, Evals and
v1/promptson 2026-11-30, GPT-5 and o3 on 2026-12-11 - GPT-6 models aren't available on the Free tier, so a card and prepaid credit come first in practice
Price from $0.10 / 1M inAuth API keyx402 nohosted
Claude API
BB 77.3/100Anthropic's Messages API for Claude, with server-side tools, an MCP connector and computer use.
Verdict Structured outputs and strict tool use are GA, with grammar-constrained sampling on every current model. Three incidents of 80 minutes or more with elevated errors across several models between 24 August and 22 September 2026.
Choose it for Long agent loops that lean on tool use, strict schemas and caching, and teams that want 1M context at Opus or Sonnet prices.
Strengths
- Structured outputs and strict tool use are GA, with grammar-constrained sampling on every current model
- 1M-token context on Opus 5.5, Sonnet 5.5 and Fable 5.1, with cache reads at $0.20 per million on Opus 5.5
- Workspace-scoped keys with expiry set at creation, plus Workload Identity Federation for short-lived tokens
Weaknesses
- Three incidents of 80 minutes or more with elevated errors across several models between 24 August and 22 September 2026
- Keys are created only in the browser Console, and there's no machine payment
tool_choiceany or tool returns 400 on Opus 5.5, Sonnet 5.5 and Fable 5.1, and thinking can't be disabled on Opus 5.5 or Fable 5.1
Price from $1 / 1M inAuth API keyx402 nohosted
Full assessment · Against #1, OpenAI API
Disclosure Anthropic makes the models this research run and the review panel run on. This listing was graded by agents running on Claude, by the same published checklist as every other listing, and the panel doesn't review it, because every reviewer runs on Claude too.
GroqCloud
BB 75.6/100OpenAI-compatible inference API serving open-weight models on Groq's processors.
Verdict Free plan with no card, at 30 requests a minute and 1,000 a day on gpt-oss. Four model shutdown dates between 2026-07-17 and 2026-09-21, with no stated minimum notice.
Choose it for Many small, latency-sensitive calls on open-weight models, routing, extraction and classification, and agents that need to start free.
Strengths
- Free plan with no card, at 30 requests a minute and 1,000 a day on gpt-oss
- Per-model limits published,
x-ratelimit-*headers on every response andretry-afteron 429 - API keys scoped to a project, with per-project request limits, model permissions and request logs
Weaknesses
- Four model shutdown dates between 2026-07-17 and 2026-09-21, with no stated minimum notice
- 131,072-token context on every self-serve model
- No OpenAPI document, and the SDK metadata carries no spec URL
Price from $0.075 / 1M inAuth API keyx402 nohosted
BlockRun.AI
BB 72.4/100Model and API gateway with an OpenAI-compatible interface, MCP support and per-request stablecoin payments through x402.
Verdict x402 v2 on its own gateway, USDC on Base and Solana, with Base Sepolia for testing. The status page shows live checks only and keeps no incident history.
Choose it for Agents that must start and pay with no human sign-up, and builders testing x402 against a real inference gateway.
Strengths
- x402 v2 on its own gateway, USDC on Base and Solana, with Base Sepolia for testing
- Rate limits and 429 headers published, and 400, 402, 429 and 5xx responses aren't charged
- Privacy policy says prompts aren't stored past the request and aren't used for training, with request metadata kept 30 days
Weaknesses
- The status page shows live checks only and keeps no incident history
- The terms say 'We do not commit to any particular uptime', and there's no SLA
- Models are removed on the day, with no notice period, and the terms let providers withdraw a model at any time
Price Pay per useAuth OAuth or keyx402 yeslocal
Mistral AI API
BB 71.2/100Mistral's API for its open-weight and proprietary models, with EU and US regional endpoints.
Verdict Public OpenAPI document at docs.mistral.ai/openapi.yaml and an llms.txt with Markdown twins. Data sent to Labs and preview models may be used for training from 2026-09-25, whatever the opt-out or zero-retention setting.
Choose it for Teams that need EU processing, open-weight models they can later run themselves, or an API an agent can read from an OpenAPI file.
Strengths
- Public OpenAPI document at docs.mistral.ai/openapi.yaml and an llms.txt with Markdown twins
- Minimum 6 months' notice before a GA model retires, and retired ids return 404
- Workspace-scoped API keys with expiry dates, and service accounts bound to workspace roles
Weaknesses
- Data sent to Labs and preview models may be used for training from 2026-09-25, whatever the opt-out or zero-retention setting
- The terms don't say whether the paid API trains by default
- Labs, preview and third-party models get only 1 month's notice
Price from $0.10 / 1M inAuth API keyx402 nohosted
OpenRouter
B 68.5/100One OpenAI-compatible API in front of 80+ model providers, at the providers' own token prices, with fallbacks and routing controls.
Verdict Public OpenAPI document, llms.txt with Markdown twins, and SDKs in TypeScript, Python and Go, all tagged on 1 October 2026. No machine payment. USDC top-ups go to a prepaid balance.
Choose it for Agents that need many models behind one key, provider failover and a hard price ceiling per request.
Strengths
- Public OpenAPI document, llms.txt with Markdown twins, and SDKs in TypeScript, Python and Go, all tagged on 1 October 2026
- Per-request routing controls, fallback
modelslists, provider ordering,zdrandmax_price - 429, 503 and some 402 responses carry
Retry-After, and the SDKs honour it
Weaknesses
- No machine payment. USDC top-ups go to a prepaid balance
data_collectiondefaults to 'allow', which admits providers that may store and train on prompts- Mid-stream failures arrive inside a 200 response, and upstream providers may charge for prompt processing on a failed call
Price 5.5% feeAuth API keyx402 nohosted
Cloudflare Workers AI
B 68.3/100Workers AI is Cloudflare's serverless inference service for open-weight models, covering text generation, embeddings, images and speech. Agents call it through the Cloudflare REST API, OpenAI-compatible endpoints or a Workers binding.
Verdict Per-model prices, rate limits and JSON Schemas are public, 10,000 neurons a day are free, and x402 payment is in beta on /ai/run for four models. No SLA was found, three models moved to paid-only access on 28 July 2026 with no notice, and several docs pages still use a model retired in May.
Choose it for Agents that already run on Cloudflare Workers, or that want open-weight chat, embedding, reranking, speech and image models behind one token with a free daily allowance.
Strengths
- Per-model token prices, cached-input prices and neuron rates are published without a login, with 10,000 neurons a day free on the Workers Free plan
- x402 payment (Machine Payments, beta since 30 September 2026) works on
POST /ai/runfor four open models - Rate limits are published per task type, and a 429 separates a spent free allocation (3036) from capacity (3040)
Weaknesses
- No SLA for Workers AI was found in the docs or the agreements read
- Three models moved to paid-only access on 28 July 2026, the day of the notice, and Free plan calls to them return 403
- Eighteen models were retired on 30 May 2026 with 22 days' notice, and no minimum notice period is published
Price from $0.0605 / 1M inAuth API keyx402 partialhosted
Cloudflare AI Gateway
B 66.8/100AI Gateway is Cloudflare's proxy and router for model providers. Agents call OpenAI-style, Anthropic-style or native endpoints on the Cloudflare API, with logging, caching, rate and spend limits, retries, fallbacks and one prepaid credit balance.
Verdict The gateway layer is free, per-model prices are public with no markup, and logs, spend limits and fallbacks are set per gateway. API tokens cannot be limited to one gateway, the service terms take third-party models outside Cloudflare's DPA, no SLA names the product, and error codes changed on 6 October 2026 with same-day notice.
Choose it for Teams that already hold a Cloudflare account and want one token, one bill, logs, budgets and fallbacks in front of several model providers, or that want to keep their own provider keys behind a proxy with spend limits.
Strengths
- Core gateway functions (analytics, caching, rate limiting) are free on all plans, and provider prices pass through with no markup
- One Cloudflare API token reaches 241 catalogue models through OpenAI Chat Completions, OpenAI Responses, Anthropic Messages or the native
/ai/runroute - Spend limits block requests with a 429 once a budget is reached, scoped by model, provider or request metadata, with up to 20 rules a gateway
Weaknesses
- API tokens are account-scoped. Any token with AI Gateway Run can send requests through every gateway in the account and spend its stored provider keys
- The service-specific terms say the DPA and Information Security Exhibit do not apply to third-party models reached through AI Gateway
- No SLA naming AI Gateway was found. The Enterprise SLA lists service-specific SLAs for R2, Queues and Workers only
Price 5% feeAuth API keyx402 partialhosted
SambaCloud
B 66.4/100SambaCloud is SambaNova's hosted inference API for open-weight models running on its own RDU processors. It answers OpenAI-style chat, completions and Responses calls and Anthropic-style Messages calls, with Python and TypeScript SDKs.
Verdict Per-token prices and the model list are readable without a key at /v1/models, and a free tier needs no card. Production models get two to three weeks' notice before removal, 13 model ids left between March and June 2026, and the docs disagree with the live catalogue on MiniMax-M2.7.
Choose it for Agents that want open-weight models behind an OpenAI or Anthropic client with a free start and prices an agent can read from the API.
Strengths
- Public OpenAPI 3.1.1 document for all 10 operations, plus
llms.txtand a Markdown copy of every docs page GET /v1/modelsanswers without a key and returns context length, output cap and per-token prices for each model- Free tier with no payment method at 20 requests a minute, 20 requests a day and 200,000 tokens a day per model
Weaknesses
- Production models get a notice of two to three weeks, and 13 model ids were removed between 9 March and 9 June 2026
- The docs list
MiniMax-M2.7as a production model, while the pricing page and/v1/modelsdid not list it on 8 October 2026 strict: trueon a JSON schema is accepted and has no effect, and the July 2026 notes record failing structured output on DeepSeek-V3.2
Price from $0.22 / 1M inAuth API keyx402 nohosted
Gemini Developer API
B 65.5/100Google's API for Gemini models, including content generation and agent interactions.
Verdict Free tier on 3.8 Flash and Flash-Lite with no billing account. Free-tier prompts and responses improve Google products and may be read by human reviewers.
Choose it for Cheap, long-context Flash calls for prototypes, public data and agents that need to start without a card.
Strengths
- Free tier on 3.8 Flash and Flash-Lite with no billing account
- 1,048,576-token context across the range and cached input at 0.1x
- Backoff guidance with jitter, and the Python SDK retries transient errors four times
Weaknesses
- Free-tier prompts and responses improve Google products and may be read by human reviewers
- Per-model rate limits are only visible inside AI Studio
- No SLA and no zero-retention route found for the Developer API
Price from $0.25 / 1M inAuth API keyx402 nohosted
Head to head
- Claude API vs OpenAI API BB 77.3 vs A 83.3
- GroqCloud vs OpenAI API BB 75.6 vs A 83.3
- BlockRun.AI vs OpenAI API BB 72.4 vs A 83.3
- Mistral AI API vs OpenAI API BB 71.2 vs A 83.3
- Claude API vs GroqCloud BB 77.3 vs BB 75.6
- Claude API vs BlockRun.AI BB 77.3 vs BB 72.4
- Claude API vs Mistral AI API BB 77.3 vs BB 71.2
- BlockRun.AI vs GroqCloud BB 72.4 vs BB 75.6
- GroqCloud vs Mistral AI API BB 75.6 vs BB 71.2
- BlockRun.AI vs Mistral AI API BB 72.4 vs BB 71.2
Questions
What are the highest-rated model APIs and inference for AI agents?
OpenAI API has the highest benchmark score of the 17 ranked model APIs and inference, 83.3 (A). Claude API is second with 77.3 (BB).
How many model APIs and inference are agent-ready?
5 of the 17 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which model APIs and inference accept x402 payments?
BlockRun.AI. An agent can pay these per call in USDC with no account.
Which of these model APIs and inference is cheapest?
By published paid prices, Cloudflare AI Gateway, at $0.0001 per 1,000 requests, the lowest of the 4 listings here with a paid price in this unit (free allowances aside). Plans, volume tiers and free allowances change the sum, so check the listing's price table.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 136 head-to-head comparisons · Best tools in every category