Best of · Voice & speech

Best conversational voice-agent APIs

The 10 highest-scoring of 14 conversational voice-agent APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.

  • 14 ranked
  • 1 agent-ready
  • 14 hosted endpoints
  • Updated 9 October 2026

Top three

Picks by need

Worked out from the scores, prices and facts, so they change when the research does.

Highest score overall

ElevenLabs Agents API + MCP BB

BB, 71.3/100 on the benchmark.

Also Retell AI API + MCP, B, 69.1/100.

Reliability

Vapi API + MCP B

70/100 on reliability, against 60 for the overall leader.

Agent ergonomics

Vocily AI D

84/100 on agent ergonomics, against 77 for the overall leader.

Maintenance & community

Vapi API + MCP B

88/100 on maintenance & community, against 85 for the overall leader.

Self-hosting under an open licence

Bolna API + MCP D

self-hosted, MIT licence.

The shortlist

#ToolGradeBest forPriceWhere
1 ElevenLabs Agents API + MCP
ElevenLabs
BB 71.3 Teams that want a managed pipeline with ElevenLabs voices, a choice of LLM, wide telephony integration and web and mobile SDKs. $22 / mo hosted
2 Retell AI API + MCP
Retell AI
B 69.1 Teams that want a self-serve phone agent with clear per-component pricing, scoped keys and a choice of LLM and voice from Retell's menu. $2 / mo hosted
3 Deepgram Voice Agent API
Deepgram
B 68.1 Developers who already run telephony (Twilio, Amazon Connect, Genesys, AudioCodes) and want a cheap managed pipeline with their choice of LLM. Pay per use hosted and local
4 AssemblyAI Voice Agent API
AssemblyAI, Inc.
B 64.4 A developer who wants one managed pipeline at a flat rate and is content with AssemblyAI's own speech models and voices, with an optional model of their own. Pay per use hosted
5 Bland AI API + MCP
Bland AI
B 63.8 Teams that want one predictable per-minute bill for outbound and inbound phone agents, and personal agents that need their own US number. $299 / mo hosted and local
6 Vapi API + MCP
Vapi
B 63.5 Developers who want to choose every provider, bring their own keys and tune cost against latency. $29 / mo hosted and local
7 OpenAI Realtime API
OpenAI
B 63.3 Teams building their own voice agent on one speech-to-speech model who want browser, server and phone transports from one vendor and can run a backend for secrets and tools. Pay per use hosted
8 Ultravox Realtime API
Ultravox (Fixie.ai)
C 58.5 Cost-sensitive voice agents that bring their own telephony and want a speech-native model. $100 / mo hosted
9 Gemini Live API
Google
C 57.1 Teams building their own voice or vision assistant on a single speech-to-speech model with function calling and Search grounding, who can run a backend for tokens and audio transport. $14 / 1k req hosted
10 Hume EVI (Empathic Voice Interface)
Hume AI
C 56.7 Consumer and coaching products where the caller's tone matters and English is enough on EVI 3. $14 / mo hosted

4 more are ranked in the full table.

How to choose

  1. Speech-to-speech or pipelineCheck whether the product runs one speech-to-speech model or a separate recognition, language and synthesis chain, since each changes latency and what you can swap.
  2. Handling of interruptionsAsk how the platform stops speaking when a caller interrupts, because a slow stop leaves the agent talking over the customer.
  3. Tool calls during a live callCheck how a tool call is confirmed to the caller and what happens when it fails, since a long pause or a wrong booking is what the caller hears.
  4. Cost per completed callWork out the cost of a finished call from per-minute, language model and telephony charges together, because a low per-minute rate can hide a costly language model.

How the benchmark tests this category. The same booking and support tasks on every platform. We measure response latency, how interruptions are handled, tool-call success, recovery after a misunderstanding and the total cost per completed call, and record whether the product runs a direct speech-to-speech model or an STT, model and TTS pipeline, since those are different setups.

Each one in detail

#1

ElevenLabs Agents API + MCP

BB 71.3/100

ElevenAgents (formerly Conversational AI) runs hosted voice agents as a pipeline of a fine-tuned ElevenLabs ASR model, an LLM of your choice or your own, ElevenLabs TTS and a proprietary turn-taking model.

Verdict API keys scoped to endpoint groups, with per-key credit limits and service accounts. $0.08 a minute excludes LLM tokens and carrier minutes.

Choose it for Teams that want a managed pipeline with ElevenLabs voices, a choice of LLM, wide telephony integration and web and mobile SDKs.

Strengths

  • API keys scoped to endpoint groups, with per-key credit limits and service accounts
  • Hosted MCP server that signs in with OAuth, with EU, India and Singapore endpoints
  • Public OpenAPI, llms.txt and an errors page with codes, fixes and request IDs

Weaknesses

  • $0.08 a minute excludes LLM tokens and carrier minutes
  • At least six incidents since July where agent calls failed or didn't start
  • Conversation data kept 2 years by default

Price $22 / moAuth OAuth or keyx402 nohosted

Full assessment

#2

Retell AI API + MCP

B 69.1/100

Hosted platform for phone and web voice agents, built as single-prompt agents or node-based conversation flows.

Verdict Published per-minute price for every component, from $0.07 a minute all in. Call data kept indefinitely unless retention is set per agent.

Choose it for Teams that want a self-serve phone agent with clear per-component pricing, scoped keys and a choice of LLM and voice from Retell's menu.

Strengths

  • Published per-minute price for every component, from $0.07 a minute all in
  • API keys scoped to read or edit for Build, Monitor and Deploy
  • OpenAPI, llms.txt and a deprecation feed with RSS

Weaknesses

  • Call data kept indefinitely unless retention is set per agent
  • MCP tools carry no read-only or destructive annotations
  • Web and phone calls disrupted for 69 minutes on 5 September 2026

Price $2 / moAuth API keyx402 nohosted

Full assessment · Against #1, ElevenLabs Agents API + MCP

#3

Deepgram Voice Agent API

B 68.1/100

One WebSocket that runs Deepgram STT (Flux or Nova-3), a managed or bring-your-own LLM and Deepgram or third-party TTS, with turn-taking, barge-in and function calling.

Verdict $0.075 a minute all in on Standard, $0.050 with your own LLM and TTS, $200 of credit with no card. No phone numbers or SIP, so you bridge Twilio or another carrier yourself.

Choose it for Developers who already run telephony (Twilio, Amazon Connect, Genesys, AudioCodes) and want a cheap managed pipeline with their choice of LLM.

Strengths

  • $0.075 a minute all in on Standard, $0.050 with your own LLM and TTS, $200 of credit with no card
  • AsyncAPI description of the agent socket plus a public OpenAPI file
  • Role-based API keys with expiry and 30-second browser tokens

Weaknesses

  • No phone numbers or SIP, so you bridge Twilio or another carrier yourself
  • Fifteen status incidents since July, four of them an hour or longer on parts the agent uses
  • Call audio is kept for model improvement unless mip_opt_out is set

Price Pay per useAuth API keyx402 nohosted and local

Full assessment · Against #1, ElevenLabs Agents API + MCP

#4

AssemblyAI Voice Agent API

B 64.4/100

AssemblyAI's Voice Agent API runs a spoken conversation over one WebSocket, combining its own speech-to-text, language model and text-to-speech. Agents are stored through a REST API and reached from a browser, a server or an inbound Twilio SIP number.

Verdict One $0.075 a minute rate covers speech-to-text, the managed model, speech output, recordings and hosting, and the socket has an AsyncAPI file, typed error codes and a 30-second resume window. No concurrent-session limit is published, the status page has no Voice Agent component, and phone calls are inbound only through a Twilio trunk the customer owns.

Choose it for A developer who wants one managed pipeline at a flat rate and is content with AssemblyAI's own speech models and voices, with an optional model of their own.

Strengths

  • One published rate of $4.50 an hour ($0.075 a minute), billed per second, with $50 of credit and no card for new accounts
  • AsyncAPI 3.0 file for the socket, an OpenAPI file for the token endpoint, llms.txt and a Markdown copy of every docs page
  • Session errors carry a code, a message and sometimes the offending field, and the docs name the three codes that are safe to retry

Weaknesses

  • No number is published for the concurrent-session limit that the concurrency_exceeded error enforces
  • status.assemblyai.com lists the Asynchronous API, Streaming API and LLM Gateway, with no Voice Agent component
  • Telephony is inbound only, over a SIP trunk on a Twilio account the customer owns

Price Pay per useAuth API keyx402 nohosted

Full assessment · Against #1, ElevenLabs Agents API + MCP

#5

Bland AI API + MCP

B 63.8/100

Phone-first voice-agent platform that runs its own speech recognition, language model, TTS and telephony, billed as one per-minute rate.

Verdict One per-minute rate covering STT, LLM, TTS and telephony, $0.14 on Start and $0.12 on Build. No OpenAPI document.

Choose it for Teams that want one predictable per-minute bill for outbound and inbound phone agents, and personal agents that need their own US number.

Strengths

  • One per-minute rate covering STT, LLM, TTS and telephony, $0.14 on Start and $0.12 on Build
  • Start plan with 2 credits and an inbound number, no card
  • Hosted MCP server with 42 tools labelled read, write or destructive, with confirmation on destructive ones

Weaknesses

  • No OpenAPI document
  • Start caps at 10 concurrent calls and 100 calls a day
  • Failed calls and outbound attempts cost $0.015 each

Price $299 / moAuth API keyx402 nohosted and local

Full assessment · Against #1, ElevenLabs Agents API + MCP

#6

Vapi API + MCP

B 63.5/100

Developer platform for phone and web voice agents.

Verdict Voice agents can combine supported transcription, model and voice providers, including customer-supplied endpoints. Per-minute cost depends on the selected providers.

Choose it for Developers who want to choose every provider, bring their own keys and tune cost against latency.

Strengths

  • Any mix of transcriber, model and voice provider, or your own keys and endpoints
  • OpenAPI file, llms.txt and a weekly changelog
  • Public keys limited to allowed origins and assistants

Weaknesses

  • Real per-minute cost depends on provider choices
  • Usage only has 4 concurrent lines and 14 days of retention
  • MCP tools that place calls or buy numbers carry no annotations

Price $29 / moAuth API keyx402 nohosted and local

Full assessment · Against #1, ElevenLabs Agents API + MCP

#7

OpenAI Realtime API

B 63.3/100

OpenAI's Realtime API runs spoken conversations with speech-to-speech models such as gpt-realtime-2.1. Clients connect over WebRTC, WebSocket or SIP, and sessions support interruptions, function calling and remote MCP tools.

Verdict A generally available speech-to-speech API with WebRTC, WebSocket and SIP transports, a public OpenAPI file that types its events, and per-token prices. Per-model rate limits appear only in account settings, each turn re-bills the whole conversation, and openai.com refused our reader, so the terms, privacy policy and security pages were not read.

Choose it for Teams building their own voice agent on one speech-to-speech model who want browser, server and phone transports from one vendor and can run a backend for secrets and tools.

Strengths

  • The OpenAPI 3.1 file in openai/openai-openapi covers nine Realtime paths and types 13 client events and 48 server events as schemas.
  • Browser clients use client secrets that expire after 10 seconds to 2 hours, 10 minutes by default, so the API key stays on a server.
  • Remote MCP tools accept an allowed_tools list, and require_approval defaults to always in the OpenAPI file.

Weaknesses

  • openai.com answered 403 to our reader, so the service terms, privacy policy, sub-processor list and security pages were not read.
  • Realtime rate limits are shown only in account settings. The public table gives numbers for GPT-6 model families only.
  • Every response re-sends the whole conversation, so later turns cost more. Audio is $32 in and $64 out per 1M tokens on gpt-realtime-2.1.

Price Pay per useAuth API keyx402 nohosted

Full assessment · Against #1, ElevenLabs Agents API + MCP

#8

Ultravox Realtime API

C 58.5/100

Hosted voice agents on Ultravox, an open-weight model that takes speech directly with no speech-to-text step and answers through a TTS voice.

Verdict Flat $0.05 a minute including model and built-in voices. No MCP server.

Choose it for Cost-sensitive voice agents that bring their own telephony and want a speech-native model.

Strengths

  • Flat $0.05 a minute including model and built-in voices
  • Speech-native model with no separate STT stage, v0.7 weights public under MIT
  • 429 and 503 responses with Retry-After and backoff guidance

Weaknesses

  • No MCP server
  • News page, deprecation guide and Python client last updated in December 2025
  • Plain API keys with no scopes

Price $100 / moAuth API keyx402 nohosted

Full assessment · Against #1, ElevenLabs Agents API + MCP

#9

Gemini Live API

C 57.1/100

Google's Live API runs real-time spoken conversations with Gemini audio-to-audio models over a stateful WebSocket. It takes audio, images and text, speaks back, and supports interruptions, function calling and Google Search grounding.

Verdict A speech-to-speech API with per-token prices published, a free tier and a generally available model, gemini-3.8-live, since 15 September 2026. The documented raw WebSocket connection carries the API key in the URL, connections reset about every 10 minutes, and no SLA or readable incident history was found for the Developer API.

Choose it for Teams building their own voice or vision assistant on a single speech-to-speech model with function calling and Search grounding, who can run a backend for tokens and audio transport.

Strengths

  • gemini-3.8-live went generally available on 15 September 2026 with asynchronous function calling as the default and three response scheduling modes.
  • Ephemeral tokens can be single use, expire in 30 minutes by default and be locked to a model and session configuration.
  • Prices are public per million tokens with per-minute equivalents, $0.005 a minute of audio in and $0.018 a minute of audio out.

Weaknesses

  • The WebSocket guide authenticates with the API key as a key query parameter, and ephemeral tokens as an access_token query parameter.
  • The status page at aistudio.google.com/status renders in the browser only, so no incident history could be read, and no SLA was found for the Developer API.
  • No machine-readable contract for the WebSocket messages was found. The Discovery document types only the setup and token schemas.

Price $14 / 1k reqAuth API keyx402 nohosted

Full assessment · Against #1, ElevenLabs Agents API + MCP

#10

Hume EVI (Empathic Voice Interface)

C 56.7/100

Hume's hosted voice-agent service, accessed through a WebSocket API.

Verdict Speech-language model that reads the caller's tone and answers with matching prosody. One account-wide key, sent in the WebSocket and Twilio webhook URLs. Hume ends API access on 13 November 2026.

Choose it for Consumer and coaching products where the caller's tone matters and English is enough on EVI 3.

Strengths

  • Speech-language model that reads the caller's tone and answers with matching prosody
  • 4 to 7 cents a minute including Hume's own model
  • External LLMs or a custom endpoint for tool calling

Weaknesses

  • Hume ends access to the TTS and EVI APIs on 13 November 2026 and deletes account data after that date
  • One account-wide key, sent in the WebSocket and Twilio webhook URLs
  • Privacy page contradicts itself on training use of EVI data

Price $14 / moAuth OAuth or keyx402 nohosted

Full assessment · Against #1, ElevenLabs Agents API + MCP

Head to head

All 91 comparisons in this category

Questions

What are the highest-rated conversational voice-agent APIs for AI agents?

ElevenLabs Agents API + MCP has the highest benchmark score of the 14 ranked conversational voice-agent APIs, 71.3 (BB). Retell AI API + MCP is second with 69.1 (B).

How many conversational voice-agent APIs are agent-ready?

1 of the 14 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.

Which conversational voice-agent APIs accept x402 payments?

None of the ranked listings here accepts x402 for its main call yet.

How is this list ranked?

By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.

How this list is made

The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.

Full ranked table · 91 head-to-head comparisons · Best tools in every category

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.