Best of · Voice & speech
Best conversational voice-agent APIs
The 10 highest-scoring of 14 conversational voice-agent APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 14 ranked
- 1 agent-ready
- 14 hosted endpoints
- Updated 9 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Highest score overall
ElevenLabs Agents API + MCP BB
BB, 71.3/100 on the benchmark.
Also Retell AI API + MCP, B, 69.1/100.
Maintenance & community
88/100 on maintenance & community, against 85 for the overall leader.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | ElevenLabs Agents API + MCP ElevenLabs |
BB 71.3 | Teams that want a managed pipeline with ElevenLabs voices, a choice of LLM, wide telephony integration and web and mobile SDKs. | $22 / mo | hosted |
| 2 | Retell AI API + MCP Retell AI |
B 69.1 | Teams that want a self-serve phone agent with clear per-component pricing, scoped keys and a choice of LLM and voice from Retell's menu. | $2 / mo | hosted |
| 3 | Deepgram Voice Agent API Deepgram |
B 68.1 | Developers who already run telephony (Twilio, Amazon Connect, Genesys, AudioCodes) and want a cheap managed pipeline with their choice of LLM. | Pay per use | hosted and local |
| 4 | AssemblyAI Voice Agent API AssemblyAI, Inc. |
B 64.4 | A developer who wants one managed pipeline at a flat rate and is content with AssemblyAI's own speech models and voices, with an optional model of their own. | Pay per use | hosted |
| 5 | Bland AI API + MCP Bland AI |
B 63.8 | Teams that want one predictable per-minute bill for outbound and inbound phone agents, and personal agents that need their own US number. | $299 / mo | hosted and local |
| 6 | Vapi API + MCP Vapi |
B 63.5 | Developers who want to choose every provider, bring their own keys and tune cost against latency. | $29 / mo | hosted and local |
| 7 | OpenAI Realtime API OpenAI |
B 63.3 | Teams building their own voice agent on one speech-to-speech model who want browser, server and phone transports from one vendor and can run a backend for secrets and tools. | Pay per use | hosted |
| 8 | Ultravox Realtime API Ultravox (Fixie.ai) |
C 58.5 | Cost-sensitive voice agents that bring their own telephony and want a speech-native model. | $100 / mo | hosted |
| 9 | Gemini Live API |
C 57.1 | Teams building their own voice or vision assistant on a single speech-to-speech model with function calling and Search grounding, who can run a backend for tokens and audio transport. | $14 / 1k req | hosted |
| 10 | Hume EVI (Empathic Voice Interface) Hume AI |
C 56.7 | Consumer and coaching products where the caller's tone matters and English is enough on EVI 3. | $14 / mo | hosted |
4 more are ranked in the full table.
How to choose
- Speech-to-speech or pipelineCheck whether the product runs one speech-to-speech model or a separate recognition, language and synthesis chain, since each changes latency and what you can swap.
- Handling of interruptionsAsk how the platform stops speaking when a caller interrupts, because a slow stop leaves the agent talking over the customer.
- Tool calls during a live callCheck how a tool call is confirmed to the caller and what happens when it fails, since a long pause or a wrong booking is what the caller hears.
- Cost per completed callWork out the cost of a finished call from per-minute, language model and telephony charges together, because a low per-minute rate can hide a costly language model.
How the benchmark tests this category. The same booking and support tasks on every platform. We measure response latency, how interruptions are handled, tool-call success, recovery after a misunderstanding and the total cost per completed call, and record whether the product runs a direct speech-to-speech model or an STT, model and TTS pipeline, since those are different setups.
Each one in detail
ElevenLabs Agents API + MCP
BB 71.3/100ElevenAgents (formerly Conversational AI) runs hosted voice agents as a pipeline of a fine-tuned ElevenLabs ASR model, an LLM of your choice or your own, ElevenLabs TTS and a proprietary turn-taking model.
Verdict API keys scoped to endpoint groups, with per-key credit limits and service accounts. $0.08 a minute excludes LLM tokens and carrier minutes.
Choose it for Teams that want a managed pipeline with ElevenLabs voices, a choice of LLM, wide telephony integration and web and mobile SDKs.
Strengths
- API keys scoped to endpoint groups, with per-key credit limits and service accounts
- Hosted MCP server that signs in with OAuth, with EU, India and Singapore endpoints
- Public OpenAPI, llms.txt and an errors page with codes, fixes and request IDs
Weaknesses
- $0.08 a minute excludes LLM tokens and carrier minutes
- At least six incidents since July where agent calls failed or didn't start
- Conversation data kept 2 years by default
Price $22 / moAuth OAuth or keyx402 nohosted
Retell AI API + MCP
B 69.1/100Hosted platform for phone and web voice agents, built as single-prompt agents or node-based conversation flows.
Verdict Published per-minute price for every component, from $0.07 a minute all in. Call data kept indefinitely unless retention is set per agent.
Choose it for Teams that want a self-serve phone agent with clear per-component pricing, scoped keys and a choice of LLM and voice from Retell's menu.
Strengths
- Published per-minute price for every component, from $0.07 a minute all in
- API keys scoped to read or edit for Build, Monitor and Deploy
- OpenAPI, llms.txt and a deprecation feed with RSS
Weaknesses
- Call data kept indefinitely unless retention is set per agent
- MCP tools carry no read-only or destructive annotations
- Web and phone calls disrupted for 69 minutes on 5 September 2026
Price $2 / moAuth API keyx402 nohosted
Deepgram Voice Agent API
B 68.1/100One WebSocket that runs Deepgram STT (Flux or Nova-3), a managed or bring-your-own LLM and Deepgram or third-party TTS, with turn-taking, barge-in and function calling.
Verdict $0.075 a minute all in on Standard, $0.050 with your own LLM and TTS, $200 of credit with no card. No phone numbers or SIP, so you bridge Twilio or another carrier yourself.
Choose it for Developers who already run telephony (Twilio, Amazon Connect, Genesys, AudioCodes) and want a cheap managed pipeline with their choice of LLM.
Strengths
- $0.075 a minute all in on Standard, $0.050 with your own LLM and TTS, $200 of credit with no card
- AsyncAPI description of the agent socket plus a public OpenAPI file
- Role-based API keys with expiry and 30-second browser tokens
Weaknesses
- No phone numbers or SIP, so you bridge Twilio or another carrier yourself
- Fifteen status incidents since July, four of them an hour or longer on parts the agent uses
- Call audio is kept for model improvement unless
mip_opt_outis set
Price Pay per useAuth API keyx402 nohosted and local
AssemblyAI Voice Agent API
B 64.4/100AssemblyAI's Voice Agent API runs a spoken conversation over one WebSocket, combining its own speech-to-text, language model and text-to-speech. Agents are stored through a REST API and reached from a browser, a server or an inbound Twilio SIP number.
Verdict One $0.075 a minute rate covers speech-to-text, the managed model, speech output, recordings and hosting, and the socket has an AsyncAPI file, typed error codes and a 30-second resume window. No concurrent-session limit is published, the status page has no Voice Agent component, and phone calls are inbound only through a Twilio trunk the customer owns.
Choose it for A developer who wants one managed pipeline at a flat rate and is content with AssemblyAI's own speech models and voices, with an optional model of their own.
Strengths
- One published rate of $4.50 an hour ($0.075 a minute), billed per second, with $50 of credit and no card for new accounts
- AsyncAPI 3.0 file for the socket, an OpenAPI file for the token endpoint, llms.txt and a Markdown copy of every docs page
- Session errors carry a code, a message and sometimes the offending field, and the docs name the three codes that are safe to retry
Weaknesses
- No number is published for the concurrent-session limit that the
concurrency_exceedederror enforces - status.assemblyai.com lists the Asynchronous API, Streaming API and LLM Gateway, with no Voice Agent component
- Telephony is inbound only, over a SIP trunk on a Twilio account the customer owns
Price Pay per useAuth API keyx402 nohosted
Bland AI API + MCP
B 63.8/100Phone-first voice-agent platform that runs its own speech recognition, language model, TTS and telephony, billed as one per-minute rate.
Verdict One per-minute rate covering STT, LLM, TTS and telephony, $0.14 on Start and $0.12 on Build. No OpenAPI document.
Choose it for Teams that want one predictable per-minute bill for outbound and inbound phone agents, and personal agents that need their own US number.
Strengths
- One per-minute rate covering STT, LLM, TTS and telephony, $0.14 on Start and $0.12 on Build
- Start plan with 2 credits and an inbound number, no card
- Hosted MCP server with 42 tools labelled read, write or destructive, with confirmation on destructive ones
Weaknesses
- No OpenAPI document
- Start caps at 10 concurrent calls and 100 calls a day
- Failed calls and outbound attempts cost $0.015 each
Price $299 / moAuth API keyx402 nohosted and local
Vapi API + MCP
B 63.5/100Developer platform for phone and web voice agents.
Verdict Voice agents can combine supported transcription, model and voice providers, including customer-supplied endpoints. Per-minute cost depends on the selected providers.
Choose it for Developers who want to choose every provider, bring their own keys and tune cost against latency.
Strengths
- Any mix of transcriber, model and voice provider, or your own keys and endpoints
- OpenAPI file, llms.txt and a weekly changelog
- Public keys limited to allowed origins and assistants
Weaknesses
- Real per-minute cost depends on provider choices
- Usage only has 4 concurrent lines and 14 days of retention
- MCP tools that place calls or buy numbers carry no annotations
Price $29 / moAuth API keyx402 nohosted and local
OpenAI Realtime API
B 63.3/100OpenAI's Realtime API runs spoken conversations with speech-to-speech models such as gpt-realtime-2.1. Clients connect over WebRTC, WebSocket or SIP, and sessions support interruptions, function calling and remote MCP tools.
Verdict A generally available speech-to-speech API with WebRTC, WebSocket and SIP transports, a public OpenAPI file that types its events, and per-token prices. Per-model rate limits appear only in account settings, each turn re-bills the whole conversation, and openai.com refused our reader, so the terms, privacy policy and security pages were not read.
Choose it for Teams building their own voice agent on one speech-to-speech model who want browser, server and phone transports from one vendor and can run a backend for secrets and tools.
Strengths
- The OpenAPI 3.1 file in
openai/openai-openapicovers nine Realtime paths and types 13 client events and 48 server events as schemas. - Browser clients use client secrets that expire after 10 seconds to 2 hours, 10 minutes by default, so the API key stays on a server.
- Remote MCP tools accept an
allowed_toolslist, andrequire_approvaldefaults toalwaysin the OpenAPI file.
Weaknesses
- openai.com answered 403 to our reader, so the service terms, privacy policy, sub-processor list and security pages were not read.
- Realtime rate limits are shown only in account settings. The public table gives numbers for GPT-6 model families only.
- Every response re-sends the whole conversation, so later turns cost more. Audio is $32 in and $64 out per 1M tokens on
gpt-realtime-2.1.
Price Pay per useAuth API keyx402 nohosted
Ultravox Realtime API
C 58.5/100Hosted voice agents on Ultravox, an open-weight model that takes speech directly with no speech-to-text step and answers through a TTS voice.
Verdict Flat $0.05 a minute including model and built-in voices. No MCP server.
Choose it for Cost-sensitive voice agents that bring their own telephony and want a speech-native model.
Strengths
- Flat $0.05 a minute including model and built-in voices
- Speech-native model with no separate STT stage, v0.7 weights public under MIT
- 429 and 503 responses with Retry-After and backoff guidance
Weaknesses
- No MCP server
- News page, deprecation guide and Python client last updated in December 2025
- Plain API keys with no scopes
Price $100 / moAuth API keyx402 nohosted
Gemini Live API
C 57.1/100Google's Live API runs real-time spoken conversations with Gemini audio-to-audio models over a stateful WebSocket. It takes audio, images and text, speaks back, and supports interruptions, function calling and Google Search grounding.
Verdict A speech-to-speech API with per-token prices published, a free tier and a generally available model, gemini-3.8-live, since 15 September 2026. The documented raw WebSocket connection carries the API key in the URL, connections reset about every 10 minutes, and no SLA or readable incident history was found for the Developer API.
Choose it for Teams building their own voice or vision assistant on a single speech-to-speech model with function calling and Search grounding, who can run a backend for tokens and audio transport.
Strengths
gemini-3.8-livewent generally available on 15 September 2026 with asynchronous function calling as the default and three response scheduling modes.- Ephemeral tokens can be single use, expire in 30 minutes by default and be locked to a model and session configuration.
- Prices are public per million tokens with per-minute equivalents, $0.005 a minute of audio in and $0.018 a minute of audio out.
Weaknesses
- The WebSocket guide authenticates with the API key as a
keyquery parameter, and ephemeral tokens as anaccess_tokenquery parameter. - The status page at aistudio.google.com/status renders in the browser only, so no incident history could be read, and no SLA was found for the Developer API.
- No machine-readable contract for the WebSocket messages was found. The Discovery document types only the setup and token schemas.
Price $14 / 1k reqAuth API keyx402 nohosted
Hume EVI (Empathic Voice Interface)
C 56.7/100Hume's hosted voice-agent service, accessed through a WebSocket API.
Verdict Speech-language model that reads the caller's tone and answers with matching prosody. One account-wide key, sent in the WebSocket and Twilio webhook URLs. Hume ends API access on 13 November 2026.
Choose it for Consumer and coaching products where the caller's tone matters and English is enough on EVI 3.
Strengths
- Speech-language model that reads the caller's tone and answers with matching prosody
- 4 to 7 cents a minute including Hume's own model
- External LLMs or a custom endpoint for tool calling
Weaknesses
- Hume ends access to the TTS and EVI APIs on 13 November 2026 and deletes account data after that date
- One account-wide key, sent in the WebSocket and Twilio webhook URLs
- Privacy page contradicts itself on training use of EVI data
Price $14 / moAuth OAuth or keyx402 nohosted
Head to head
- ElevenLabs Agents API + MCP vs Retell AI API + MCP BB 71.3 vs B 69.1
- Deepgram Voice Agent API vs ElevenLabs Agents API + MCP B 68.1 vs BB 71.3
- AssemblyAI Voice Agent API vs ElevenLabs Agents API + MCP B 64.4 vs BB 71.3
- Bland AI API + MCP vs ElevenLabs Agents API + MCP B 63.8 vs BB 71.3
- Deepgram Voice Agent API vs Retell AI API + MCP B 68.1 vs B 69.1
- AssemblyAI Voice Agent API vs Retell AI API + MCP B 64.4 vs B 69.1
- Bland AI API + MCP vs Retell AI API + MCP B 63.8 vs B 69.1
- AssemblyAI Voice Agent API vs Deepgram Voice Agent API B 64.4 vs B 68.1
- Bland AI API + MCP vs Deepgram Voice Agent API B 63.8 vs B 68.1
- AssemblyAI Voice Agent API vs Bland AI API + MCP B 64.4 vs B 63.8
Questions
What are the highest-rated conversational voice-agent APIs for AI agents?
ElevenLabs Agents API + MCP has the highest benchmark score of the 14 ranked conversational voice-agent APIs, 71.3 (BB). Retell AI API + MCP is second with 69.1 (B).
How many conversational voice-agent APIs are agent-ready?
1 of the 14 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which conversational voice-agent APIs accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 91 head-to-head comparisons · Best tools in every category