Best of · Voice & speech
Best text-to-speech APIs for AI agents
All 10 ranked text-to-speech APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 10 ranked
- 5 agent-ready
- 10 hosted endpoints
- Updated 9 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Highest score overall
Amazon Polly BB
BB, 75.6/100 on the benchmark.
Also ElevenLabs Text to Speech API + MCP, BB, 75.2/100.
Schema & documentation
Deepgram Text-to-Speech (Aura-2, Flux TTS) BB
95/100 on schema & documentation, against 90 for the overall leader.
Agent ergonomics
ElevenLabs Text to Speech API + MCP BB
92/100 on agent ergonomics, against 82 for the overall leader.
Security & auth
Azure AI Speech text-to-speech BB
90/100 on security & auth, against 80 for the overall leader.
Maintenance & community
Azure AI Speech text-to-speech BB
80/100 on maintenance & community, against 50 for the overall leader.
Transparency & trust
Azure AI Speech text-to-speech BB
85/100 on transparency & trust, against 77 for the overall leader.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | Amazon Polly Amazon Web Services |
BB 75.6 | Operators already on AWS who want predictable, cheap speech for prompts, notifications and IVR, or generative voices with bidirectional streaming. | Pay per use | hosted |
| 2 | ElevenLabs Text to Speech API + MCP ElevenLabs |
BB 75.2 | Best when an agent needs many languages, many voices or the most expressive models from one vendor, and the operator can accept training by default or pay for Enterprise. | $6 / mo | hosted and local |
| 3 | Deepgram Text-to-Speech (Aura-2, Flux TTS) Deepgram |
BB 72.7 | English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs. | Pay per use | hosted and local |
| 4 | Azure AI Speech text-to-speech Microsoft Azure |
BB 71.4 | Operators on Azure who need SSML control, many languages and voices, and enterprise access control. | $960 / mo | hosted |
| 5 | Murf TTS API + MCP Murf |
BB 70.6 | Budget voice-agent output where 2 to 5 concurrent streams are enough, or for studio-style renders on Gen2. | Freemium | hosted and local |
| 6 | Cartesia Sonic TTS API + MCP Cartesia |
B 63.8 | Real-time voice agents that stream LLM output straight into speech and need stable, pinnable models. | $5 / mo | hosted and local |
| 7 | Soniox Text-to-Speech Soniox |
B 63.7 | Multilingual agents that need any voice in any of 60-plus languages, short turns and strict data handling. | Pay per use | hosted |
| 8 | Fish Audio TTS API Fish Audio |
C 60.7 | Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes. | Pay per use | hosted |
| 9 | Rime TTS API + MCP Rime Labs |
C 55.8 | English and a handful of other languages in live voice agents where the operator cares most about data terms and price. | Freemium | hosted |
| 10 | Resemble AI Text-to-Speech API Resemble AI |
D 50.3 | Operators who already use Resemble for detection or watermarking and want one vendor for both. | $350 / mo | hosted |
How to choose
- Time to first audioCheck the time to first audio when streaming, because a voice agent that waits for the full clip before speaking leaves the caller in silence.
- Pronunciation and SSML controlCheck whether SSML or a pronunciation dictionary can fix names, numbers and acronyms, because a misread account code or product name is hard for a listener to correct.
- Voice consistency over lengthCheck whether the voice keeps its tone and pace across a long passage, since drift over a long narration can make it sound like separate speakers.
- Cost of a fixed scriptCheck the price per character or per minute of output, and whether custom voices carry an extra fee, because the bill grows with every word.
How the benchmark tests this category. One fixed script through every API. We measure time to first audio, check pronunciation of names, numbers and acronyms, judge naturalness blind, listen for drift across a long passage, and price the whole script.
Each one in detail
Amazon Polly
BB 75.6/100AWS's speech synthesis API with four engines (standard, neural, long-form and generative) and about 110 voices in 42 languages and variants.
Verdict IAM policies scope access per action and resource, and CloudTrail logs each call. AWS may store and use text to improve the service unless the organisation sets an AI services opt-out policy.
Choose it for Operators already on AWS who want predictable, cheap speech for prompts, notifications and IVR, or generative voices with bidirectional streaming.
Strengths
- IAM policies scope access per action and resource, and CloudTrail logs each call
- Quotas published per operation and engine, with backoff and jitter guidance for throttling
- Covered by the Amazon Machine Learning Language SLA
Weaknesses
- AWS may store and use text to improve the service unless the organisation sets an AI services opt-out policy
- A new account needs a card, and the monthly free characters only apply to accounts opened before 2025-07-15
- Neural, long-form and generative synthesis is limited to 8 requests a second by default
Price Pay per useAuth API keyx402 nohosted
ElevenLabs Text to Speech API + MCP
BB 75.2/100ElevenLabs text-to-speech over REST, HTTP streaming and WebSocket.
Verdict Keys can be limited to chosen endpoints and given a credit quota, and service accounts hold keys that don't belong to a person. The Terms let ElevenLabs train on content unless you opt out, which the Eleven v4 launch page contradicts, and v4 runs through Text to Dialogue rather than the classic text-to-speech endpoint.
Choose it for Best when an agent needs many languages, many voices or the most expressive models from one vendor, and the operator can accept training by default or pay for Enterprise.
Strengths
- Keys can be limited to chosen endpoints and given a credit quota, and service accounts hold keys that don't belong to a person
- Error reference with about 90 codes, each carrying
type,code,messageandrequest_id - OpenAPI file, llms.txt and Markdown twins of every docs page
Weaknesses
- Content may be used for training unless you opt out under Data use, though the Eleven v4 launch page says it isn't used without consent
- Zero retention (
enable_logging=false) is Enterprise only, and TTS text and audio are kept by default - No SLA published for self-serve plans
Price $6 / moAuth OAuth or keyx402 nohosted and local
Deepgram Text-to-Speech (Aura-2, Flux TTS)
BB 72.7/100Deepgram's text-to-speech API for generating spoken audio.
Verdict OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages. Requests can be kept for training unless each one sets mip_opt_out=true.
Choose it for English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs.
Strengths
- OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages
- Flux TTS's Interrupt event returns
text_spokenandtext_remainingon barge-in - Keys carry roles and scopes, and browser tokens live 30 seconds
Weaknesses
- Requests can be kept for training unless each one sets
mip_opt_out=true - Flux TTS is English only and capped at 5 concurrent streams in the EU, Australia and India below Enterprise
- No SSML, and Flux TTS strips other vendors' tags with a warning
Price Pay per useAuth API keyx402 nohosted and local
Azure AI Speech text-to-speech
BB 71.4/100Azure's text-to-speech service for generating spoken audio.
Verdict Real-time synthesis keeps neither the input text nor the output audio. An Azure subscription needs a card, even for the free F0 tier.
Choose it for Operators on Azure who need SSML control, many languages and voices, and enterprise access control.
Strengths
- Real-time synthesis keeps neither the input text nor the output audio
- Full SSML with speaking styles, prosody, phonemes, lexicons and up to 50 voice or audio tags a request
- Microsoft Entra ID with role-based access, or two rotatable keys
Weaknesses
- An Azure subscription needs a card, even for the free F0 tier
- No llms.txt and no OpenAPI file for text-to-speech found
- 429s often reflect busy capacity for a voice in a region, which a quota increase doesn't fix
Price $960 / moAuth OAuth or keyx402 nohosted
Murf TTS API + MCP
BB 70.6/100Murf's API for generating speech from text.
Verdict Falcon 2 costs 1 cent per 1,000 characters on pay as you go, with a $2 minimum purchase. Falcon 2 concurrency is 2 outside US-East on free and pay as you go, 5 on US-East.
Choose it for Budget voice-agent output where 2 to 5 concurrent streams are enough, or for studio-style renders on Gen2.
Strengths
- Falcon 2 costs 1 cent per 1,000 characters on pay as you go, with a $2 minimum purchase
- OpenAPI and AsyncAPI specs, llms.txt and Markdown twins of every docs page
- Errors page lists 7 HTTP codes, 4 WebSocket errors and 8 warnings, each with cause and fix, and tells you to back off on 429
Weaknesses
- Falcon 2 concurrency is 2 outside US-East on free and pay as you go, 5 on US-East
- Gen2 streaming was deprecated on 2026-08-16 with 27 days' notice
- The security page says all data stays in AWS us-east-2, while Falcon 2 runs in 12 regions and the subprocessor list names four other clouds
Price FreemiumAuth API keyx402 nohosted and local
Cartesia Sonic TTS API + MCP
B 63.8/100Sonic text-to-speech, currently sonic-3.6 in 44 languages, over /tts/bytes, /tts/sse and a WebSocket that takes streamed LLM text with contexts and word timestamps.
Verdict WebSocket input accepts streamed LLM text with contexts, flushing and word timestamps. Five TTS incidents on the status page between 29 July and 21 August 2026.
Choose it for Real-time voice agents that stream LLM output straight into speech and need stable, pinnable models.
Strengths
- WebSocket input accepts streamed LLM text with contexts, flushing and word timestamps
- Dated model snapshots such as
sonic-3.6-2026-08-27never change, andCartesia-Versionpins the API shape - OpenAPI and AsyncAPI files per API version listed in llms.txt
Weaknesses
- Five TTS incidents on the status page between 29 July and 21 August 2026
- TTS concurrency is 2 on Free, 3 on Pro, 5 on Startup and 15 on Scale
- The Terms allow training on inputs and outputs unless you opt out by form
Price $5 / moAuth OAuth or keyx402 nohosted and local
Soniox Text-to-Speech
B 63.7/100Streaming and REST text-to-speech (tts-rt-v2) in 60+ languages, where every voice speaks every language.
Verdict The security page states that content is not stored by default or used for training. Audio output is capped at two minutes per request or stream.
Choose it for Multilingual agents that need any voice in any of 60-plus languages, short turns and strict data handling.
Strengths
- Nothing is stored unless you ask and nothing trains on your content, per the security page
- One error format with a stable
error_type,request_idandmore_infolink - Separate TTS REST and TTS Real-time status components in four regions, 100 per cent uptime shown over 90 days
Weaknesses
- Audio stops at 2 minutes per request or stream and the cap can't be raised
- 3 concurrent requests and 100 requests a minute by default
tts-rt-v1was removed 20 days after its deprecation notice
Price Pay per useAuth API keyx402 nohosted
Fish Audio TTS API
C 60.7/100Fish Audio's API turns text into speech with the s2.1-pro model in 83 languages, over a REST endpoint, a timestamped stream and a WebSocket that accepts text as it is produced.
Verdict The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.
Choose it for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.
Strengths
- Public OpenAPI 3.1 file, llms.txt, Markdown pages and two installable agent skills for the SDKs and the raw API
- WebSocket input takes text as it is produced, with
flushandstopevents and a variant that returns word timestamps s2.1-pro-freeruns the production model at $0 under fair-use limits, through 30 November 2026
Weaknesses
- The terms allow Usage Data and Content to train models, with no opt-out found
- An unrecognised
modelheader falls back to paids2.1-prowithout an error - The Text-to-Speech API status component shows downtime on 11 days in 90, the longest 58 minutes on 6 August 2026, mostly on the free model
Price Pay per useAuth OAuth or keyx402 nohosted
Rime TTS API + MCP
C 55.8/100Rime's streaming text-to-speech for voice agents.
Verdict Only character counts are kept by default, and customer data isn't used for training without an opt-in. No OpenAPI or AsyncAPI file and no official SDKs.
Choose it for English and a handful of other languages in live voice agents where the operator cares most about data terms and price.
Strengths
- Only character counts are kept by default, and customer data isn't used for training without an opt-in
- Coda at $0.05 and Mist v3 at $0.03 per 1,000 characters, with free minutes and no card
- 20 concurrent generations on Starter
Weaknesses
- No OpenAPI or AsyncAPI file and no official SDKs
- Errors are plain-text messages without machine-readable codes
- Arcana's cloud retirement came 17 days after the announcement
Price FreemiumAuth API keyx402 nohosted
Resemble AI Text-to-Speech API
D 50.3/100Resemble's TTS API on its current Resemble Ultra model, which the changelog says is powered by xAI.
Verdict OpenAPI file in JSON and YAML, llms.txt and Markdown pages. Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published.
Choose it for Operators who already use Resemble for detection or watermarking and want one vendor for both.
Strengths
- OpenAPI file in JSON and YAML, llms.txt and Markdown pages
- SSML with prompt, temperature and exaggeration controls and inline tags such as
[laugh] - Status page checks Ultra HTTP synthesis directly, 100 per cent over 90 days
Weaknesses
- Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published
- WebSocket streaming only on Business at $1,000 a month
- Errors are
success: falseand a message, no status codes listed
Price $350 / moAuth API keyx402 nohosted
Head to head
- Amazon Polly vs ElevenLabs Text to Speech API + MCP BB 75.6 vs BB 75.2
- Amazon Polly vs Deepgram Text-to-Speech (Aura-2, Flux TTS) BB 75.6 vs BB 72.7
- Amazon Polly vs Azure AI Speech text-to-speech BB 75.6 vs BB 71.4
- Amazon Polly vs Murf TTS API + MCP BB 75.6 vs BB 70.6
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs ElevenLabs Text to Speech API + MCP BB 72.7 vs BB 75.2
- Azure AI Speech text-to-speech vs ElevenLabs Text to Speech API + MCP BB 71.4 vs BB 75.2
- ElevenLabs Text to Speech API + MCP vs Murf TTS API + MCP BB 75.2 vs BB 70.6
- Azure AI Speech text-to-speech vs Deepgram Text-to-Speech (Aura-2, Flux TTS) BB 71.4 vs BB 72.7
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Murf TTS API + MCP BB 72.7 vs BB 70.6
- Azure AI Speech text-to-speech vs Murf TTS API + MCP BB 71.4 vs BB 70.6
Questions
What are the highest-rated text-to-speech APIs for AI agents?
Amazon Polly has the highest benchmark score of the 10 ranked text-to-speech APIs, 75.6 (BB). ElevenLabs Text to Speech API + MCP is second with 75.2 (BB).
How many text-to-speech APIs are agent-ready?
5 of the 10 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which text-to-speech APIs accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 55 head-to-head comparisons · Best tools in every category