# Best text-to-speech APIs for AI agents > Amazon Polly (BB), ElevenLabs Text to Speech API + MCP (BB) and Deepgram Text-to-Speech (Aura-2, Flux TTS) (BB) lead the 10 ranked text-to-speech APIs. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Canonical: https://www.anchorterminal.com/best/text-to-speech/ - Markdown: https://www.anchorterminal.com/best/text-to-speech/index.md (~5,400 tokens) - Slim: https://www.anchorterminal.com/best/text-to-speech/index.min.md (~1,530 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/best/text-to-speech/index.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 All 10 ranked text-to-speech APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 10 · agent-ready (BB or better): 5 · accept x402: 0 · hosted endpoints: 10 - Full ranked table: https://www.anchorterminal.com/categories/text-to-speech.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/text-to-speech/index.md (55) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Amazon Polly](https://www.anchorterminal.com/tools/amazon-polly.md) | BB | 75.6 | Operators already on AWS who want predictable, cheap speech for prompts, notifications and IVR, or generative voices with bidirectional streaming. | Pay per use | hosted | | 2 | [ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/tools/elevenlabs-tts.md) | BB | 75.2 | Best when an agent needs many languages, many voices or the most expressive models from one vendor, and the operator can accept training by default or pay for Enterprise. | $6 / mo | hosted and local | | 3 | [Deepgram Text-to-Speech (Aura-2, Flux TTS)](https://www.anchorterminal.com/tools/deepgram-tts.md) | BB | 72.7 | English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs. | Pay per use | hosted and local | | 4 | [Azure AI Speech text-to-speech](https://www.anchorterminal.com/tools/azure-text-to-speech.md) | BB | 71.4 | Operators on Azure who need SSML control, many languages and voices, and enterprise access control. | $960 / mo | hosted | | 5 | [Murf TTS API + MCP](https://www.anchorterminal.com/tools/murf-tts.md) | BB | 70.6 | Budget voice-agent output where 2 to 5 concurrent streams are enough, or for studio-style renders on Gen2. | Freemium | hosted and local | | 6 | [Cartesia Sonic TTS API + MCP](https://www.anchorterminal.com/tools/cartesia-tts.md) | B | 63.8 | Real-time voice agents that stream LLM output straight into speech and need stable, pinnable models. | $5 / mo | hosted and local | | 7 | [Soniox Text-to-Speech](https://www.anchorterminal.com/tools/soniox-tts.md) | B | 63.7 | Multilingual agents that need any voice in any of 60-plus languages, short turns and strict data handling. | Pay per use | hosted | | 8 | [Fish Audio TTS API](https://www.anchorterminal.com/tools/fish-audio-tts.md) | C | 60.7 | Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes. | Pay per use | hosted | | 9 | [Rime TTS API + MCP](https://www.anchorterminal.com/tools/rime-tts.md) | C | 55.8 | English and a handful of other languages in live voice agents where the operator cares most about data terms and price. | Freemium | hosted | | 10 | [Resemble AI Text-to-Speech API](https://www.anchorterminal.com/tools/resemble-ai-tts.md) | D | 50.3 | Operators who already use Resemble for detection or watermarking and want one vendor for both. | $350 / mo | hosted | ## Picks by need - Highest score overall: [Amazon Polly](https://www.anchorterminal.com/tools/amazon-polly.md), BB, 75.6/100 on the benchmark. Also [ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/tools/elevenlabs-tts.md), BB, 75.2/100. - Schema & documentation: [Deepgram Text-to-Speech (Aura-2, Flux TTS)](https://www.anchorterminal.com/tools/deepgram-tts.md), 95/100 on schema & documentation, against 90 for the overall leader. - Agent ergonomics: [ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/tools/elevenlabs-tts.md), 92/100 on agent ergonomics, against 82 for the overall leader. - Security & auth: [Azure AI Speech text-to-speech](https://www.anchorterminal.com/tools/azure-text-to-speech.md), 90/100 on security & auth, against 80 for the overall leader. - Maintenance & community: [Azure AI Speech text-to-speech](https://www.anchorterminal.com/tools/azure-text-to-speech.md), 80/100 on maintenance & community, against 50 for the overall leader. - Transparency & trust: [Azure AI Speech text-to-speech](https://www.anchorterminal.com/tools/azure-text-to-speech.md), 85/100 on transparency & trust, against 77 for the overall leader. - A hosted MCP endpoint: [ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/tools/elevenlabs-tts.md), remote MCP server, nothing to install. ## How to choose - Time to first audio: Check the time to first audio when streaming, because a voice agent that waits for the full clip before speaking leaves the caller in silence. - Pronunciation and SSML control: Check whether SSML or a pronunciation dictionary can fix names, numbers and acronyms, because a misread account code or product name is hard for a listener to correct. - Voice consistency over length: Check whether the voice keeps its tone and pace across a long passage, since drift over a long narration can make it sound like separate speakers. - Cost of a fixed script: Check the price per character or per minute of output, and whether custom voices carry an extra fee, because the bill grows with every word. - How the benchmark tests this category: One fixed script through every API. We measure time to first audio, check pronunciation of names, numbers and acronyms, judge naturalness blind, listen for drift across a long passage, and price the whole script. ## Each one in detail ### 1. Amazon Polly, BB 75.6/100 AWS's speech synthesis API with four engines (standard, neural, long-form and generative) and about 110 voices in 42 languages and variants. - Verdict: IAM policies scope access per action and resource, and CloudTrail logs each call. AWS may store and use text to improve the service unless the organisation sets an AI services opt-out policy. - Choose it for: Operators already on AWS who want predictable, cheap speech for prompts, notifications and IVR, or generative voices with bidirectional streaming. - Strength: IAM policies scope access per action and resource, and CloudTrail logs each call - Strength: Quotas published per operation and engine, with backoff and jitter guidance for throttling - Strength: Covered by the Amazon Machine Learning Language SLA - Weakness: AWS may store and use text to improve the service unless the organisation sets an AI services opt-out policy - Weakness: A new account needs a card, and the monthly free characters only apply to accounts opened before 2025-07-15 - Weakness: Neural, long-form and generative synthesis is limited to 8 requests a second by default - Price: Pay per use · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/amazon-polly.md ### 2. ElevenLabs Text to Speech API + MCP, BB 75.2/100 ElevenLabs text-to-speech over REST, HTTP streaming and WebSocket. - Verdict: Keys can be limited to chosen endpoints and given a credit quota, and service accounts hold keys that don't belong to a person. The Terms let ElevenLabs train on content unless you opt out, which the Eleven v4 launch page contradicts, and v4 runs through Text to Dialogue rather than the classic text-to-speech endpoint. - Choose it for: Best when an agent needs many languages, many voices or the most expressive models from one vendor, and the operator can accept training by default or pay for Enterprise. - Strength: Keys can be limited to chosen endpoints and given a credit quota, and service accounts hold keys that don't belong to a person - Strength: Error reference with about 90 codes, each carrying `type`, `code`, `message` and `request_id` - Strength: OpenAPI file, llms.txt and Markdown twins of every docs page - Weakness: Content may be used for training unless you opt out under Data use, though the Eleven v4 launch page says it isn't used without consent - Weakness: Zero retention (`enable_logging=false`) is Enterprise only, and TTS text and audio are kept by default - Weakness: No SLA published for self-serve plans - Price: $6 / mo · Auth: OAuth or key · x402: no · Where: hosted and local - Full assessment: https://www.anchorterminal.com/tools/elevenlabs-tts.md - Against #1: https://www.anchorterminal.com/compare/amazon-polly-vs-elevenlabs-tts.md ### 3. Deepgram Text-to-Speech (Aura-2, Flux TTS), BB 72.7/100 Deepgram's text-to-speech API for generating spoken audio. - Verdict: OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages. Requests can be kept for training unless each one sets `mip_opt_out=true`. - Choose it for: English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs. - Strength: OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages - Strength: Flux TTS's Interrupt event returns `text_spoken` and `text_remaining` on barge-in - Strength: Keys carry roles and scopes, and browser tokens live 30 seconds - Weakness: Requests can be kept for training unless each one sets `mip_opt_out=true` - Weakness: Flux TTS is English only and capped at 5 concurrent streams in the EU, Australia and India below Enterprise - Weakness: No SSML, and Flux TTS strips other vendors' tags with a warning - Price: Pay per use · Auth: API key · x402: no · Where: hosted and local - Full assessment: https://www.anchorterminal.com/tools/deepgram-tts.md - Against #1: https://www.anchorterminal.com/compare/amazon-polly-vs-deepgram-tts.md ### 4. Azure AI Speech text-to-speech, BB 71.4/100 Azure's text-to-speech service for generating spoken audio. - Verdict: Real-time synthesis keeps neither the input text nor the output audio. An Azure subscription needs a card, even for the free F0 tier. - Choose it for: Operators on Azure who need SSML control, many languages and voices, and enterprise access control. - Strength: Real-time synthesis keeps neither the input text nor the output audio - Strength: Full SSML with speaking styles, prosody, phonemes, lexicons and up to 50 voice or audio tags a request - Strength: Microsoft Entra ID with role-based access, or two rotatable keys - Weakness: An Azure subscription needs a card, even for the free F0 tier - Weakness: No llms.txt and no OpenAPI file for text-to-speech found - Weakness: 429s often reflect busy capacity for a voice in a region, which a quota increase doesn't fix - Price: $960 / mo · Auth: OAuth or key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/azure-text-to-speech.md - Against #1: https://www.anchorterminal.com/compare/amazon-polly-vs-azure-text-to-speech.md ### 5. Murf TTS API + MCP, BB 70.6/100 Murf's API for generating speech from text. - Verdict: Falcon 2 costs 1 cent per 1,000 characters on pay as you go, with a $2 minimum purchase. Falcon 2 concurrency is 2 outside US-East on free and pay as you go, 5 on US-East. - Choose it for: Budget voice-agent output where 2 to 5 concurrent streams are enough, or for studio-style renders on Gen2. - Strength: Falcon 2 costs 1 cent per 1,000 characters on pay as you go, with a $2 minimum purchase - Strength: OpenAPI and AsyncAPI specs, llms.txt and Markdown twins of every docs page - Strength: Errors page lists 7 HTTP codes, 4 WebSocket errors and 8 warnings, each with cause and fix, and tells you to back off on 429 - Weakness: Falcon 2 concurrency is 2 outside US-East on free and pay as you go, 5 on US-East - Weakness: Gen2 streaming was deprecated on 2026-08-16 with 27 days' notice - Weakness: The security page says all data stays in AWS us-east-2, while Falcon 2 runs in 12 regions and the subprocessor list names four other clouds - Price: Freemium · Auth: API key · x402: no · Where: hosted and local - Full assessment: https://www.anchorterminal.com/tools/murf-tts.md - Against #1: https://www.anchorterminal.com/compare/amazon-polly-vs-murf-tts.md ### 6. Cartesia Sonic TTS API + MCP, B 63.8/100 Sonic text-to-speech, currently `sonic-3.6` in 44 languages, over `/tts/bytes`, `/tts/sse` and a WebSocket that takes streamed LLM text with contexts and word timestamps. - Verdict: WebSocket input accepts streamed LLM text with contexts, flushing and word timestamps. Five TTS incidents on the status page between 29 July and 21 August 2026. - Choose it for: Real-time voice agents that stream LLM output straight into speech and need stable, pinnable models. - Strength: WebSocket input accepts streamed LLM text with contexts, flushing and word timestamps - Strength: Dated model snapshots such as `sonic-3.6-2026-08-27` never change, and `Cartesia-Version` pins the API shape - Strength: OpenAPI and AsyncAPI files per API version listed in llms.txt - Weakness: Five TTS incidents on the status page between 29 July and 21 August 2026 - Weakness: TTS concurrency is 2 on Free, 3 on Pro, 5 on Startup and 15 on Scale - Weakness: The Terms allow training on inputs and outputs unless you opt out by form - Price: $5 / mo · Auth: OAuth or key · x402: no · Where: hosted and local - Full assessment: https://www.anchorterminal.com/tools/cartesia-tts.md - Against #1: https://www.anchorterminal.com/compare/amazon-polly-vs-cartesia-tts.md ### 7. Soniox Text-to-Speech, B 63.7/100 Streaming and REST text-to-speech (`tts-rt-v2`) in 60+ languages, where every voice speaks every language. - Verdict: The security page states that content is not stored by default or used for training. Audio output is capped at two minutes per request or stream. - Choose it for: Multilingual agents that need any voice in any of 60-plus languages, short turns and strict data handling. - Strength: Nothing is stored unless you ask and nothing trains on your content, per the security page - Strength: One error format with a stable `error_type`, `request_id` and `more_info` link - Strength: Separate TTS REST and TTS Real-time status components in four regions, 100 per cent uptime shown over 90 days - Weakness: Audio stops at 2 minutes per request or stream and the cap can't be raised - Weakness: 3 concurrent requests and 100 requests a minute by default - Weakness: `tts-rt-v1` was removed 20 days after its deprecation notice - Price: Pay per use · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/soniox-tts.md - Against #1: https://www.anchorterminal.com/compare/amazon-polly-vs-soniox-tts.md ### 8. Fish Audio TTS API, C 60.7/100 Fish Audio's API turns text into speech with the `s2.1-pro` model in 83 languages, over a REST endpoint, a timestamped stream and a WebSocket that accepts text as it is produced. - Verdict: The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation. - Choose it for: Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes. - Strength: Public OpenAPI 3.1 file, llms.txt, Markdown pages and two installable agent skills for the SDKs and the raw API - Strength: WebSocket input takes text as it is produced, with `flush` and `stop` events and a variant that returns word timestamps - Strength: `s2.1-pro-free` runs the production model at $0 under fair-use limits, through 30 November 2026 - Weakness: The terms allow Usage Data and Content to train models, with no opt-out found - Weakness: An unrecognised `model` header falls back to paid `s2.1-pro` without an error - Weakness: The Text-to-Speech API status component shows downtime on 11 days in 90, the longest 58 minutes on 6 August 2026, mostly on the free model - Price: Pay per use · Auth: OAuth or key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/fish-audio-tts.md - Against #1: https://www.anchorterminal.com/compare/amazon-polly-vs-fish-audio-tts.md ### 9. Rime TTS API + MCP, C 55.8/100 Rime's streaming text-to-speech for voice agents. - Verdict: Only character counts are kept by default, and customer data isn't used for training without an opt-in. No OpenAPI or AsyncAPI file and no official SDKs. - Choose it for: English and a handful of other languages in live voice agents where the operator cares most about data terms and price. - Strength: Only character counts are kept by default, and customer data isn't used for training without an opt-in - Strength: Coda at $0.05 and Mist v3 at $0.03 per 1,000 characters, with free minutes and no card - Strength: 20 concurrent generations on Starter - Weakness: No OpenAPI or AsyncAPI file and no official SDKs - Weakness: Errors are plain-text messages without machine-readable codes - Weakness: Arcana's cloud retirement came 17 days after the announcement - Price: Freemium · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/rime-tts.md - Against #1: https://www.anchorterminal.com/compare/amazon-polly-vs-rime-tts.md ### 10. Resemble AI Text-to-Speech API, D 50.3/100 Resemble's TTS API on its current Resemble Ultra model, which the changelog says is powered by xAI. - Verdict: OpenAPI file in JSON and YAML, llms.txt and Markdown pages. Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published. - Choose it for: Operators who already use Resemble for detection or watermarking and want one vendor for both. - Strength: OpenAPI file in JSON and YAML, llms.txt and Markdown pages - Strength: SSML with prompt, temperature and exaggeration controls and inline tags such as `[laugh]` - Strength: Status page checks Ultra HTTP synthesis directly, 100 per cent over 90 days - Weakness: Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published - Weakness: WebSocket streaming only on Business at $1,000 a month - Weakness: Errors are `success: false` and a message, no status codes listed - Price: $350 / mo · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/resemble-ai-tts.md - Against #1: https://www.anchorterminal.com/compare/amazon-polly-vs-resemble-ai-tts.md ## Head to head - [Amazon Polly vs ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/compare/amazon-polly-vs-elevenlabs-tts.md) - [Amazon Polly vs Deepgram Text-to-Speech (Aura-2, Flux TTS)](https://www.anchorterminal.com/compare/amazon-polly-vs-deepgram-tts.md) - [Amazon Polly vs Azure AI Speech text-to-speech](https://www.anchorterminal.com/compare/amazon-polly-vs-azure-text-to-speech.md) - [Amazon Polly vs Murf TTS API + MCP](https://www.anchorterminal.com/compare/amazon-polly-vs-murf-tts.md) - [Deepgram Text-to-Speech (Aura-2, Flux TTS) vs ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/compare/deepgram-tts-vs-elevenlabs-tts.md) - [Azure AI Speech text-to-speech vs ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/compare/azure-text-to-speech-vs-elevenlabs-tts.md) - [ElevenLabs Text to Speech API + MCP vs Murf TTS API + MCP](https://www.anchorterminal.com/compare/elevenlabs-tts-vs-murf-tts.md) - [Azure AI Speech text-to-speech vs Deepgram Text-to-Speech (Aura-2, Flux TTS)](https://www.anchorterminal.com/compare/azure-text-to-speech-vs-deepgram-tts.md) - [Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Murf TTS API + MCP](https://www.anchorterminal.com/compare/deepgram-tts-vs-murf-tts.md) - [Azure AI Speech text-to-speech vs Murf TTS API + MCP](https://www.anchorterminal.com/compare/azure-text-to-speech-vs-murf-tts.md) ## Questions ### What are the highest-rated text-to-speech APIs for AI agents? Amazon Polly has the highest benchmark score of the 10 ranked text-to-speech APIs, 75.6 (BB). ElevenLabs Text to Speech API + MCP is second with 75.2 (BB). ### How many text-to-speech APIs are agent-ready? 5 of the 10 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark. ### Which text-to-speech APIs accept x402 payments? None of the ranked listings here accepts x402 for its main call yet. ### How is this list ranked? By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026. ## How this list is made The order is the Anchor benchmark score, the same number as on each listing. Each listing is graded from public evidence against the benchmark checklist, and the picks are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.