# Best text-to-speech APIs for AI agents (slim) > Amazon Polly (BB), ElevenLabs Text to Speech API + MCP (BB) and Deepgram Text-to-Speech (Aura-2, Flux TTS) (BB) lead the 10 ranked text-to-speech APIs. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Full: https://www.anchorterminal.com/best/text-to-speech/index.md (~5,400 tokens) · this version ~1,530 tokens · JSON https://www.anchorterminal.com/best/text-to-speech/index.json · canonical https://www.anchorterminal.com/best/text-to-speech/ - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 All 10 ranked text-to-speech APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 10 · agent-ready (BB or better): 5 · accept x402: 0 · hosted endpoints: 10 - Full ranked table: https://www.anchorterminal.com/categories/text-to-speech.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/text-to-speech/index.md (55) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Amazon Polly](https://www.anchorterminal.com/tools/amazon-polly.md) | BB | 75.6 | Operators already on AWS who want predictable, cheap speech for prompts, notifications and IVR, or generative voices with bidirectional streaming. | Pay per use | hosted | | 2 | [ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/tools/elevenlabs-tts.md) | BB | 75.2 | Best when an agent needs many languages, many voices or the most expressive models from one vendor, and the operator can accept training by default or pay for Enterprise. | $6 / mo | hosted and local | | 3 | [Deepgram Text-to-Speech (Aura-2, Flux TTS)](https://www.anchorterminal.com/tools/deepgram-tts.md) | BB | 72.7 | English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs. | Pay per use | hosted and local | | 4 | [Azure AI Speech text-to-speech](https://www.anchorterminal.com/tools/azure-text-to-speech.md) | BB | 71.4 | Operators on Azure who need SSML control, many languages and voices, and enterprise access control. | $960 / mo | hosted | | 5 | [Murf TTS API + MCP](https://www.anchorterminal.com/tools/murf-tts.md) | BB | 70.6 | Budget voice-agent output where 2 to 5 concurrent streams are enough, or for studio-style renders on Gen2. | Freemium | hosted and local | | 6 | [Cartesia Sonic TTS API + MCP](https://www.anchorterminal.com/tools/cartesia-tts.md) | B | 63.8 | Real-time voice agents that stream LLM output straight into speech and need stable, pinnable models. | $5 / mo | hosted and local | | 7 | [Soniox Text-to-Speech](https://www.anchorterminal.com/tools/soniox-tts.md) | B | 63.7 | Multilingual agents that need any voice in any of 60-plus languages, short turns and strict data handling. | Pay per use | hosted | | 8 | [Fish Audio TTS API](https://www.anchorterminal.com/tools/fish-audio-tts.md) | C | 60.7 | Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes. | Pay per use | hosted | | 9 | [Rime TTS API + MCP](https://www.anchorterminal.com/tools/rime-tts.md) | C | 55.8 | English and a handful of other languages in live voice agents where the operator cares most about data terms and price. | Freemium | hosted | | 10 | [Resemble AI Text-to-Speech API](https://www.anchorterminal.com/tools/resemble-ai-tts.md) | D | 50.3 | Operators who already use Resemble for detection or watermarking and want one vendor for both. | $350 / mo | hosted | ## Picks by need - Highest score overall: [Amazon Polly](https://www.anchorterminal.com/tools/amazon-polly.md), BB, 75.6/100 on the benchmark. Also [ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/tools/elevenlabs-tts.md), BB, 75.2/100. - Schema & documentation: [Deepgram Text-to-Speech (Aura-2, Flux TTS)](https://www.anchorterminal.com/tools/deepgram-tts.md), 95/100 on schema & documentation, against 90 for the overall leader. - Agent ergonomics: [ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/tools/elevenlabs-tts.md), 92/100 on agent ergonomics, against 82 for the overall leader. - Security & auth: [Azure AI Speech text-to-speech](https://www.anchorterminal.com/tools/azure-text-to-speech.md), 90/100 on security & auth, against 80 for the overall leader. - Maintenance & community: [Azure AI Speech text-to-speech](https://www.anchorterminal.com/tools/azure-text-to-speech.md), 80/100 on maintenance & community, against 50 for the overall leader. - Transparency & trust: [Azure AI Speech text-to-speech](https://www.anchorterminal.com/tools/azure-text-to-speech.md), 85/100 on transparency & trust, against 77 for the overall leader. - A hosted MCP endpoint: [ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/tools/elevenlabs-tts.md), remote MCP server, nothing to install. ## How to choose - Time to first audio: Check the time to first audio when streaming, because a voice agent that waits for the full clip before speaking leaves the caller in silence. - Pronunciation and SSML control: Check whether SSML or a pronunciation dictionary can fix names, numbers and acronyms, because a misread account code or product name is hard for a listener to correct. - Voice consistency over length: Check whether the voice keeps its tone and pace across a long passage, since drift over a long narration can make it sound like separate speakers. - Cost of a fixed script: Check the price per character or per minute of output, and whether custom voices carry an extra fee, because the bill grows with every word. - How the benchmark tests this category: One fixed script through every API. We measure time to first audio, check pronunciation of names, numbers and acronyms, judge naturalness blind, listen for drift across a long passage, and price the whole script. Each listing's verdict, strengths and weaknesses: https://www.anchorterminal.com/best/text-to-speech/index.md