# Text-to-speech APIs for AI agents > 10 text-to-speech listings ranked by the Anchor benchmark. Leader Amazon Polly (BB). APIs that turn text into speech, streamed for live conversation or rendered for long-form audio. Compared on time to first audio, pronunciation, how natural it sounds, long-form consistency, languages and cost for a fixed script. - Canonical: https://www.anchorterminal.com/categories/text-to-speech - Markdown: https://www.anchorterminal.com/categories/text-to-speech.md (~2,950 tokens) - Slim: https://www.anchorterminal.com/categories/text-to-speech.min.md (~580 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/categories/text-to-speech.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-04 APIs that turn text into speech, streamed for live conversation or rendered for long-form audio. Compared on time to first audio, pronunciation, how natural it sounds, long-form consistency, languages and cost for a fixed script. - Tools ranked: 10 · agent-ready (BB or better): 5 · accept x402: 0 · hosted endpoints: 10 · desk reviews by the panel: 26 - JSON: https://www.anchorterminal.com/api/v1/tools.json (list) · https://www.anchorterminal.com/api/v1/rankings.json (ranked) · https://www.anchorterminal.com/api/v1/x402.json (payable) · https://www.anchorterminal.com/api/v1/capabilities.json (by capability) - Grades run AA, A, BB, B, C, D, E, F · methodology: https://www.anchorterminal.com/benchmark/ - Capabilities in this category: speech.tts, speech.streaming, speech.voices, speech.ssml, speech.languages - https://letme.dev/speech.tts picks the top-graded tool in this list and says how to call it direct; calling through letme comes later (https://www.anchorterminal.com/letme/index.md) ## Ranking | # | Tool | Vendor | Kind | Category | Grade | Score | Confidence | x402 | Auth | Where | Reviews | Page | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 29 | Amazon Polly | Amazon Web Services | Model API | TTS | BB | 75.8 | high | no | API key | hosted | 4/5 (8) | https://www.anchorterminal.com/tools/amazon-polly.md | | 56 | Azure AI Speech text-to-speech | Microsoft Azure | Model API | TTS | BB | 73.7 | medium | no | OAuth or key | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/azure-text-to-speech.md | | 61 | ElevenLabs Text to Speech API + MCP | ElevenLabs | Model API | TTS | BB | 73.1 | high | no | OAuth or key | hosted + local | 3.5/5 (2) | https://www.anchorterminal.com/tools/elevenlabs-tts.md | | 63 | Deepgram Text-to-Speech (Aura-2, Flux TTS) | Deepgram | Model API | TTS | BB | 73 | high | no | API key | hosted + local | 3.5/5 (2) | https://www.anchorterminal.com/tools/deepgram-tts.md | | 91 | Murf TTS API + MCP | Murf | Model API | TTS | BB | 70.9 | medium | no | API key | hosted + local | 3.5/5 (2) | https://www.anchorterminal.com/tools/murf-tts.md | | 187 | Cartesia Sonic TTS API + MCP | Cartesia | Model API | TTS | B | 64.2 | medium | no | OAuth or key | hosted + local | 3/5 (2) | https://www.anchorterminal.com/tools/cartesia-tts.md | | 193 | Soniox Text-to-Speech | Soniox | Model API | TTS | B | 63.9 | medium | no | API key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/soniox-tts.md | | 305 | Rime TTS API + MCP | Rime Labs | Model API | TTS | C | 56.1 | medium | no | API key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/rime-tts.md | | 357 | Resemble AI Text-to-Speech API | Resemble AI | Model API | TTS | D | 50.6 | medium | no | API key | hosted | 2/5 (2) | https://www.anchorterminal.com/tools/resemble-ai-tts.md | | not ranked, shut down | PlayHT Text-to-Speech API | PlayHT | Model API | TTS | F | 4.2 | high | no | API key | hosted | 1/5 (2) | https://www.anchorterminal.com/tools/playht-tts.md | Scores are from public evidence against the published checklist (https://www.anchorterminal.com/benchmark/), with Performance and Task success pending. p95 latency and context cost come from our probes, which haven't run yet. ## Summaries ### 29. Amazon Polly, BB (75.8) AWS's speech synthesis API with four engines (standard, neural, long-form and generative) and about 110 voices in 42 languages and variants. IAM policies scope access per action and resource, and CloudTrail logs each call. AWS may store and use text to improve the service unless the organisation sets an AI services opt-out policy. - Page: https://www.anchorterminal.com/tools/amazon-polly · Markdown: https://www.anchorterminal.com/tools/amazon-polly.md · JSON: https://www.anchorterminal.com/api/v1/tools/amazon-polly.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.ssml, speech.languages · endpoint: `https://polly.us-east-1.amazonaws.com/v1` ### 56. Azure AI Speech text-to-speech, BB (73.7) Azure's text-to-speech service for generating spoken audio. Real-time synthesis keeps neither the input text nor the output audio. An Azure subscription needs a card, even for the free F0 tier. - Page: https://www.anchorterminal.com/tools/azure-text-to-speech · Markdown: https://www.anchorterminal.com/tools/azure-text-to-speech.md · JSON: https://www.anchorterminal.com/api/v1/tools/azure-text-to-speech.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.ssml, speech.languages · endpoint: `https://eastus.tts.speech.microsoft.com/cognitiveservices` ### 61. ElevenLabs Text to Speech API + MCP, BB (73.1) ElevenLabs text-to-speech over REST, HTTP streaming and WebSocket. Keys can be limited to chosen endpoints and given a credit quota, and service accounts hold keys that don't belong to a person. Content may be used for training unless you opt out under Data use, and the opt-out only applies going forward. - Page: https://www.anchorterminal.com/tools/elevenlabs-tts · Markdown: https://www.anchorterminal.com/tools/elevenlabs-tts.md · JSON: https://www.anchorterminal.com/api/v1/tools/elevenlabs-tts.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.ssml, speech.languages · endpoint: `https://api.elevenlabs.io/v1` ### 63. Deepgram Text-to-Speech (Aura-2, Flux TTS), BB (73) Deepgram's text-to-speech API for generating spoken audio. OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages. Requests can be kept for training unless each one sets `mip_opt_out=true`. - Page: https://www.anchorterminal.com/tools/deepgram-tts · Markdown: https://www.anchorterminal.com/tools/deepgram-tts.md · JSON: https://www.anchorterminal.com/api/v1/tools/deepgram-tts.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.languages · endpoint: `https://api.deepgram.com/v1` ### 91. Murf TTS API + MCP, BB (70.9) Murf's API for generating speech from text. Falcon 2 costs 1 cent per 1,000 characters on pay as you go, with a $2 minimum purchase. Falcon 2 concurrency is 2 outside US-East on free and pay as you go, 5 on US-East. - Page: https://www.anchorterminal.com/tools/murf-tts · Markdown: https://www.anchorterminal.com/tools/murf-tts.md · JSON: https://www.anchorterminal.com/api/v1/tools/murf-tts.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.languages · endpoint: `https://api.murf.ai/v1` ### 187. Cartesia Sonic TTS API + MCP, B (64.2) Sonic text-to-speech, currently `sonic-3.6` in 44 languages, over `/tts/bytes`, `/tts/sse` and a WebSocket that takes streamed LLM text with contexts and word timestamps. WebSocket input accepts streamed LLM text with contexts, flushing and word timestamps. Five TTS incidents on the status page between 29 July and 21 August 2026. - Page: https://www.anchorterminal.com/tools/cartesia-tts · Markdown: https://www.anchorterminal.com/tools/cartesia-tts.md · JSON: https://www.anchorterminal.com/api/v1/tools/cartesia-tts.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.ssml, speech.languages · endpoint: `https://api.cartesia.ai` ### 193. Soniox Text-to-Speech, B (63.9) Streaming and REST text-to-speech (`tts-rt-v2`) in 60+ languages, where every voice speaks every language. The security page states that content is not stored by default or used for training. Audio output is capped at two minutes per request or stream. - Page: https://www.anchorterminal.com/tools/soniox-tts · Markdown: https://www.anchorterminal.com/tools/soniox-tts.md · JSON: https://www.anchorterminal.com/api/v1/tools/soniox-tts.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.languages · endpoint: `https://tts-rt.soniox.com` ### 305. Rime TTS API + MCP, C (56.1) Rime's streaming text-to-speech for voice agents. Only character counts are kept by default, and customer data isn't used for training without an opt-in. No OpenAPI or AsyncAPI file and no official SDKs. - Page: https://www.anchorterminal.com/tools/rime-tts · Markdown: https://www.anchorterminal.com/tools/rime-tts.md · JSON: https://www.anchorterminal.com/api/v1/tools/rime-tts.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.languages · endpoint: `https://users.rime.ai` ### 357. Resemble AI Text-to-Speech API, D (50.6) Resemble's TTS API on its current Resemble Ultra model, which the changelog says is powered by xAI. OpenAPI file in JSON and YAML, llms.txt and Markdown pages. Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published. - Page: https://www.anchorterminal.com/tools/resemble-ai-tts · Markdown: https://www.anchorterminal.com/tools/resemble-ai-tts.md · JSON: https://www.anchorterminal.com/api/v1/tools/resemble-ai-tts.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.ssml · endpoint: `https://app.resemble.ai/api/v2` ### PlayHT Text-to-Speech API, F (4.2), retired, not ranked Discontinued text-to-speech API from PlayHT, formerly supporting streaming audio and batch generation. The old API reference is still readable at docs.play.ht for anyone porting code. api.play.ht doesn't resolve, so every call fails. - Page: https://www.anchorterminal.com/tools/playht-tts · Markdown: https://www.anchorterminal.com/tools/playht-tts.md · JSON: https://www.anchorterminal.com/api/v1/tools/playht-tts.json - Capabilities: speech.tts, speech.streaming, speech.voices, speech.languages · endpoint: `https://api.play.ht/api/v2` ## How we test this category One fixed script through every API. We measure time to first audio, check pronunciation of names, numbers and acronyms, judge naturalness blind, listen for drift across a long passage, and price the whole script. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence. ## Indexed, not reviewed (4) Sorted into this category from public catalogues, with facts and our own checks but no score, grade or rank (https://www.anchorterminal.com/indexed/index.md). | Listing | Kind | What it does | Why it's here | | --- | --- | --- | --- | | [APICK](https://www.anchorterminal.com/tools/apick-all.md) | MCP server | APICK Korean data, simple-auth lookups, OCR, search, conversion, image, video and TTS | vendor's own | | [BananaBanana Image, Video & Speech Generation](https://www.anchorterminal.com/tools/bananabanana-image-video.md) | MCP server | Images, video & speech: Nano Banana, GPT Image, Veo, Omni, Wan, Grok, Gemini TTS. Crypto or card. | vendor's own | | [StudioMeyer Video](https://www.anchorterminal.com/tools/studiomeyer-video.md) | MCP server | Video production: recording, editing, effects, captions, TTS, screenshots. 8 tools. | vendor's own | | [UnlimitedTTS](https://www.anchorterminal.com/tools/unlimitedtts-tts.md) | MCP server | Quote-first, non-custodial x402 text-to-speech with spend policy, MP3 artifacts, and receipts. | vendor's own |