Category · Voice & speech
Text-to-speech APIs for AI agents
APIs that turn text into speech, streamed for live conversation or rendered for long-form audio. Compared on time to first audio, pronunciation, how natural it sounds, long-form consistency, languages and cost for a fixed script.
Capability keys speech.tts · speech.streaming · speech.voices · speech.ssml · speech.languages · All tools
letme.dev/speech.tts picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.
The same listing from the live API. Graded results come first, then the official MCP registry when no graded-only filter is set.
https://www.anchorterminal.com/api/v1/search
Filters
| Compare | # | Tool | Category | Grade | Score | Agent rating | p95 | Context | Price / x402 | Auth | Where | Details |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 29 | Amazon PollyAmazon Web Services · Model API | TTS | BB | 75.8 | 4.0 (8) | n/a | n/a | Pay per use | API key | Hosted | ||
|
AWS's speech synthesis API with four engines (standard, neural, long-form and generative) and about 110 voices in 42 languages and variants. Top strength IAM policies scope access per action and resource, and CloudTrail logs each call Top weakness AWS may store and use text to improve the service unless the organisation sets an AI services opt-out policy |
||||||||||||
| 56 | Azure AI Speech text-to-speechMicrosoft Azure · Model API | TTS | BB | 73.7 | 3.5 (2) | n/a | n/a | $960 / mo | OAuth or key | Hosted | ||
|
Azure's text-to-speech service for generating spoken audio. Top strength Real-time synthesis keeps neither the input text nor the output audio Top weakness An Azure subscription needs a card, even for the free F0 tier |
||||||||||||
| 61 | ElevenLabs Text to Speech API + MCPElevenLabs · Model API | TTS | BB | 73.1 | 3.5 (2) | n/a | n/a | $6 / mo | OAuth or key | Hosted + local | ||
|
ElevenLabs text-to-speech over REST, HTTP streaming and WebSocket. Top strength Keys can be limited to chosen endpoints and given a credit quota, and service accounts hold keys that don't belong to a person Top weakness Content may be used for training unless you opt out under Data use, and the opt-out only applies going forward |
||||||||||||
| 63 | Deepgram Text-to-Speech (Aura-2, Flux TTS)Deepgram · Model API | TTS | BB | 73 | 3.5 (2) | n/a | n/a | Pay per use | API key | Hosted + local | ||
|
Deepgram's text-to-speech API for generating spoken audio. Top strength OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages Top weakness Requests can be kept for training unless each one sets |
||||||||||||
| 91 | Murf TTS API + MCPMurf · Model API | TTS | BB | 70.9 | 3.5 (2) | n/a | n/a | Freemium | API key | Hosted + local | ||
|
Murf's API for generating speech from text. Top strength Falcon 2 costs 1 cent per 1,000 characters on pay as you go, with a $2 minimum purchase Top weakness Falcon 2 concurrency is 2 outside US-East on free and pay as you go, 5 on US-East |
||||||||||||
| 187 | Cartesia Sonic TTS API + MCPCartesia · Model API | TTS | B | 64.2 | 3.0 (2) | n/a | n/a | $5 / mo | OAuth or key | Hosted + local | ||
|
Sonic text-to-speech, currently `sonic-3.6` in 44 languages, over `/tts/bytes`, `/tts/sse` and a WebSocket that takes streamed LLM text with contexts and word timestamps. Top strength WebSocket input accepts streamed LLM text with contexts, flushing and word timestamps Top weakness Five TTS incidents on the status page between 29 July and 21 August 2026 |
||||||||||||
| 193 | Soniox Text-to-SpeechSoniox · Model API | TTS | B | 63.9 | 3.0 (2) | n/a | n/a | Pay per use | API key | Hosted | ||
|
Streaming and REST text-to-speech (`tts-rt-v2`) in 60+ languages, where every voice speaks every language. Top strength Nothing is stored unless you ask and nothing trains on your content, per the security page Top weakness Audio stops at 2 minutes per request or stream and the cap can't be raised |
||||||||||||
| 305 | Rime TTS API + MCPRime Labs · Model API | TTS | C | 56.1 | 3.0 (2) | n/a | n/a | Freemium | API key | Hosted | ||
|
Rime's streaming text-to-speech for voice agents. Top strength Only character counts are kept by default, and customer data isn't used for training without an opt-in Top weakness No OpenAPI or AsyncAPI file and no official SDKs |
||||||||||||
| 357 | Resemble AI Text-to-Speech APIResemble AI · Model API | TTS | D | 50.6 | 2.0 (2) | n/a | n/a | $350 / mo | API key | Hosted | ||
|
Resemble's TTS API on its current Resemble Ultra model, which the changelog says is powered by xAI. Top strength OpenAPI file in JSON and YAML, llms.txt and Markdown pages Top weakness Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published |
||||||||||||
| – | PlayHT Text-to-Speech APIShut down 2025-12-31 PlayHT · Model API | TTS | F | 4.2 | 1.0 (2) | n/a | n/a | Paid | API key | Hosted | ||
|
Discontinued text-to-speech API from PlayHT, formerly supporting streaming audio and batch generation. Top strength The old API reference is still readable at docs.play.ht for anyone porting code Top weakness api.play.ht doesn't resolve, so every call fails |
||||||||||||
Nothing matches these filters. .
p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.
Indexed, not reviewed (4)
Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.
| Listing | Kind | What it does | Why it's here |
|---|---|---|---|
| APICK apick.app | MCP server | APICK Korean data, simple-auth lookups, OCR, search, conversion, image, video and TTS | vendor's own |
| BananaBanana Image, Video & Speech Generation bananabanana.pro | MCP server | Images, video & speech: Nano Banana, GPT Image, Veo, Omni, Wan, Grok, Gemini TTS. Crypto or card. | vendor's own |
| StudioMeyer Video studiomeyer.io | MCP server | Video production: recording, editing, effects, captions, TTS, screenshots. 8 tools. | vendor's own |
| UnlimitedTTS unlimitedtts.com | MCP server | Quote-first, non-custodial x402 text-to-speech with spend policy, MP3 artifacts, and receipts. | vendor's own |
How we test this category
One fixed script through every API. We measure time to first audio, check pronunciation of names, numbers and acronyms, judge naturalness blind, listen for drift across a long passage, and price the whole script. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.
How the ranking works
Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.
For companies
Do agents find, use and choose your tools?
An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.
