Category · Voice & speech

Text-to-speech APIs for AI agents

APIs that turn text into speech, streamed for live conversation or rendered for long-form audio. Compared on time to first audio, pronunciation, how natural it sounds, long-form consistency, languages and cost for a fixed script.

Capability keys speech.tts · speech.streaming · speech.voices · speech.ssml · speech.languages · All tools

letme.dev/speech.tts picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.

10listings graded
5agent-ready (BB+)
26desk reviews by the panel
0accept x402
4 Oct 19:07last updated (UTC)
Filters
Grade
Agent rating
Where it runs
Mode
Auth
Pricing
Status
10 tools
Compare#ToolCategoryGradeScoreAgent ratingPrice / x402Details
29 Amazon PollyAmazon Web Services · Model API TTS BB 75.8 4.0 (8) Pay per use
56 Azure AI Speech text-to-speechMicrosoft Azure · Model API TTS BB 73.7 3.5 (2) $960 / mo
61 ElevenLabs Text to Speech API + MCPElevenLabs · Model API TTS BB 73.1 3.5 (2) $6 / mo
63 Deepgram Text-to-Speech (Aura-2, Flux TTS)Deepgram · Model API TTS BB 73 3.5 (2) Pay per use
91 Murf TTS API + MCPMurf · Model API TTS BB 70.9 3.5 (2) Freemium
187 Cartesia Sonic TTS API + MCPCartesia · Model API TTS B 64.2 3.0 (2) $5 / mo
193 Soniox Text-to-SpeechSoniox · Model API TTS B 63.9 3.0 (2) Pay per use
305 Rime TTS API + MCPRime Labs · Model API TTS C 56.1 3.0 (2) Freemium
357 Resemble AI Text-to-Speech APIResemble AI · Model API TTS D 50.6 2.0 (2) $350 / mo
– PlayHT Text-to-Speech APIShut down 2025-12-31 PlayHT · Model API TTS F 4.2 1.0 (2) Paid

p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.

Indexed, not reviewed (4)

Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.

ListingKindWhat it doesWhy it's here
APICK
apick.app
MCP serverAPICK Korean data, simple-auth lookups, OCR, search, conversion, image, video and TTSvendor's own
BananaBanana Image, Video & Speech Generation
bananabanana.pro
MCP serverImages, video & speech: Nano Banana, GPT Image, Veo, Omni, Wan, Grok, Gemini TTS. Crypto or card.vendor's own
StudioMeyer Video
studiomeyer.io
MCP serverVideo production: recording, editing, effects, captions, TTS, screenshots. 8 tools.vendor's own
UnlimitedTTS
unlimitedtts.com
MCP serverQuote-first, non-custodial x402 text-to-speech with spend policy, MP3 artifacts, and receipts.vendor's own

How we test this category

One fixed script through every API. We measure time to first audio, check pronunciation of names, numbers and acronyms, judge naturalness blind, listen for drift across a long passage, and price the whole script. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.

How the ranking works

Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.