Best of · Voice & speech

Best text-to-speech APIs for AI agents

All 10 ranked text-to-speech APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.

  • 10 ranked
  • 5 agent-ready
  • 10 hosted endpoints
  • Updated 9 October 2026

Top three

Picks by need

Worked out from the scores, prices and facts, so they change when the research does.

Highest score overall

Amazon Polly BB

BB, 75.6/100 on the benchmark.

Also ElevenLabs Text to Speech API + MCP, BB, 75.2/100.

Schema & documentation

Deepgram Text-to-Speech (Aura-2, Flux TTS) BB

95/100 on schema & documentation, against 90 for the overall leader.

Agent ergonomics

ElevenLabs Text to Speech API + MCP BB

92/100 on agent ergonomics, against 82 for the overall leader.

Security & auth

Azure AI Speech text-to-speech BB

90/100 on security & auth, against 80 for the overall leader.

Maintenance & community

Azure AI Speech text-to-speech BB

80/100 on maintenance & community, against 50 for the overall leader.

Transparency & trust

Azure AI Speech text-to-speech BB

85/100 on transparency & trust, against 77 for the overall leader.

A hosted MCP endpoint

ElevenLabs Text to Speech API + MCP BB

remote MCP server, nothing to install.

The shortlist

#ToolGradeBest forPriceWhere
1 Amazon Polly
Amazon Web Services
BB 75.6 Operators already on AWS who want predictable, cheap speech for prompts, notifications and IVR, or generative voices with bidirectional streaming. Pay per use hosted
2 ElevenLabs Text to Speech API + MCP
ElevenLabs
BB 75.2 Best when an agent needs many languages, many voices or the most expressive models from one vendor, and the operator can accept training by default or pay for Enterprise. $6 / mo hosted and local
3 Deepgram Text-to-Speech (Aura-2, Flux TTS)
Deepgram
BB 72.7 English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs. Pay per use hosted and local
4 Azure AI Speech text-to-speech
Microsoft Azure
BB 71.4 Operators on Azure who need SSML control, many languages and voices, and enterprise access control. $960 / mo hosted
5 Murf TTS API + MCP
Murf
BB 70.6 Budget voice-agent output where 2 to 5 concurrent streams are enough, or for studio-style renders on Gen2. Freemium hosted and local
6 Cartesia Sonic TTS API + MCP
Cartesia
B 63.8 Real-time voice agents that stream LLM output straight into speech and need stable, pinnable models. $5 / mo hosted and local
7 Soniox Text-to-Speech
Soniox
B 63.7 Multilingual agents that need any voice in any of 60-plus languages, short turns and strict data handling. Pay per use hosted
8 Fish Audio TTS API
Fish Audio
C 60.7 Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes. Pay per use hosted
9 Rime TTS API + MCP
Rime Labs
C 55.8 English and a handful of other languages in live voice agents where the operator cares most about data terms and price. Freemium hosted
10 Resemble AI Text-to-Speech API
Resemble AI
D 50.3 Operators who already use Resemble for detection or watermarking and want one vendor for both. $350 / mo hosted

How to choose

  1. Time to first audioCheck the time to first audio when streaming, because a voice agent that waits for the full clip before speaking leaves the caller in silence.
  2. Pronunciation and SSML controlCheck whether SSML or a pronunciation dictionary can fix names, numbers and acronyms, because a misread account code or product name is hard for a listener to correct.
  3. Voice consistency over lengthCheck whether the voice keeps its tone and pace across a long passage, since drift over a long narration can make it sound like separate speakers.
  4. Cost of a fixed scriptCheck the price per character or per minute of output, and whether custom voices carry an extra fee, because the bill grows with every word.

How the benchmark tests this category. One fixed script through every API. We measure time to first audio, check pronunciation of names, numbers and acronyms, judge naturalness blind, listen for drift across a long passage, and price the whole script.

Each one in detail

#1

Amazon Polly

BB 75.6/100

AWS's speech synthesis API with four engines (standard, neural, long-form and generative) and about 110 voices in 42 languages and variants.

Verdict IAM policies scope access per action and resource, and CloudTrail logs each call. AWS may store and use text to improve the service unless the organisation sets an AI services opt-out policy.

Choose it for Operators already on AWS who want predictable, cheap speech for prompts, notifications and IVR, or generative voices with bidirectional streaming.

Strengths

  • IAM policies scope access per action and resource, and CloudTrail logs each call
  • Quotas published per operation and engine, with backoff and jitter guidance for throttling
  • Covered by the Amazon Machine Learning Language SLA

Weaknesses

  • AWS may store and use text to improve the service unless the organisation sets an AI services opt-out policy
  • A new account needs a card, and the monthly free characters only apply to accounts opened before 2025-07-15
  • Neural, long-form and generative synthesis is limited to 8 requests a second by default

Price Pay per useAuth API keyx402 nohosted

Full assessment

#2

ElevenLabs Text to Speech API + MCP

BB 75.2/100

ElevenLabs text-to-speech over REST, HTTP streaming and WebSocket.

Verdict Keys can be limited to chosen endpoints and given a credit quota, and service accounts hold keys that don't belong to a person. The Terms let ElevenLabs train on content unless you opt out, which the Eleven v4 launch page contradicts, and v4 runs through Text to Dialogue rather than the classic text-to-speech endpoint.

Choose it for Best when an agent needs many languages, many voices or the most expressive models from one vendor, and the operator can accept training by default or pay for Enterprise.

Strengths

  • Keys can be limited to chosen endpoints and given a credit quota, and service accounts hold keys that don't belong to a person
  • Error reference with about 90 codes, each carrying type, code, message and request_id
  • OpenAPI file, llms.txt and Markdown twins of every docs page

Weaknesses

  • Content may be used for training unless you opt out under Data use, though the Eleven v4 launch page says it isn't used without consent
  • Zero retention (enable_logging=false) is Enterprise only, and TTS text and audio are kept by default
  • No SLA published for self-serve plans

Price $6 / moAuth OAuth or keyx402 nohosted and local

Full assessment · Against #1, Amazon Polly

#3

Deepgram Text-to-Speech (Aura-2, Flux TTS)

BB 72.7/100

Deepgram's text-to-speech API for generating spoken audio.

Verdict OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages. Requests can be kept for training unless each one sets mip_opt_out=true.

Choose it for English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs.

Strengths

  • OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages
  • Flux TTS's Interrupt event returns text_spoken and text_remaining on barge-in
  • Keys carry roles and scopes, and browser tokens live 30 seconds

Weaknesses

  • Requests can be kept for training unless each one sets mip_opt_out=true
  • Flux TTS is English only and capped at 5 concurrent streams in the EU, Australia and India below Enterprise
  • No SSML, and Flux TTS strips other vendors' tags with a warning

Price Pay per useAuth API keyx402 nohosted and local

Full assessment · Against #1, Amazon Polly

#4

Azure AI Speech text-to-speech

BB 71.4/100

Azure's text-to-speech service for generating spoken audio.

Verdict Real-time synthesis keeps neither the input text nor the output audio. An Azure subscription needs a card, even for the free F0 tier.

Choose it for Operators on Azure who need SSML control, many languages and voices, and enterprise access control.

Strengths

  • Real-time synthesis keeps neither the input text nor the output audio
  • Full SSML with speaking styles, prosody, phonemes, lexicons and up to 50 voice or audio tags a request
  • Microsoft Entra ID with role-based access, or two rotatable keys

Weaknesses

  • An Azure subscription needs a card, even for the free F0 tier
  • No llms.txt and no OpenAPI file for text-to-speech found
  • 429s often reflect busy capacity for a voice in a region, which a quota increase doesn't fix

Price $960 / moAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Amazon Polly

#5

Murf TTS API + MCP

BB 70.6/100

Murf's API for generating speech from text.

Verdict Falcon 2 costs 1 cent per 1,000 characters on pay as you go, with a $2 minimum purchase. Falcon 2 concurrency is 2 outside US-East on free and pay as you go, 5 on US-East.

Choose it for Budget voice-agent output where 2 to 5 concurrent streams are enough, or for studio-style renders on Gen2.

Strengths

  • Falcon 2 costs 1 cent per 1,000 characters on pay as you go, with a $2 minimum purchase
  • OpenAPI and AsyncAPI specs, llms.txt and Markdown twins of every docs page
  • Errors page lists 7 HTTP codes, 4 WebSocket errors and 8 warnings, each with cause and fix, and tells you to back off on 429

Weaknesses

  • Falcon 2 concurrency is 2 outside US-East on free and pay as you go, 5 on US-East
  • Gen2 streaming was deprecated on 2026-08-16 with 27 days' notice
  • The security page says all data stays in AWS us-east-2, while Falcon 2 runs in 12 regions and the subprocessor list names four other clouds

Price FreemiumAuth API keyx402 nohosted and local

Full assessment · Against #1, Amazon Polly

#6

Cartesia Sonic TTS API + MCP

B 63.8/100

Sonic text-to-speech, currently sonic-3.6 in 44 languages, over /tts/bytes, /tts/sse and a WebSocket that takes streamed LLM text with contexts and word timestamps.

Verdict WebSocket input accepts streamed LLM text with contexts, flushing and word timestamps. Five TTS incidents on the status page between 29 July and 21 August 2026.

Choose it for Real-time voice agents that stream LLM output straight into speech and need stable, pinnable models.

Strengths

  • WebSocket input accepts streamed LLM text with contexts, flushing and word timestamps
  • Dated model snapshots such as sonic-3.6-2026-08-27 never change, and Cartesia-Version pins the API shape
  • OpenAPI and AsyncAPI files per API version listed in llms.txt

Weaknesses

  • Five TTS incidents on the status page between 29 July and 21 August 2026
  • TTS concurrency is 2 on Free, 3 on Pro, 5 on Startup and 15 on Scale
  • The Terms allow training on inputs and outputs unless you opt out by form

Price $5 / moAuth OAuth or keyx402 nohosted and local

Full assessment · Against #1, Amazon Polly

#7

Soniox Text-to-Speech

B 63.7/100

Streaming and REST text-to-speech (tts-rt-v2) in 60+ languages, where every voice speaks every language.

Verdict The security page states that content is not stored by default or used for training. Audio output is capped at two minutes per request or stream.

Choose it for Multilingual agents that need any voice in any of 60-plus languages, short turns and strict data handling.

Strengths

  • Nothing is stored unless you ask and nothing trains on your content, per the security page
  • One error format with a stable error_type, request_id and more_info link
  • Separate TTS REST and TTS Real-time status components in four regions, 100 per cent uptime shown over 90 days

Weaknesses

  • Audio stops at 2 minutes per request or stream and the cap can't be raised
  • 3 concurrent requests and 100 requests a minute by default
  • tts-rt-v1 was removed 20 days after its deprecation notice

Price Pay per useAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Polly

#8

Fish Audio TTS API

C 60.7/100

Fish Audio's API turns text into speech with the s2.1-pro model in 83 languages, over a REST endpoint, a timestamped stream and a WebSocket that accepts text as it is produced.

Verdict The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.

Choose it for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.

Strengths

  • Public OpenAPI 3.1 file, llms.txt, Markdown pages and two installable agent skills for the SDKs and the raw API
  • WebSocket input takes text as it is produced, with flush and stop events and a variant that returns word timestamps
  • s2.1-pro-free runs the production model at $0 under fair-use limits, through 30 November 2026

Weaknesses

  • The terms allow Usage Data and Content to train models, with no opt-out found
  • An unrecognised model header falls back to paid s2.1-pro without an error
  • The Text-to-Speech API status component shows downtime on 11 days in 90, the longest 58 minutes on 6 August 2026, mostly on the free model

Price Pay per useAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Amazon Polly

#9

Rime TTS API + MCP

C 55.8/100

Rime's streaming text-to-speech for voice agents.

Verdict Only character counts are kept by default, and customer data isn't used for training without an opt-in. No OpenAPI or AsyncAPI file and no official SDKs.

Choose it for English and a handful of other languages in live voice agents where the operator cares most about data terms and price.

Strengths

  • Only character counts are kept by default, and customer data isn't used for training without an opt-in
  • Coda at $0.05 and Mist v3 at $0.03 per 1,000 characters, with free minutes and no card
  • 20 concurrent generations on Starter

Weaknesses

  • No OpenAPI or AsyncAPI file and no official SDKs
  • Errors are plain-text messages without machine-readable codes
  • Arcana's cloud retirement came 17 days after the announcement

Price FreemiumAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Polly

#10

Resemble AI Text-to-Speech API

D 50.3/100

Resemble's TTS API on its current Resemble Ultra model, which the changelog says is powered by xAI.

Verdict OpenAPI file in JSON and YAML, llms.txt and Markdown pages. Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published.

Choose it for Operators who already use Resemble for detection or watermarking and want one vendor for both.

Strengths

  • OpenAPI file in JSON and YAML, llms.txt and Markdown pages
  • SSML with prompt, temperature and exaggeration controls and inline tags such as [laugh]
  • Status page checks Ultra HTTP synthesis directly, 100 per cent over 90 days

Weaknesses

  • Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published
  • WebSocket streaming only on Business at $1,000 a month
  • Errors are success: false and a message, no status codes listed

Price $350 / moAuth API keyx402 nohosted

Full assessment · Against #1, Amazon Polly

Head to head

All 55 comparisons in this category

Questions

What are the highest-rated text-to-speech APIs for AI agents?

Amazon Polly has the highest benchmark score of the 10 ranked text-to-speech APIs, 75.6 (BB). ElevenLabs Text to Speech API + MCP is second with 75.2 (BB).

How many text-to-speech APIs are agent-ready?

5 of the 10 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.

Which text-to-speech APIs accept x402 payments?

None of the ranked listings here accepts x402 for its main call yet.

How is this list ranked?

By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.

How this list is made

The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.

Full ranked table · 55 head-to-head comparisons · Best tools in every category

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.