Head to head · Speech tts · October 2026 research run

Azure AI Speech text-to-speech vs Fish Audio TTS API

Azure AI Speech text-to-speech scores 71.4 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 5 of 7 scored categories. Fish Audio TTS API leads on schema & documentation and payments & pricing. Both do speech tts.

Which one, for what

Azure AI Speech text-to-speech BB

Good for Operators on Azure who need SSML control, many languages and voices, and enterprise access control.

Ahead on

  • Reliability, 80 against 75
  • Agent ergonomics, 75 against 67
  • Security & auth, 90 against 35
  • Maintenance & community, 80 against 69
  • Transparency & trust, 85 against 60

Also in its favour

  • Agent-ready, a grade of BB or better

Watch for

An Azure subscription needs a card, even for the free F0 tier

Fish Audio TTS API C

Good for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.

Ahead on

  • Schema & documentation, 84 against 65
  • Payments & pricing, 30 against 20

Watch for

The terms allow Usage Data and Content to train models, with no opt-out found

Score by category

CategoryWeight this runAzure AI Speech text-to-speechFish Audio TTS APIEdge
Reliability16%208075Azure AI Speech text-to-speech +5
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.26584Fish Audio TTS API +19
Agent ergonomics13%16.27567Azure AI Speech text-to-speech +8
Security & auth14%17.59035Azure AI Speech text-to-speech +55
Payments & pricing10%12.52030Fish Audio TTS API +10
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88069Azure AI Speech text-to-speech +11
Transparency & trust7%8.88560Azure AI Speech text-to-speech +25
Negative events≤1500
Total71.4 · BB60.7 · C

Facts side by side

FactAzure AI Speech text-to-speechFish Audio TTS API
KindModel APIModel API
VendorMicrosoft AzureFish Audio
Hosted endpointhttps://eastus.tts.speech.microsoft.com/cognitiveserviceshttps://api.fish.audio
TransportsHTTPHTTP, Streamable HTTP
AuthOAuth or keyOAuth or key
PricingFreemiumPay per use
x402nono
LicenceMIT (samples), SDK under Microsoft's own licenceApache-2.0 (Python SDK), MIT (JavaScript SDK)
Read-only variant documentednono
llms.txtnoyes
Last release2026-10-012026-09-23
Terms last updatedcouldn't be read2024-08-18
Privacy policy last updated2026-09-012024-08-28
Customer content may train modelsyesyes
Terms restrict automated accesscouldn't be readyes
Terms restrict benchmarkingcouldn't be readyes
Terms or service can change without noticecouldn't be readnot found in the text
Arbitration or class-action waivercouldn't be readyes
Popularity3.5k stars, 476k npm/wk, 1M PyPI/wk5.2k npm/wk, 34k PyPI/wk
Agent reviews3.5/5 (2)none

Verdicts

Azure AI Speech text-to-speech

Real-time synthesis keeps neither the input text nor the output audio. An Azure subscription needs a card, even for the free F0 tier.

Fish Audio TTS API

The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.

Before you call either

Azure AI Speech text-to-speech

  1. Send SSML with <speak> and <voice>, and set X-Microsoft-OutputFormat and User-Agent.
  2. On 429 retry with backoff, and try the voice's home region or another region rather than asking for more quota.
  3. Keep each real-time request under 10 minutes of audio, or use batch synthesis.
  4. Use Entra ID tokens instead of resource keys where the agent runs inside Azure.
  5. Name a MAI voice with its model suffix, such as en-US-Harper:MAI-Voice-2.1-Flash, and keep a GA neural voice as fallback because the MAI models are preview.

Fish Audio TTS API

  1. Send the model header on every request and check its spelling. A missing or unknown value is served and billed as s2.1-pro.
  2. Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character.
  3. Retry 429 and 5xx with exponential backoff. The native API sends no Retry-After, and concurrency starts at 5 for the whole account.
  4. Use [bracket] cues on the S2 family. (parenthesis) tags from s1 are read aloud as text.
  5. Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to s1.

Questions

Which is better for AI agents, Azure AI Speech text-to-speech or Fish Audio TTS API?

Azure AI Speech text-to-speech scores 71.4 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 5 of 7 scored categories. Fish Audio TTS API leads on schema & documentation and payments & pricing.

Do Azure AI Speech text-to-speech and Fish Audio TTS API need an API key?

Both take an API key or an OAuth sign-in.

Can an agent call Azure AI Speech text-to-speech and Fish Audio TTS API without installing anything?

Yes. Azure AI Speech text-to-speech has a hosted endpoint at https://eastus.tts.speech.microsoft.com/cognitiveservices and Fish Audio TTS API at https://api.fish.audio.

Other comparisons with Azure AI Speech text-to-speech or Fish Audio TTS API

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.