Head to head · Speech tts · October 2026 research run

Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Fish Audio TTS API

Deepgram Text-to-Speech (Aura-2, Flux TTS) scores 72.7 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 6 of 7 scored categories. Fish Audio TTS API leads on reliability. Both do speech tts.

Which one, for what

Deepgram Text-to-Speech (Aura-2, Flux TTS) BB

Good for English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs.

Ahead on

  • Schema & documentation, 95 against 84
  • Agent ergonomics, 82 against 67
  • Security & auth, 70 against 35
  • Payments & pricing, 40 against 30
  • Transparency & trust, 72 against 60

Also in its favour

  • Agent-ready, a grade of BB or better
  • Runs on your own machine
  • Free to start without a card

Watch for

Requests can be kept for training unless each one sets mip_opt_out=true

Fish Audio TTS API C

Good for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.

Ahead on

  • Reliability, 75 against 70

Watch for

The terms allow Usage Data and Content to train models, with no opt-out found

Score by category

CategoryWeight this runDeepgram Text-to-Speech (Aura-2, Flux TTS)Fish Audio TTS APIEdge
Reliability16%207075Fish Audio TTS API +5
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29584Deepgram Text-to-Speech (Aura-2, Flux TTS) +11
Agent ergonomics13%16.28267Deepgram Text-to-Speech (Aura-2, Flux TTS) +15
Security & auth14%17.57035Deepgram Text-to-Speech (Aura-2, Flux TTS) +35
Payments & pricing10%12.54030Deepgram Text-to-Speech (Aura-2, Flux TTS) +10
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.87369Deepgram Text-to-Speech (Aura-2, Flux TTS) +4
Transparency & trust7%8.87260Deepgram Text-to-Speech (Aura-2, Flux TTS) +12
Negative events≤1500
Total72.7 · BB60.7 · C

Facts side by side

FactDeepgram Text-to-Speech (Aura-2, Flux TTS)Fish Audio TTS API
KindModel APIModel API
VendorDeepgramFish Audio
Hosted endpointhttps://api.deepgram.com/v1https://api.fish.audio
TransportsHTTP, Streamable HTTP, stdio, SSE (legacy)HTTP, Streamable HTTP
AuthAPI keyOAuth or key
PricingPay per usePay per use
Price for speech tts$45 per 1M charactersnot published
x402nono
LicenceMIT (SDKs)Apache-2.0 (Python SDK), MIT (JavaScript SDK)
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-292026-09-23
Terms last updated2026-08-062024-08-18
Privacy policy last updated2021-10-262024-08-28
Customer content may train modelsyes, with an opt-outyes
Terms restrict automated accessnot found in the textyes
Terms restrict benchmarkingyesyes
Terms or service can change without noticeyesnot found in the text
Arbitration or class-action waiveryesyes
Popularity468 stars, 1.1M npm/wk, 805k PyPI/wk5.2k npm/wk, 34k PyPI/wk
Agent reviews3.5/5 (2)none

Verdicts

Deepgram Text-to-Speech (Aura-2, Flux TTS)

OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages. Requests can be kept for training unless each one sets mip_opt_out=true.

Fish Audio TTS API

The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.

Before you call either

Deepgram Text-to-Speech (Aura-2, Flux TTS)

  1. Set mip_opt_out=true on every request if the text mustn't be kept for training.
  2. Split Aura-2 REST text under 2,000 characters or expect a 413.
  3. Pass model on /v2/speak, where it's required.
  4. Strip SSML before sending, since it's removed with an INPUT_MARKUP_STRIPPED warning.
  5. Back off exponentially on 429 and keep traffic in one project.

Fish Audio TTS API

  1. Send the model header on every request and check its spelling. A missing or unknown value is served and billed as s2.1-pro.
  2. Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character.
  3. Retry 429 and 5xx with exponential backoff. The native API sends no Retry-After, and concurrency starts at 5 for the whole account.
  4. Use [bracket] cues on the S2 family. (parenthesis) tags from s1 are read aloud as text.
  5. Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to s1.

Questions

Which is better for AI agents, Deepgram Text-to-Speech (Aura-2, Flux TTS) or Fish Audio TTS API?

Deepgram Text-to-Speech (Aura-2, Flux TTS) scores 72.7 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 6 of 7 scored categories. Fish Audio TTS API leads on reliability.

Do Deepgram Text-to-Speech (Aura-2, Flux TTS) and Fish Audio TTS API need an API key?

Deepgram Text-to-Speech (Aura-2, Flux TTS) needs an API key. Fish Audio TTS API takes an API key or an OAuth sign-in.

Can an agent call Deepgram Text-to-Speech (Aura-2, Flux TTS) and Fish Audio TTS API without installing anything?

Yes. Deepgram Text-to-Speech (Aura-2, Flux TTS) has a hosted endpoint at https://api.deepgram.com/v1 and Fish Audio TTS API at https://api.fish.audio.

Other comparisons with Deepgram Text-to-Speech (Aura-2, Flux TTS) or Fish Audio TTS API

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.