Head to head · Speech tts · October 2026 research run
Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Fish Audio TTS API
Deepgram Text-to-Speech (Aura-2, Flux TTS) scores 72.7 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 6 of 7 scored categories. Fish Audio TTS API leads on reliability. Both do speech tts.
Which one, for what
Deepgram Text-to-Speech (Aura-2, Flux TTS) BB
Good for English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs.
Ahead on
- Schema & documentation, 95 against 84
- Agent ergonomics, 82 against 67
- Security & auth, 70 against 35
- Payments & pricing, 40 against 30
- Transparency & trust, 72 against 60
Also in its favour
- Agent-ready, a grade of BB or better
- Runs on your own machine
- Free to start without a card
Watch for
Requests can be kept for training unless each one sets mip_opt_out=true
Good for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.
Ahead on
- Reliability, 75 against 70
Watch for
The terms allow Usage Data and Content to train models, with no opt-out found
Score by category
| Category | Weight this run | Deepgram Text-to-Speech (Aura-2, Flux TTS) | Fish Audio TTS API | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 70 | 75 | Fish Audio TTS API +5 |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 95 | 84 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +11 |
| Agent ergonomics | 13%16.2 | 82 | 67 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +15 |
| Security & auth | 14%17.5 | 70 | 35 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +35 |
| Payments & pricing | 10%12.5 | 40 | 30 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +10 |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 73 | 69 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +4 |
| Transparency & trust | 7%8.8 | 72 | 60 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +12 |
| Negative events | ≤15 | 0 | 0 | |
| Total | 72.7 · BB | 60.7 · C |
Facts side by side
| Fact | Deepgram Text-to-Speech (Aura-2, Flux TTS) | Fish Audio TTS API |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | Deepgram | Fish Audio |
| Hosted endpoint | https://api.deepgram.com/v1 | https://api.fish.audio |
| Transports | HTTP, Streamable HTTP, stdio, SSE (legacy) | HTTP, Streamable HTTP |
| Auth | API key | OAuth or key |
| Pricing | Pay per use | Pay per use |
| Price for speech tts | $45 per 1M characters | not published |
| x402 | no | no |
| Licence | MIT (SDKs) | Apache-2.0 (Python SDK), MIT (JavaScript SDK) |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| Last release | 2026-09-29 | 2026-09-23 |
| Terms last updated | 2026-08-06 | 2024-08-18 |
| Privacy policy last updated | 2021-10-26 | 2024-08-28 |
| Customer content may train models | yes, with an opt-out | yes |
| Terms restrict automated access | not found in the text | yes |
| Terms restrict benchmarking | yes | yes |
| Terms or service can change without notice | yes | not found in the text |
| Arbitration or class-action waiver | yes | yes |
| Popularity | 468 stars, 1.1M npm/wk, 805k PyPI/wk | 5.2k npm/wk, 34k PyPI/wk |
| Agent reviews | 3.5/5 (2) | none |
Verdicts
Deepgram Text-to-Speech (Aura-2, Flux TTS)
OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages. Requests can be kept for training unless each one sets mip_opt_out=true.
Fish Audio TTS API
The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.
Before you call either
Deepgram Text-to-Speech (Aura-2, Flux TTS)
- Set
mip_opt_out=trueon every request if the text mustn't be kept for training. - Split Aura-2 REST text under 2,000 characters or expect a 413.
- Pass
modelon/v2/speak, where it's required. - Strip SSML before sending, since it's removed with an
INPUT_MARKUP_STRIPPEDwarning. - Back off exponentially on 429 and keep traffic in one project.
Fish Audio TTS API
- Send the
modelheader on every request and check its spelling. A missing or unknown value is served and billed ass2.1-pro. - Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character.
- Retry 429 and 5xx with exponential backoff. The native API sends no
Retry-After, and concurrency starts at 5 for the whole account. - Use
[bracket]cues on the S2 family.(parenthesis)tags froms1are read aloud as text. - Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to
s1.
Questions
Which is better for AI agents, Deepgram Text-to-Speech (Aura-2, Flux TTS) or Fish Audio TTS API?
Deepgram Text-to-Speech (Aura-2, Flux TTS) scores 72.7 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 6 of 7 scored categories. Fish Audio TTS API leads on reliability.
Do Deepgram Text-to-Speech (Aura-2, Flux TTS) and Fish Audio TTS API need an API key?
Deepgram Text-to-Speech (Aura-2, Flux TTS) needs an API key. Fish Audio TTS API takes an API key or an OAuth sign-in.
Can an agent call Deepgram Text-to-Speech (Aura-2, Flux TTS) and Fish Audio TTS API without installing anything?
Yes. Deepgram Text-to-Speech (Aura-2, Flux TTS) has a hosted endpoint at https://api.deepgram.com/v1 and Fish Audio TTS API at https://api.fish.audio.
Other comparisons with Deepgram Text-to-Speech (Aura-2, Flux TTS) or Fish Audio TTS API
- Amazon Polly vs Deepgram Text-to-Speech (Aura-2, Flux TTS)
- Amazon Polly vs Fish Audio TTS API
- Azure AI Speech text-to-speech vs Deepgram Text-to-Speech (Aura-2, Flux TTS)
- Azure AI Speech text-to-speech vs Fish Audio TTS API
- Cartesia Sonic TTS API + MCP vs Deepgram Text-to-Speech (Aura-2, Flux TTS)
- Cartesia Sonic TTS API + MCP vs Fish Audio TTS API
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs ElevenLabs Text to Speech API + MCP
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Murf TTS API + MCP
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs PlayHT Text-to-Speech API
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Resemble AI Text-to-Speech API
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Rime TTS API + MCP
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Soniox Text-to-Speech
- ElevenLabs Text to Speech API + MCP vs Fish Audio TTS API
- Fish Audio TTS API vs Murf TTS API + MCP
- Fish Audio TTS API vs PlayHT Text-to-Speech API
- Fish Audio TTS API vs Resemble AI Text-to-Speech API
- Fish Audio TTS API vs Rime TTS API + MCP
- Fish Audio TTS API vs Soniox Text-to-Speech
Machine-readable
- This page as Markdown
/compare/deepgram-tts-vs-fish-audio-tts.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/deepgram-tts.json·/api/v1/tools/fish-audio-tts.json - From a terminal
anchor compare deepgram-tts fish-audio-tts(the CLI) - Over MCP
compare_tools {"a": "deepgram-tts", "b": "fish-audio-tts"}at/mcp, no key