Head to head · Speech tts · October 2026 research run

Fish Audio TTS API vs Soniox Text-to-Speech

Soniox Text-to-Speech scores 63.7 (B) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 4 of 7 scored categories. Fish Audio TTS API leads on schema & documentation, payments & pricing and maintenance & community. Both do speech tts.

Which one, for what

Fish Audio TTS API C

Good for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.

Ahead on

  • Schema & documentation, 84 against 60
  • Payments & pricing, 30 against 20
  • Maintenance & community, 69 against 50

Watch for

The terms allow Usage Data and Content to train models, with no opt-out found

Soniox Text-to-Speech B

Good for Multilingual agents that need any voice in any of 60-plus languages, short turns and strict data handling.

Ahead on

  • Reliability, 83 against 75
  • Security & auth, 75 against 35
  • Transparency & trust, 72 against 60

Watch for

Audio stops at 2 minutes per request or stream and the cap can't be raised

Score by category

CategoryWeight this runFish Audio TTS APISoniox Text-to-SpeechEdge
Reliability16%207583Soniox Text-to-Speech +8
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28460Fish Audio TTS API +24
Agent ergonomics13%16.26768Soniox Text-to-Speech +1
Security & auth14%17.53575Soniox Text-to-Speech +40
Payments & pricing10%12.53020Fish Audio TTS API +10
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86950Fish Audio TTS API +19
Transparency & trust7%8.86072Soniox Text-to-Speech +12
Negative events≤1500
Total60.7 · C63.7 · B

Facts side by side

FactFish Audio TTS APISoniox Text-to-Speech
KindModel APIModel API
VendorFish AudioSoniox
Hosted endpointhttps://api.fish.audiohttps://tts-rt.soniox.com
TransportsHTTP, Streamable HTTPHTTP
AuthOAuth or keyAPI key
PricingPay per usePay per use
Price for speech ttsnot published$0.0117 per minute of audio
x402nono
LicenceApache-2.0 (Python SDK), MIT (JavaScript SDK)Apache-2.0 (Python SDK)
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-232026-08-11
Terms last updated2024-08-182026-06-29
Privacy policy last updated2024-08-282026-06-29
Customer content may train modelsyesnot found in the text
Terms restrict automated accessyesnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textyes
Arbitration or class-action waiveryesnot found in the text
Popularity5.2k npm/wk, 34k PyPI/wk12 stars, 22k npm/wk
Agent reviewsnone3/5 (2)

Verdicts

Fish Audio TTS API

The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.

Soniox Text-to-Speech

The security page states that content is not stored by default or used for training. Audio output is capped at two minutes per request or stream.

Before you call either

Fish Audio TTS API

  1. Send the model header on every request and check its spelling. A missing or unknown value is served and billed as s2.1-pro.
  2. Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character.
  3. Retry 429 and 5xx with exponential backoff. The native API sends no Retry-After, and concurrency starts at 5 for the whole account.
  4. Use [bracket] cues on the S2 family. (parenthesis) tags from s1 are read aloud as text.
  5. Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to s1.

Soniox Text-to-Speech

  1. Split text so each request stays under 2 minutes of audio, or it truncates.
  2. Branch on error_type, not the message, and back off on limit_exceeded.
  3. Use a temporary API key for browser clients.
  4. Use bracketed audio tags such as [whispering] instead of SSML.
  5. Pick the regional host (EU, Japan, India) that matches your data residency.

Questions

Which is better for AI agents, Fish Audio TTS API or Soniox Text-to-Speech?

Soniox Text-to-Speech scores 63.7 (B) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 4 of 7 scored categories. Fish Audio TTS API leads on schema & documentation, payments & pricing and maintenance & community.

Do Fish Audio TTS API and Soniox Text-to-Speech need an API key?

Fish Audio TTS API takes an API key or an OAuth sign-in. Soniox Text-to-Speech needs an API key.

Can an agent call Fish Audio TTS API and Soniox Text-to-Speech without installing anything?

Yes. Fish Audio TTS API has a hosted endpoint at https://api.fish.audio and Soniox Text-to-Speech at https://tts-rt.soniox.com.

Other comparisons with Fish Audio TTS API or Soniox Text-to-Speech

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.