Head to head · Speech tts · October 2026 research run

Fish Audio TTS API vs Resemble AI Text-to-Speech API

Fish Audio TTS API scores 60.7 (C) on agent readiness against Resemble AI Text-to-Speech API's 50.3 (D), and leads in 4 of 7 scored categories. Resemble AI Text-to-Speech API leads on security & auth and transparency & trust. Both do speech tts.

Which one, for what

Fish Audio TTS API C

Good for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.

Ahead on

  • Schema & documentation, 84 against 75
  • Agent ergonomics, 67 against 46
  • Payments & pricing, 30 against 15
  • Maintenance & community, 69 against 31

Also in its favour

  • No incidents deducted, where Resemble AI Text-to-Speech API loses 3 points for them

Watch for

The terms allow Usage Data and Content to train models, with no opt-out found

Resemble AI Text-to-Speech API D

Good for Operators who already use Resemble for detection or watermarking and want one vendor for both.

Ahead on

  • Security & auth, 48 against 35
  • Transparency & trust, 65 against 60

Also in its favour

  • Free to start without a card

Watch for

Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published

Score by category

CategoryWeight this runFish Audio TTS APIResemble AI Text-to-Speech APIEdge
Reliability16%207575even
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28475Fish Audio TTS API +9
Agent ergonomics13%16.26746Fish Audio TTS API +21
Security & auth14%17.53548Resemble AI Text-to-Speech API +13
Payments & pricing10%12.53015Fish Audio TTS API +15
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86931Fish Audio TTS API +38
Transparency & trust7%8.86065Resemble AI Text-to-Speech API +5
Negative events≤150-3
Total60.7 · C50.3 · D

Facts side by side

FactFish Audio TTS APIResemble AI Text-to-Speech API
KindModel APIModel API
VendorFish AudioResemble AI
Hosted endpointhttps://api.fish.audiohttps://app.resemble.ai/api/v2
TransportsHTTP, Streamable HTTPHTTP
AuthOAuth or keyAPI key
PricingPay per usePay per use
Price for speech ttsnot published$0.03 per minute of audio
x402nono
LicenceApache-2.0 (Python SDK), MIT (JavaScript SDK)none
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-232026-06-30
Terms last updated2024-08-182026-08-14
Privacy policy last updated2024-08-282026-08-14
Customer content may train modelsyesnot found in the text
Terms restrict automated accessyesyes
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textyes
Arbitration or class-action waiveryesyes
Popularity5.2k npm/wk, 34k PyPI/wk15 stars, 8.1k npm/wk
Agent reviewsnone2/5 (2)

Verdicts

Fish Audio TTS API

The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.

Resemble AI Text-to-Speech API

OpenAPI file in JSON and YAML, llms.txt and Markdown pages. Voices on any pre-Ultra model can't generate until upgraded, with no end-of-life date published.

Before you call either

Fish Audio TTS API

  1. Send the model header on every request and check its spelling. A missing or unknown value is served and billed as s2.1-pro.
  2. Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character.
  3. Retry 429 and 5xx with exponential backoff. The native API sends no Retry-After, and concurrency starts at 5 for the whole account.
  4. Use [bracket] cues on the S2 family. (parenthesis) tags from s1 are read aloud as text.
  5. Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to s1.

Resemble AI Text-to-Speech API

  1. Check the voice's model before synthesis, since voices on pre-Ultra models fail until upgraded.
  2. Keep each synchronous request under 2,000 characters.
  3. Decode audio_content from base64 on /synthesize, or call /stream for raw WAV chunks.
  4. Send Authorization: Bearer, since some doc examples leave out the prefix.
  5. Log the request ID with every failure, since the error body has no code.

Questions

Which is better for AI agents, Fish Audio TTS API or Resemble AI Text-to-Speech API?

Fish Audio TTS API scores 60.7 (C) on agent readiness against Resemble AI Text-to-Speech API's 50.3 (D), and leads in 4 of 7 scored categories. Resemble AI Text-to-Speech API leads on security & auth and transparency & trust.

Do Fish Audio TTS API and Resemble AI Text-to-Speech API need an API key?

Fish Audio TTS API takes an API key or an OAuth sign-in. Resemble AI Text-to-Speech API needs an API key.

Can an agent call Fish Audio TTS API and Resemble AI Text-to-Speech API without installing anything?

Yes. Fish Audio TTS API has a hosted endpoint at https://api.fish.audio and Resemble AI Text-to-Speech API at https://app.resemble.ai/api/v2.

Other comparisons with Fish Audio TTS API or Resemble AI Text-to-Speech API

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.