Head to head · Text-to-speech · October 2026 research run

Fish Audio TTS API vs HeyGen Voice

HeyGen Voice scores 71.3 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 4 of 7 scored categories. Both do text-to-speech.

Best text-to-speech APIs for AI agents · All 66 tts comparisons

Which one, for what

Fish Audio TTS API C

Best for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.

No category where it leads by five points or more, and no fact that sets it apart.

Watch for

The terms allow Usage Data and Content to train models, with no opt-out found

HeyGen Voice BB

Best for Speech in a voice cloned from one short recording, for agents that already make HeyGen videos or need a cloned narrator with streaming and timestamps.

Ahead on

  • Schema & documentation, 94 against 84
  • Agent ergonomics, 83 against 67
  • Security & auth, 69 against 35
  • Transparency & trust, 69 against 60

Also in its favour

  • Agent-ready, a grade of BB or better

Watch for

Speaks only voices cloned in the caller's workspace. Stock and designed voices go through a separate endpoint and engines

Score by category

CategoryWeight this runFish Audio TTS APIHeyGen VoiceEdge
Reliability16%207575even
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28494HeyGen Voice +10
Agent ergonomics13%16.26783HeyGen Voice +16
Security & auth14%17.53569HeyGen Voice +34
Payments & pricing10%12.53030even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86965Fish Audio TTS API +4
Transparency & trust7%8.86069HeyGen Voice +9
Negative events≤1500
TotalC 60.7/100BB 71.3/100

Facts side by side

FactFish Audio TTS APIHeyGen Voice
KindModel APIModel API
VendorFish AudioHeyGen
Hosted endpointhttps://api.fish.audiohttps://api.heygen.com
TransportsHTTP, Streamable HTTPHTTP
AuthOAuth or keyAPI key
PricingPay per usePay per use
x402nono
LicenceApache-2.0 (Python SDK), MIT (JavaScript SDK)Closed service under HeyGen's Terms of Service (HeyGen Technology, Inc., last updated 23 July 2026).
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-23none
Terms last updated2024-08-18
Privacy policy last updated2024-08-28
Customer content may train modelsyes
Terms restrict automated accessyes
Terms restrict benchmarkingyes
Terms or service can change without noticenot found in the text
Arbitration or class-action waiveryes
Popularity5.2k npm/wk, 34k PyPI/wknone

Verdicts

Fish Audio TTS API

The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.

HeyGen Voice

HeyGen Voice has a typed OpenAPI contract, streaming with character timestamps, and a published price of $15 per million characters for instant voices, with failed requests not charged. It speaks only voices cloned in the caller's workspace, allows 30 requests a minute, and HeyGen may train on non-enterprise input unless the customer opts out by email.

Before you call either

Fish Audio TTS API

  1. Send the model header on every request and check its spelling. A missing or unknown value is served and billed as s2.1-pro.
  2. Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character.
  3. Retry 429 and 5xx with exponential backoff. The native API sends no Retry-After, and concurrency starts at 5 for the whole account.
  4. Use [bracket] cues on the S2 family. (parenthesis) tags from s1 are read aloud as text.
  5. Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to s1.

HeyGen Voice

  1. Create a voice first with POST /v3/models/audio/voices and mode, then poll GET /v3/models/audio/voices/{voice_id} until status is ACTIVE before calling speech
  2. Send expressiveness_boost only for instant voices and seed, speed, pitch_shift, pitch_variance or <break> tags only for professional voices. Mixing them returns 400 invalid_parameter
  3. On the stream, decode and play each base64 WAV part in part_index order. Do not concatenate the bytes
  4. Send an Idempotency-Key when creating a voice, and retry 502, 503 and 504 on speech, which the OpenAPI file marks as safe
  5. Use POST /v3/voices/speech for stock or designed voices. It is a different engine and price

Questions

Which is better for AI agents, Fish Audio TTS API or HeyGen Voice?

HeyGen Voice scores 71.3 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 4 of 7 scored categories.

Do Fish Audio TTS API and HeyGen Voice need an API key?

Fish Audio TTS API takes an API key or an OAuth sign-in. HeyGen Voice needs an API key.

Can an agent call Fish Audio TTS API and HeyGen Voice without installing anything?

Yes. Fish Audio TTS API has a hosted endpoint at https://api.fish.audio and HeyGen Voice at https://api.heygen.com.

Other comparisons with Fish Audio TTS API or HeyGen Voice

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.