Head to head · Text-to-speech · October 2026 research run
Fish Audio TTS API vs HeyGen Voice
HeyGen Voice scores 71.3 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 4 of 7 scored categories. Both do text-to-speech.
Best text-to-speech APIs for AI agents · All 66 tts comparisons
Which one, for what
Best for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.
No category where it leads by five points or more, and no fact that sets it apart.
Watch for
The terms allow Usage Data and Content to train models, with no opt-out found
HeyGen Voice BB
Best for Speech in a voice cloned from one short recording, for agents that already make HeyGen videos or need a cloned narrator with streaming and timestamps.
Ahead on
- Schema & documentation, 94 against 84
- Agent ergonomics, 83 against 67
- Security & auth, 69 against 35
- Transparency & trust, 69 against 60
Also in its favour
- Agent-ready, a grade of BB or better
Watch for
Speaks only voices cloned in the caller's workspace. Stock and designed voices go through a separate endpoint and engines
Score by category
| Category | Weight this run | Fish Audio TTS API | HeyGen Voice | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 75 | 75 | even |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 84 | 94 | HeyGen Voice +10 |
| Agent ergonomics | 13%16.2 | 67 | 83 | HeyGen Voice +16 |
| Security & auth | 14%17.5 | 35 | 69 | HeyGen Voice +34 |
| Payments & pricing | 10%12.5 | 30 | 30 | even |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 69 | 65 | Fish Audio TTS API +4 |
| Transparency & trust | 7%8.8 | 60 | 69 | HeyGen Voice +9 |
| Negative events | ≤15 | 0 | 0 | |
| Total | C 60.7/100 | BB 71.3/100 |
Facts side by side
| Fact | Fish Audio TTS API | HeyGen Voice |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | Fish Audio | HeyGen |
| Hosted endpoint | https://api.fish.audio | https://api.heygen.com |
| Transports | HTTP, Streamable HTTP | HTTP |
| Auth | OAuth or key | API key |
| Pricing | Pay per use | Pay per use |
| x402 | no | no |
| Licence | Apache-2.0 (Python SDK), MIT (JavaScript SDK) | Closed service under HeyGen's Terms of Service (HeyGen Technology, Inc., last updated 23 July 2026). |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| Last release | 2026-09-23 | none |
| Terms last updated | 2024-08-18 | |
| Privacy policy last updated | 2024-08-28 | |
| Customer content may train models | yes | |
| Terms restrict automated access | yes | |
| Terms restrict benchmarking | yes | |
| Terms or service can change without notice | not found in the text | |
| Arbitration or class-action waiver | yes | |
| Popularity | 5.2k npm/wk, 34k PyPI/wk | none |
Verdicts
Fish Audio TTS API
The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.
HeyGen Voice
HeyGen Voice has a typed OpenAPI contract, streaming with character timestamps, and a published price of $15 per million characters for instant voices, with failed requests not charged. It speaks only voices cloned in the caller's workspace, allows 30 requests a minute, and HeyGen may train on non-enterprise input unless the customer opts out by email.
Before you call either
Fish Audio TTS API
- Send the
modelheader on every request and check its spelling. A missing or unknown value is served and billed ass2.1-pro. - Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character.
- Retry 429 and 5xx with exponential backoff. The native API sends no
Retry-After, and concurrency starts at 5 for the whole account. - Use
[bracket]cues on the S2 family.(parenthesis)tags froms1are read aloud as text. - Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to
s1.
HeyGen Voice
- Create a voice first with POST /v3/models/audio/voices and
mode, then poll GET /v3/models/audio/voices/{voice_id} untilstatusis ACTIVE before calling speech - Send
expressiveness_boostonly for instant voices andseed,speed,pitch_shift,pitch_varianceor<break>tags only for professional voices. Mixing them returns 400 invalid_parameter - On the stream, decode and play each base64 WAV part in
part_indexorder. Do not concatenate the bytes - Send an Idempotency-Key when creating a voice, and retry 502, 503 and 504 on speech, which the OpenAPI file marks as safe
- Use POST /v3/voices/speech for stock or designed voices. It is a different engine and price
Questions
Which is better for AI agents, Fish Audio TTS API or HeyGen Voice?
HeyGen Voice scores 71.3 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in 4 of 7 scored categories.
Do Fish Audio TTS API and HeyGen Voice need an API key?
Fish Audio TTS API takes an API key or an OAuth sign-in. HeyGen Voice needs an API key.
Can an agent call Fish Audio TTS API and HeyGen Voice without installing anything?
Yes. Fish Audio TTS API has a hosted endpoint at https://api.fish.audio and HeyGen Voice at https://api.heygen.com.
Other comparisons with Fish Audio TTS API or HeyGen Voice
- Amazon Polly vs Fish Audio TTS API
- Amazon Polly vs HeyGen Voice
- Azure AI Speech text-to-speech vs Fish Audio TTS API
- Azure AI Speech text-to-speech vs HeyGen Voice
- Cartesia Sonic TTS API + MCP vs Fish Audio TTS API
- Cartesia Sonic TTS API + MCP vs HeyGen Voice
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Fish Audio TTS API
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs HeyGen Voice
- ElevenLabs Text to Speech API + MCP vs Fish Audio TTS API
- ElevenLabs Text to Speech API + MCP vs HeyGen Voice
- Fish Audio TTS API vs Murf TTS API + MCP
- Fish Audio TTS API vs PlayHT Text-to-Speech API
- Fish Audio TTS API vs Resemble AI Text-to-Speech API
- Fish Audio TTS API vs Rime TTS API + MCP
- Fish Audio TTS API vs Soniox Text-to-Speech
- HeyGen Voice vs Murf TTS API + MCP
- HeyGen Voice vs PlayHT Text-to-Speech API
- HeyGen Voice vs Resemble AI Text-to-Speech API
- HeyGen Voice vs Rime TTS API + MCP
- HeyGen Voice vs Soniox Text-to-Speech
Machine-readable
- This page as Markdown
/compare/fish-audio-tts-vs-heygen-voice.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/fish-audio-tts.json·/api/v1/tools/heygen-voice.json - From a terminal
anchor compare fish-audio-tts heygen-voice(the CLI) - Over MCP
compare_tools {"a": "fish-audio-tts", "b": "heygen-voice"}at/mcp, no key