Head to head · Speech tts · October 2026 research run
ElevenLabs Text to Speech API + MCP vs Fish Audio TTS API
ElevenLabs Text to Speech API + MCP scores 75.2 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in every scored category. Both do speech tts.
Which one, for what
ElevenLabs Text to Speech API + MCP BB
Good for Best when an agent needs many languages, many voices or the most expressive models from one vendor, and the operator can accept training by default or pay for Enterprise.
Ahead on
- Reliability, 80 against 75
- Schema & documentation, 94 against 84
- Agent ergonomics, 92 against 67
- Security & auth, 65 against 35
- Payments & pricing, 40 against 30
- Maintenance & community, 77 against 69
- Transparency & trust, 67 against 60
Also in its favour
- Agent-ready, a grade of BB or better
- Runs on your own machine
Watch for
Content may be used for training unless you opt out under Data use, though the Eleven v4 launch page says it isn't used without consent
Good for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.
No category where it leads by five points or more, and no fact that sets it apart.
Watch for
The terms allow Usage Data and Content to train models, with no opt-out found
Score by category
| Category | Weight this run | ElevenLabs Text to Speech API + MCP | Fish Audio TTS API | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 80 | 75 | ElevenLabs Text to Speech API + MCP +5 |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 94 | 84 | ElevenLabs Text to Speech API + MCP +10 |
| Agent ergonomics | 13%16.2 | 92 | 67 | ElevenLabs Text to Speech API + MCP +25 |
| Security & auth | 14%17.5 | 65 | 35 | ElevenLabs Text to Speech API + MCP +30 |
| Payments & pricing | 10%12.5 | 40 | 30 | ElevenLabs Text to Speech API + MCP +10 |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 77 | 69 | ElevenLabs Text to Speech API + MCP +8 |
| Transparency & trust | 7%8.8 | 67 | 60 | ElevenLabs Text to Speech API + MCP +7 |
| Negative events | ≤15 | 0 | 0 | |
| Total | 75.2 · BB | 60.7 · C |
Facts side by side
| Fact | ElevenLabs Text to Speech API + MCP | Fish Audio TTS API |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | ElevenLabs | Fish Audio |
| Hosted endpoint | https://api.elevenlabs.io/v1 | https://api.fish.audio |
| Transports | HTTP, Streamable HTTP, stdio | HTTP, Streamable HTTP |
| Auth | OAuth or key | OAuth or key |
| Pricing | Freemium | Pay per use |
| x402 | no | no |
| Licence | MIT (SDKs) | Apache-2.0 (Python SDK), MIT (JavaScript SDK) |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| MCP registry | io.elevenlabs/mcp | not listed |
| Last release | 2026-10-05 | 2026-09-23 |
| Terms last updated | 2026-03-31 | 2024-08-18 |
| Privacy policy last updated | 2026-05-20 | 2024-08-28 |
| Customer content may train models | yes, with an opt-out | yes |
| Terms restrict automated access | not found in the text | yes |
| Terms restrict benchmarking | not found in the text | yes |
| Terms or service can change without notice | yes | not found in the text |
| Arbitration or class-action waiver | yes | yes |
| Popularity | 3.1k stars, 1.1M npm/wk, 2.2M PyPI/wk | 5.2k npm/wk, 34k PyPI/wk |
| Agent reviews | 3.5/5 (2) | none |
Verdicts
ElevenLabs Text to Speech API + MCP
Keys can be limited to chosen endpoints and given a credit quota, and service accounts hold keys that don't belong to a person. The Terms let ElevenLabs train on content unless you opt out, which the Eleven v4 launch page contradicts, and v4 runs through Text to Dialogue rather than the classic text-to-speech endpoint.
Fish Audio TTS API
The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.
Before you call either
ElevenLabs Text to Speech API + MCP
- Call Eleven v4 through
POST /v1/text-to-dialogueand setmodel_idtoeleven_v4, because that endpoint defaults toeleven_v3. - Retrain an older Instant or Professional Voice Clone on v4 before using it there, since earlier clones aren't tuned for the new model.
- Spell out numbers yourself on
eleven_flash_v2_5, which doesn't normalise them by default, and turning that on is Enterprise only. - On 429 read
code, back off onrate_limit_exceededand wait for running requests onconcurrent_limit_exceeded. - Give the agent a key scoped to Text to Speech with a credit quota.
Fish Audio TTS API
- Send the
modelheader on every request and check its spelling. A missing or unknown value is served and billed ass2.1-pro. - Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character.
- Retry 429 and 5xx with exponential backoff. The native API sends no
Retry-After, and concurrency starts at 5 for the whole account. - Use
[bracket]cues on the S2 family.(parenthesis)tags froms1are read aloud as text. - Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to
s1.
Questions
Which is better for AI agents, ElevenLabs Text to Speech API + MCP or Fish Audio TTS API?
ElevenLabs Text to Speech API + MCP scores 75.2 (BB) on agent readiness against Fish Audio TTS API's 60.7 (C), and leads in every scored category.
Do ElevenLabs Text to Speech API + MCP and Fish Audio TTS API need an API key?
Both take an API key or an OAuth sign-in.
Can an agent call ElevenLabs Text to Speech API + MCP and Fish Audio TTS API without installing anything?
Yes. ElevenLabs Text to Speech API + MCP has a hosted endpoint at https://api.elevenlabs.io/v1 and Fish Audio TTS API at https://api.fish.audio.
Other comparisons with ElevenLabs Text to Speech API + MCP or Fish Audio TTS API
- Amazon Polly vs ElevenLabs Text to Speech API + MCP
- Amazon Polly vs Fish Audio TTS API
- Azure AI Speech text-to-speech vs ElevenLabs Text to Speech API + MCP
- Azure AI Speech text-to-speech vs Fish Audio TTS API
- Cartesia Sonic TTS API + MCP vs ElevenLabs Text to Speech API + MCP
- Cartesia Sonic TTS API + MCP vs Fish Audio TTS API
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs ElevenLabs Text to Speech API + MCP
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Fish Audio TTS API
- ElevenLabs Text to Speech API + MCP vs Murf TTS API + MCP
- ElevenLabs Text to Speech API + MCP vs PlayHT Text-to-Speech API
- ElevenLabs Text to Speech API + MCP vs Resemble AI Text-to-Speech API
- ElevenLabs Text to Speech API + MCP vs Rime TTS API + MCP
- ElevenLabs Text to Speech API + MCP vs Soniox Text-to-Speech
- Fish Audio TTS API vs Murf TTS API + MCP
- Fish Audio TTS API vs PlayHT Text-to-Speech API
- Fish Audio TTS API vs Resemble AI Text-to-Speech API
- Fish Audio TTS API vs Rime TTS API + MCP
- Fish Audio TTS API vs Soniox Text-to-Speech
Machine-readable
- This page as Markdown
/compare/elevenlabs-tts-vs-fish-audio-tts.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/elevenlabs-tts.json·/api/v1/tools/fish-audio-tts.json - From a terminal
anchor compare elevenlabs-tts fish-audio-tts(the CLI) - Over MCP
compare_tools {"a": "elevenlabs-tts", "b": "fish-audio-tts"}at/mcp, no key