# Fish Audio TTS API (slim) > Fish Audio's API turns text into speech with the s2.1-pro model in 83 languages, over a REST endpoint, a timestamped stream and a WebSocket that accepts text as it is produced. - Full: https://www.anchorterminal.com/tools/fish-audio-tts.md (~7,800 tokens) · this version ~1,680 tokens · JSON https://www.anchorterminal.com/tools/fish-audio-tts.json · canonical https://www.anchorterminal.com/tools/fish-audio-tts - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **C · 60.7/100 · rank #453 of 842 · #8 in Text-to-speech · not agent-ready · confidence medium** Assessment: The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation. ## Facts - Kind: Model API · vendor: Fish Audio · category: Text-to-speech · legal entity: Hanabi AI Inc. · provenance 75/100 - Endpoint: `https://api.fish.audio` (HTTP, Streamable HTTP) - Auth: OAuth or key · pricing: Pay per use · x402: no · licence: Apache-2.0 (Python SDK), MIT (JavaScript SDK) - Probe metrics: not measured yet (probes haven't run) - Models: `s2.1-pro` (default), `s2.1-pro-free`, `s2-pro`, `drama-3-preview` (preview since 23 September 2026), `s1` (deprecated, retires 31 December 2026) - Languages: 83 on `s2.1-pro`, detected automatically from the text. 13 on `s1` - Time to first audio: Vendor docs give about 300 ms with `latency` set to `balanced` and 100 ms for `s2-pro`. Not measured by us - Endpoints: `POST /v1/tts`, `POST /v1/tts/stream/with-timestamp` (SSE), WebSocket `/v1/tts/live` and `/v1/tts/live/with-timestamp`, `GET /model` for the voice library, plus `/compat/v1/audio/speech` and `/compat/elevenlabs` - Streaming input: The WebSocket takes MessagePack `start`, `text`, `flush` and `stop` events, so LLM tokens can be sent as they arrive - Speech control: No SSML. Free-form `[bracket]` cues on the S2 family, `prosody.speed` from 0.5 to 2.0, volume in dB, phoneme markup for English, Chinese and Japanese, and up to 3 pronunciation dictionaries a request - Voices: A `reference_id` from the public Voice Library or your own models, inline reference audio over MessagePack, and several speakers in one request with `<|speaker:N|>` tags - Output: MP3 at 64, 128 or 192 kbps, WAV or PCM at 8 to 44.1 kHz, Opus at 48 kHz, all mono - Free tier: `s2.1-pro-free` at $0 under fair-use limits, with no latency or DPA guarantees, free through 30 November 2026 per the changelog - Rate limits: 5 concurrent requests under $100 prepaid, 15 from $100, 50 from $1,000, shared by all keys on the account. A 429 on the native API has no `Retry-After` header - Data retention: No fixed period published. The privacy policy keeps Content as long as needed to run the systems, and the compatibility page says zero-retention mode cannot be honoured - Prices: Text to speech, s2.1-pro $15 per 1M characters; Text to speech, s2.1-pro-free free per 1M characters - Scores: Reliability 75, Performance pending, Schema & documentation 84, Agent ergonomics 67, Security & auth 35, Payments & pricing 30, Task success pending, Maintenance & community 69, Transparency & trust 60 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, OpenAPI 3.1 file at docs.fish.audio/api-reference/openapi.json, and an AsyncAPI schema rendered on the WebSocket page. · Agent ergonomics, API reading. · Security & auth, Model reading of the checklist, with training and retention in place of the least-privilege and injection lines, as for the other TTS listin… · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, Read as a model. · Transparency & trust, Closed API with public terms, SDKs under Apache-2.0 and MIT, and S2-Pro weights published under Fish Audio's own research licence, which is… - Sources: 29, open questions: 9, both in the full twin - Capabilities: speech.tts, speech.streaming, speech.voices, speech.languages, voice.clone - JSON: https://www.anchorterminal.com/api/v1/tools/fish-audio-tts.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/fish-audio-tts.svg` or a link to https://www.anchorterminal.com/tools/fish-audio-tts from a page on fish.audio or one of its subdomains, or the README of github.com/fishaudio/fish-audio-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Send the `model` header on every request and check its spelling. A missing or unknown value is served and billed as `s2.1-pro`. 2. Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character. 3. Retry 429 and 5xx with exponential backoff. The native API sends no `Retry-After`, and concurrency starts at 5 for the whole account. 4. Use `[bracket]` cues on the S2 family. `(parenthesis)` tags from `s1` are read aloud as text. 5. Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to `s1`. ## Connect ```bash pip install fish-audio-sdk ``` ```bash curl --request POST https://api.fish.audio/v1/tts \ --header "Authorization: Bearer $FISH_API_KEY" \ --header "Content-Type: application/json" \ --header "model: s2.1-pro-free" \ --data '{ "text": "Hello from Fish Audio!", "format": "mp3" }' \ --output hello.mp3 ``` ```bash claude mcp add --transport http fish-audio https://api.fish.audio/mcp ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Amazon Polly | BB | 75.6 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/amazon-polly.min.md | | ElevenLabs Text to Speech API + MCP | BB | 75.2 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/elevenlabs-tts.min.md | | Deepgram Text-to-Speech (Aura-2, Flux TTS) | BB | 72.7 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/deepgram-tts.min.md | | Azure AI Speech text-to-speech | BB | 71.4 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/azure-text-to-speech.min.md | | Murf TTS API + MCP | BB | 70.6 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/murf-tts.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)