Fish Audio TTS API
by Fish Audio Model API in Text-to-speech
Hosted
Hanabi AI Inc. · fish.audio since 2023 · status page · who's behind it
Fish Audio's API turns text into speech with the s2.1-pro model in 83 languages, over a REST endpoint, a timestamped stream and a WebSocket that accepts text as it is produced.
Good for Voice agents that stream LLM text into speech at a low per-byte price, multilingual output from one model, and teams moving from OpenAI or ElevenLabs request shapes.
Is this your product? Claim this listing or verify it
More from Fish Audio Fish Audio Voice Cloning API (Cloning)
Assessment. The API accepts streamed text over a WebSocket, returns word timestamps, and prices speech at $15 per million UTF-8 bytes with a $0 model for development. The terms allow training on customer content with no opt-out, and no retention period, SLA document or security certification was found in the reviewed documentation.
Facts
- Transport
- HTTP, Streamable HTTP
- Endpoint
https://api.fish.audio- Auth
- OAuth or key
- Pricing
- Pay per use · Pay per use
- x402
- No
- Licence
- Apache-2.0 (Python SDK), MIT (JavaScript SDK)
- Packages
pypifish-audio-sdknpmfish-audio- llms.txt
- published
- Last release
- npm / week
- 5.2k
- PyPI / week
- 34k
- Models
s2.1-pro(default),s2.1-pro-free,s2-pro,drama-3-preview(preview since 23 September 2026),s1(deprecated, retires 31 December 2026)- Languages
- 83 on
s2.1-pro, detected automatically from the text. 13 ons1 - Time to first audio
- Vendor docs give about 300 ms with
latencyset tobalancedand 100 ms fors2-pro. Not measured by us - Endpoints
POST /v1/tts,POST /v1/tts/stream/with-timestamp(SSE), WebSocket/v1/tts/liveand/v1/tts/live/with-timestamp,GET /modelfor the voice library, plus/compat/v1/audio/speechand/compat/elevenlabs- Streaming input
- The WebSocket takes MessagePack
start,text,flushandstopevents, so LLM tokens can be sent as they arrive - Speech control
- No SSML. Free-form
[bracket]cues on the S2 family,prosody.speedfrom 0.5 to 2.0, volume in dB, phoneme markup for English, Chinese and Japanese, and up to 3 pronunciation dictionaries a request - Voices
- A
reference_idfrom the public Voice Library or your own models, inline reference audio over MessagePack, and several speakers in one request with<|speaker:N|>tags - Output
- MP3 at 64, 128 or 192 kbps, WAV or PCM at 8 to 44.1 kHz, Opus at 48 kHz, all mono
- Free tier
s2.1-pro-freeat $0 under fair-use limits, with no latency or DPA guarantees, free through 30 November 2026 per the changelog- Rate limits
- 5 concurrent requests under $100 prepaid, 15 from $100, 50 from $1,000, shared by all keys on the account. A 429 on the native API has no
Retry-Afterheader - Data retention
- No fixed period published. The privacy policy keeps Content as long as needed to run the systems, and the compatibility page says zero-retention mode cannot be honoured
Facts verified 2026-10-09 from vendor docs, repositories and package registries. JSON · Markdown
Strengths
- Public OpenAPI 3.1 file, llms.txt, Markdown pages and two installable agent skills for the SDKs and the raw API
- WebSocket input takes text as it is produced, with
flushandstopevents and a variant that returns word timestamps s2.1-pro-freeruns the production model at $0 under fair-use limits, through 30 November 2026- OpenAI-compatible and ElevenLabs-compatible endpoints refuse unsupported options with a 4xx and send
Retry-Afteron a 429 - S1 retirement was announced on 8 October 2026 for 31 December 2026, with a migration guide
Weaknesses
- The terms allow Usage Data and Content to train models, with no opt-out found
- An unrecognised
modelheader falls back to paids2.1-prowithout an error - The Text-to-Speech API status component shows downtime on 11 days in 90, the longest 58 minutes on 6 August 2026, mostly on the free model
- The published SDKs date from March 2026.
fish-audio0.1.0 on npm defaults to the deprecateds1 - No security.txt, certification, DPA text or sub-processor list found in the pages read
Before you call it notes for agents
- Send the
modelheader on every request and check its spelling. A missing or unknown value is served and billed ass2.1-pro. - Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character.
- Retry 429 and 5xx with exponential backoff. The native API sends no
Retry-After, and concurrency starts at 5 for the whole account. - Use
[bracket]cues on the S2 family.(parenthesis)tags froms1are read aloud as text. - Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to
s1.
Who's behind it provenance 75/100
- Legal entity namedHanabi AI Inc.20/20
- Domain agefish.audio, registered 2023-12-11 (2 years)7/15
- Endpoint on the vendor's domainapi.fish.audio15/15
- Terms of serviceread, states 6 of the 7 things a reader expects, and has 3 clauses that cost points3.1/10
- Privacy policyread, states 8 of the 8 things a reader expects10/10
- Status pagestatus.fish.audio10/10
- Changelogpublished10/10
- security.txtnot found0/10
Terms and privacy, as read
Terms of service dated 2024-08-18, states 6 of 7, 5 to know
TL;DR Dated 2024-08-18. States 6 of the 7 things a reader expects, and we didn't find a service level. To know before relying on it, model training with no opt-out found, limits on automated access, limits on benchmarking, cut-off without notice or for any reason and 1 more.
Says it may use customer content to train or improve models, and no opt-out was foundcosts points
Usage Data and Content may be used to develop, train, or enhance artificial intelligence or machine learning models that are part of Fish.Audio’s products and services, including third-party components of the Services
Content an agent sends could end up in a model. An opt-out, where the document gives one, is shown instead.
Restricts automated accesscosts points
(h) “crawls,” “scrapes,” or “spiders” any page, data, or portion of or relating to the Services or Content (through use of manual or automated means);
A rule against bots, scrapers or automated means can cover an agent, depending on how the vendor reads it.
Restricts benchmarking or competitive usecosts points
(j) use output from the Services to develop models that compete with Fish.Audio;
A clause against publishing test results or using the service to build something that competes.
Says access can be ended without notice or for any reason
We reserve the right to terminate your right to use or access the Services at any time, for any reason, in our sole discretion, and without notice.
The vendor can suspend or close an account without warning, which would stop an agent mid-task.
Requires arbitration or waives class actions
Please read the following ARBITRATION AGREEMENT carefully because it requires you to arbitrate certain disputes and claims with Fish.Audio and limits the manner in which you can seek relief from Fish.Audio.
Disputes go to an arbitrator, or a customer gives up joining a class action or a jury trial.
Gives the date it was last updated Last updated 2024-08-18
Effective date: August 18, 2024
Without a date nobody can tell which version they agreed to.
Names the governing law or courts The law of the State of California
This Agreement is governed by and will be construed under the Federal Arbitration Act, applicable federal law, and the laws of the State of California, without regard to the conflicts of laws provisions thereof.
Says where a dispute would be heard and under whose law.
States a limit on its liability Capped at $100
…(B) ANY SUBSTITUTE GOODS, SERVICES OR TECHNOLOGY, (C) ANY AMOUNT, IN THE AGGREGATE, IN EXCESS OF THE GREATER OF (I) ONE-HUNDRED ($100) DOLLARS OR (II) THE AMOUNTS PAID AND/OR PAYABLE BY YOU TO Fish.Audio IN CONNECTION WITH THE SERVICES IN THE TWELVE (12) MONTH PERIOD PRECEDING THIS APPLICABLE CLAIM OR (D) ANY MATTER B…
Says the most the vendor would owe if the service causes a loss.
Says how the agreement or account can be ended
A violation of any of the foregoing is grounds for termination of your right to use or access the Services.
Says when the vendor can cut off access and what notice it gives.
Says how changes to the terms are announced Says it gives notice of a change
We reserve the right to change this Agreement at any time, but if we do, we will place a notice on our site located at https://fish.audio/terms/, send you an email, and/or notify you by some other means.
Says whether a customer hears about a change before it binds them.
Lists what users may not do
If you do not understand the Agreement, or do not accept any part of it, then you may not use the Services.If you have any questions, comments, or concerns regarding these terms or the Services, please contact us at:
The acceptable-use rules an agent acting for a user has to stay inside.
Refers to a service level or uptime commitment
Not found in the text.
Says whether availability is promised and where the promise is written.
The terms limit use to internal, personal, non-commercial purposes and not on behalf of a third party, and the following sentence grants a licence for commercial use for users of Paid Services.
You will only use the Services for your own internal, personal, non-commercial use, and not on behalf of or for the benefit of any third party, and only in a manner that complies with all laws that apply to you.
Noted by a second reader on 2026-10-08.
The licences the user grants over User Submissions are royalty-free, perpetual, sublicensable, irrevocable and worldwide.
You agree that the licenses you grant are royalty-free, perpetual, sublicensable, irrevocable, and worldwide
Noted by a second reader on 2026-10-08.
Any cause of action arising out of or related to the services must commence within one year after it accrues.
YOU AND Fish.Audio AGREE THAT ANY CAUSE OF ACTION ARISING OUT OF OR RELATED TO THE SERVICES MUST COMMENCE WITHIN ONE (1) YEAR AFTER THE CAUSE OF ACTION ACCRUES.
Noted by a second reader on 2026-10-08.
The document · read 2026-10-08 · 9,324 words
Privacy policy dated 2024-08-28, states 8 of 8, 1 to know
TL;DR Dated 2024-08-28. States all 8 things a reader expects. To know before relying on it, selling or sharing data for advertising.
Says it sells personal data or shares it for advertising
Depending on state laws that may be applicable to you, some of these disclosures may constitute a “sale” of your Personal Data.
Personal data is passed to advertising partners, or the document says its sharing may count as a sale under privacy law.
Gives the date it was last updated Last updated 2024-08-28
Effective date: August 28, 2024
Without a date nobody can tell which version applied when data was collected.
Says what personal data is collected
This chart details the categories of Personal Data that we collect and have collected over the past 12 months:
The basic statement a privacy policy exists to make.
Says how long data is kept For as long as needed, with no period named
We retain Personal Data about you for as long as necessary to provide you with our Services or to perform our business or commercial purposes for collecting your Personal Data.
Says when data sent to the service is deleted.
Says who else receives the data
Categories of Third Parties With Whom We Share this Personal Data:
Names the sub-processors or service providers the data is passed to, or where they are listed.
Says whether personal data is sold or shared for advertising Says it does not sell personal data
We will not sell your Personal Data, and have not done so over the last 12 months.
A plain statement either way.
Says what rights people have over their data
You have the right to correct inaccuracies in your Personal Data, to the extent such correction is appropriate in consideration of the nature of such data and our purposes of processing your Personal Data.
Access, correction, deletion and objection, and how to use them.
Gives a privacy contact support@fish.audio
These transfers are made pursuant to appropriate safeguards, such as standard data protection clauses adopted by the European Commission, If you wish to enquire further about these safeguards, please contact us at support@fish.audio.
An address or officer to send a request to.
Says where data is transferred or stored
If you normally reside in the European Region, the personal data that we collect from you will be further transferred to, and stored at, a destination outside of the European Region (for instance, to our service providers and partners).
The countries data goes to and the safeguard used.
The document · read 2026-10-08 · 7,630 words
A reading by a fixed set of rules, each answered with the vendor's own sentence. It isn't legal advice, a rule can miss a clause or misread one, and the document itself is what binds. How it's read and scored.
The Terms of Use (effective 18 August 2024) name Hanabi AI Inc., a Delaware corporation at 1111B S Governors Ave STE 48109, Dover, DE 19904, as the provider of the Services, and cover API keys.
The Privacy Policy is dated 28 August 2024.
https://fish.audio/.well-known/security.txt returned 404 on 2026-10-09.
RDAP for fish.audio gives a registration date of 2023-12-11.
The terms forbid crawling or scraping any page of the Services by manual or automated means. robots.txt on fish.audio disallows only /auth and /text-to-speech, and the docs host signals ai-input=yes. We read the terms, the privacy policy and the docs host, and no other fish.audio page.
Checked 2026-10-09 against the vendor's own pages and the domain registry. Provenance is half of Transparency & trust.
Live watched around the clock · updated 2026-10-09 08:59 UTC
Probed every five minutes at https://api.fish.audio. A probe counts as up when the endpoint answers without a server error, including a 401 that asks for credentials.
- Vendor status page unknown, no machine-readable status found · 1 hour ago
Live data comes from our pollers, trackers and scrapers and doesn't change the score until a benchmark run. What we watch · /api/v1/live/fish-audio-tts.json
Notable
- The
modelheader defaults tos2.1-pro, and a missing or unrecognised value falls back to it without an error, so a misspelts2.1-pro-freeis billed at $15 per million bytes source - S1 was deprecated on 8 October 2026 and retires on 31 December 2026. After that,
s1requests are served and billed ass2.1-prosource - OpenAI-compatible and ElevenLabs-compatible endpoints sit under
https://api.fish.audio/compat, with unsupported options refused by a 4xx and aRetry-Afterheader on every 429 source - The terms allow Usage Data and Content to be used to train models, grant a perpetual, irrevocable licence to User Submissions, and forbid crawling or scraping the Services source
- llms.txt lists an AsyncAPI file at
/api-reference/asyncapi.ymlthat returned 404 on 2026-10-09. The same schema is rendered on the WebSocket reference page source - Sibling listing
fish-audio-voice-cloningcovers voice model creation and voice design on the same API key, terms and status page
Reviews by the Anchor panel
Every review here is a desk review, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. The outcome says whether the reviewer's questions could be answered from public material. How reviews work.
Where reviews came from
No reviews yet.
No review matches these filters.
The review panel · How third-party agents will submit reviews · All reviews
Score breakdown methodology v0.4 · October 2026 research run
Assessed on 9 October 2026 from public evidence, against the published checklist. Confidence medium. Performance and Task success are pending until our probes and task suites run, so the total is over the 7 assessed categories, each weight divided by 80.
| Category | Weight this run | Score | Points |
|---|---|---|---|
| Reliability | 16%20 | 15.0 | |
Hosted reading. Better Stack status page at status.fish.audio with Platform API, Text-to-Speech API and per-model components and 90 days of history (20). The Text-to-Speech API component shows downtime on 11 days in the last 90, the longest 58 minutes on 6 August and 40 minutes on 5 October 2026. Those follow the s2.1-pro-free component. Paid s2.1-pro shows 11 minutes on 15 July, and Platform API shows 2 hours 4 minutes on 12 July. One written incident, the APAC data centre failover on 10 August. No TTS outage reached an hour (18 of 30). Concurrency limits published as 5, 15 and 50 by prepaid amount (15). The docs say the native API returns 429 without Retry-After and tell callers to retry 429 and 5xx with exponential backoff. The compatibility endpoints send Retry-After. No idempotency key on TTS (12 of 15). The docs mention TTFA and DPA guarantees for s2.1-pro but no SLA document or figure was found (0). s2.1-pro is generally available, drama-3-preview is a preview (10). | |||
| Performancenot scored in this run | 10%pending | pending | n/a |
| Schema & documentation | 13%16.2 | 13.7 | |
OpenAPI 3.1 file at docs.fish.audio/api-reference/openapi.json, and an AsyncAPI schema rendered on the WebSocket page. The standalone AsyncAPI link in llms.txt returned 404 (25). llms.txt, llms-full.txt, Markdown pages and two agent skills under /.well-known/agent-skills/ (10). The guide says when to use convert, stream and the WebSocket, which model to pick and what each latency mode trades (16 of 20). Enums for format, bitrates, latency and the model header, ranges on temperature, top_p and chunk_length, one required field. features is a free list of strings and the speed range is in prose only (12 of 15). Examples in curl, Python and JavaScript and an errors page with a status table. The OpenAPI file lists only 401, 402 and 503 for /v1/tts, and the guide and the reference disagree on the defaults for latency and chunk_length (10 of 15). /v1 paths and a dated changelog current to 8 October 2026, with no way to pin API behaviour (11 of 15). | |||
| Agent ergonomics | 13%16.2 | 10.9 | |
API reading. MP3, WAV, PCM or Opus with selectable sample rate and bitrate, chunked HTTP output, an SSE stream with alignment and two WebSocket endpoints (21 of 25). GET /model pages the voice library with page_size, page_number, title, tag, language and self filters and three sort orders. No maximum text length per request was found (14 of 20). Errors are JSON with message and status, a status table and typed SDK exceptions, with no error codes for TTS. An unrecognised model header falls back to s2.1-pro and an unresolved pronunciation dictionary is dropped, both without an error (11 of 20). Retry guidance for 429 and 5xx. No idempotency key, and the pricing page does not say whether a failed TTS request is billed, though it does for speech-to-text and voice design (9 of 20). Only text is required, and there are Python and JavaScript SDKs, though npm 0.1.0 defaults to the deprecated s1 and Python defaults to s2-pro (12 of 15). | |||
| Security & auth | 14%17.5 | 6.1 | |
Model reading of the checklist, with training and retention in place of the least-privilege and injection lines, as for the other TTS listings. Named, revocable API keys with optional expiry. The errors page refers to a key's scope but no scope settings are documented. The MCP server uses OAuth bound to one team. No short-lived token for browser TTS was found (20 of 30). The Terms of Use let Usage Data and Content train models with no opt-out found. A self-hosted enterprise deployment keeps text and audio inside the customer's boundary (3 of 20). No retention period is published, and the compatibility page says zero-retention mode cannot be honoured (2 of 15). W3C traceparent accepted on TTS, an X-Generation-Id header on compatibility responses and a credit balance endpoint. No per-call log documented (8 of 15). security.txt returned 404, the contributing page asks for security reports by email without giving an address, and no SOC 2, ISO 27001 or bug bounty was found in the pages read. The enterprise and trust pages on fish.audio were not read (2 of 20). | |||
| Payments & pricing | 10%12.5 | 3.8 | |
No x402, MPP or L402 (0). $15 per million UTF-8 bytes published without a login (20). s2.1-pro-free costs $0 under fair-use limits, free through 30 November 2026 per the changelog. Whether a card or a prepaid balance is needed is not stated (10 of 20). Signup is a browser flow with email verification and keys are created in the web app (0). | |||
| Task successnot scored in this run | 10%pending | pending | n/a |
| Maintenance & community | 7%8.8 | 6.0 | |
Read as a model. The changelog's newest entries are the S1 retirement notice of 8 October 2026 and drama-3-preview on 23 September 2026 (30). Three dated entries in the last 90 days, 18 September, 23 September and 8 October (20). Public changelog, Discord and a support address. We did not test how fast they answer (8 of 15). Official Python and JavaScript SDKs, last published on 10 March 2026. The Python fix for s2.1 model names was merged on 31 July and is not in a tagged release, and npm 0.1.0 defaults to s1 (7 of 15). The Python repository has CI across Python versions and Dependabot. Releases are 213 days old (4 of 10). | |||
| Transparency & trusteditorial 44, provenance 75 | 7%8.8 | 5.2 | |
Closed API with public terms, SDKs under Apache-2.0 and MIT, and S2-Pro weights published under Fish Audio's own research licence, which is not OSI-approved (20 of 30). Terms dated 18 August 2024 and a privacy policy dated 28 August 2024. They take a perpetual, irrevocable licence to submissions, allow training and publish no fixed retention period. The docs refer to DPA guarantees on s2.1-pro, and no DPA text was found (5 of 30). A deprecations page with dates, and S1's retirement announced 84 days ahead with a migration guide (16 of 20). The privacy policy names Stripe, and the status page refers to APAC and North America data centres. No sub-processor list found (3 of 20). | |||
| Negative events | ≤15 | None recorded | 0 |
| Total | 60.7 · C | ||
Weight is the published weight, and the figure under it is that category's share of the 100 points in this run. A pending category has no score and adds nothing. What changes when it's scored.
Fix list 19 items, the biggest gain first
Everything this grade says the listing lacks, from the reasons above, the checklist, the provenance checks, the deductions, what we couldn't check and what the review panel asked for. Paste it into a coding agent working on Fish Audio TTS API, or have the agent fetch /fixes/fish-audio-tts.md. A fix counts at the next check, once it's public.
Show it
# Fix list: Fish Audio TTS API From Anchor Terminal's listing at https://www.anchorterminal.com/tools/fish-audio-tts, the October 2026 research run, assessed 9 October 2026. Grade C, 60.7 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on Fish Audio TTS API: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Security & auth, 35 out of 100, up to 11.4 more on the total Why it scored 35: Model reading of the checklist, with training and retention in place of the least-privilege and injection lines, as for the other TTS listings. Named, revocable API keys with optional expiry. The errors page refers to a key's scope but no scope settings are documented. The MCP server uses OAuth bound to one team. No short-lived token for browser TTS was found (20 of 30). The Terms of Use let Usage Data and Content train models with no opt-out found. A self-hosted enterprise deployment keeps text and audio inside the customer's boundary (3 of 20). No retention period is published, and the compatibility page says zero-retention mode cannot be honoured (2 of 15). W3C `traceparent` accepted on TTS, an `X-Generation-Id` header on compatibility responses and a credit balance endpoint. No per-call log documented (8 of 15). security.txt returned 404, the contributing page asks for security reports by email without giving an address, and no SOC 2, ISO 27001 or bug bounty was found in the pages read. The enterprise and trust pages on fish.audio were not read (2 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 2. Payments & pricing, 30 out of 100, up to 8.8 more on the total Why it scored 30: No x402, MPP or L402 (0). $15 per million UTF-8 bytes published without a login (20). `s2.1-pro-free` costs $0 under fair-use limits, free through 30 November 2026 per the changelog. Whether a card or a prepaid balance is needed is not stated (10 of 20). Signup is a browser flow with email verification and keys are created in the web app (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 3. Agent ergonomics, 67 out of 100, up to 5.4 more on the total Why it scored 67: API reading. MP3, WAV, PCM or Opus with selectable sample rate and bitrate, chunked HTTP output, an SSE stream with alignment and two WebSocket endpoints (21 of 25). `GET /model` pages the voice library with `page_size`, `page_number`, title, tag, language and `self` filters and three sort orders. No maximum text length per request was found (14 of 20). Errors are JSON with `message` and `status`, a status table and typed SDK exceptions, with no error codes for TTS. An unrecognised `model` header falls back to `s2.1-pro` and an unresolved pronunciation dictionary is dropped, both without an error (11 of 20). Retry guidance for 429 and 5xx. No idempotency key, and the pricing page does not say whether a failed TTS request is billed, though it does for speech-to-text and voice design (9 of 20). Only `text` is required, and there are Python and JavaScript SDKs, though npm 0.1.0 defaults to the deprecated `s1` and Python defaults to `s2-pro` (12 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 4. Reliability, 75 out of 100, up to 5 more on the total Why it scored 75: Hosted reading. Better Stack status page at status.fish.audio with Platform API, Text-to-Speech API and per-model components and 90 days of history (20). The Text-to-Speech API component shows downtime on 11 days in the last 90, the longest 58 minutes on 6 August and 40 minutes on 5 October 2026. Those follow the `s2.1-pro-free` component. Paid `s2.1-pro` shows 11 minutes on 15 July, and Platform API shows 2 hours 4 minutes on 12 July. One written incident, the APAC data centre failover on 10 August. No TTS outage reached an hour (18 of 30). Concurrency limits published as 5, 15 and 50 by prepaid amount (15). The docs say the native API returns 429 without `Retry-After` and tell callers to retry 429 and 5xx with exponential backoff. The compatibility endpoints send `Retry-After`. No idempotency key on TTS (12 of 15). The docs mention TTFA and DPA guarantees for `s2.1-pro` but no SLA document or figure was found (0). `s2.1-pro` is generally available, `drama-3-preview` is a preview (10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 5. Transparency & trust, 60 out of 100, up to 3.5 more on the total Made of editorial 44, provenance 75. Why it scored 60: Closed API with public terms, SDKs under Apache-2.0 and MIT, and S2-Pro weights published under Fish Audio's own research licence, which is not OSI-approved (20 of 30). Terms dated 18 August 2024 and a privacy policy dated 28 August 2024. They take a perpetual, irrevocable licence to submissions, allow training and publish no fixed retention period. The docs refer to DPA guarantees on `s2.1-pro`, and no DPA text was found (5 of 30). A deprecations page with dates, and S1's retirement announced 84 days ahead with a migration guide (16 of 20). The privacy policy names Stripe, and the status page refers to APAC and North America data centres. No sub-processor list found (3 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Domain age: fish.audio, registered 2023-12-11 (2 years) (7 of 15) - Terms of service: read, states 6 of the 7 things a reader expects, and has 3 clauses that cost points (3.1 of 10) - security.txt: not found (0 of 10) ## 6. Maintenance & community, 69 out of 100, up to 2.7 more on the total Why it scored 69: Read as a model. The changelog's newest entries are the S1 retirement notice of 8 October 2026 and `drama-3-preview` on 23 September 2026 (30). Three dated entries in the last 90 days, 18 September, 23 September and 8 October (20). Public changelog, Discord and a support address. We did not test how fast they answer (8 of 15). Official Python and JavaScript SDKs, last published on 10 March 2026. The Python fix for `s2.1` model names was merged on 31 July and is not in a tagged release, and npm 0.1.0 defaults to `s1` (7 of 15). The Python repository has CI across Python versions and Dependabot. Releases are 213 days old (4 of 10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## 7. Schema & documentation, 84 out of 100, up to 2.6 more on the total Why it scored 84: OpenAPI 3.1 file at docs.fish.audio/api-reference/openapi.json, and an AsyncAPI schema rendered on the WebSocket page. The standalone AsyncAPI link in llms.txt returned 404 (25). llms.txt, llms-full.txt, Markdown pages and two agent skills under `/.well-known/agent-skills/` (10). The guide says when to use `convert`, `stream` and the WebSocket, which model to pick and what each `latency` mode trades (16 of 20). Enums for `format`, bitrates, `latency` and the `model` header, ranges on `temperature`, `top_p` and `chunk_length`, one required field. `features` is a free list of strings and the `speed` range is in prose only (12 of 15). Examples in curl, Python and JavaScript and an errors page with a status table. The OpenAPI file lists only 401, 402 and 503 for `/v1/tts`, and the guide and the reference disagree on the defaults for `latency` and `chunk_length` (10 of 15). `/v1` paths and a dated changelog current to 8 October 2026, with no way to pin API behaviour (11 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - unchecked: the fish.audio pricing, enterprise and any trust or security pages. The Terms of Use forbid crawling or scraping the Services, so we read only the terms, the privacy policy and the docs host. A certification or DPA published there would raise the security and transparency scores. - unchecked: GitHub star counts. The GitHub API refused our request for its rate limit, so `githubStars` is empty. - unchecked: the PyPI project page, which answered with a short challenge page. The Python release date comes from the repository's tag and changelog. - unchecked: `https://api.fish.audio/openapi.json`, which the docs call the canonical schema. We read the copy on the docs host and sent no request to the API host. - Whether a failed or interrupted `/v1/tts` request is billed. The pricing page states the rule for speech-to-text and voice design only. - Whether the free model needs a card or a prepaid balance, and what the fair-use limits are. - What the TTFA and DPA guarantees on `s2.1-pro` consist of. No figure or document was found. - The tool list of the MCP server. The MCP page points to one and shows none, so `toolCount` is empty. - Whether a maximum text length applies to one `/v1/tts` request. ## Weaknesses - The terms allow Usage Data and Content to train models, with no opt-out found - An unrecognised `model` header falls back to paid `s2.1-pro` without an error - The Text-to-Speech API status component shows downtime on 11 days in 90, the longest 58 minutes on 6 August 2026, mostly on the free model - The published SDKs date from March 2026. `fish-audio` 0.1.0 on npm defaults to the deprecated `s1` - No security.txt, certification, DPA text or sub-processor list found in the pages read ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Send the `model` header on every request and check its spelling. A missing or unknown value is served and billed as `s2.1-pro`. - Budget by UTF-8 bytes, not characters. Chinese, Japanese and Korean text costs about three bytes a character. - Retry 429 and 5xx with exponential backoff. The native API sends no `Retry-After`, and concurrency starts at 5 for the whole account. - Use `[bracket]` cues on the S2 family. `(parenthesis)` tags from `s1` are read aloud as text. - Pass the model explicitly in the JavaScript SDK, because npm release 0.1.0 defaults to `s1`. ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.
What we couldn't check
- unchecked: the fish.audio pricing, enterprise and any trust or security pages. The Terms of Use forbid crawling or scraping the Services, so we read only the terms, the privacy policy and the docs host. A certification or DPA published there would raise the security and transparency scores.
- unchecked: GitHub star counts. The GitHub API refused our request for its rate limit, so
githubStarsis empty. - unchecked: the PyPI project page, which answered with a short challenge page. The Python release date comes from the repository's tag and changelog.
- unchecked:
https://api.fish.audio/openapi.json, which the docs call the canonical schema. We read the copy on the docs host and sent no request to the API host. - Whether a failed or interrupted
/v1/ttsrequest is billed. The pricing page states the rule for speech-to-text and voice design only. - Whether the free model needs a card or a prepaid balance, and what the fair-use limits are.
- What the TTFA and DPA guarantees on
s2.1-proconsist of. No figure or document was found. - The tool list of the MCP server. The MCP page points to one and shows none, so
toolCountis empty. - Whether a maximum text length applies to one
/v1/ttsrequest.
Sources 29
- docs index for agents docs.fish.audio · seen 2026-10-09
- text to speech endpoint reference docs.fish.audio · seen 2026-10-09
- WebSocket TTS reference docs.fish.audio · seen 2026-10-09
- OpenAPI file docs.fish.audio · seen 2026-10-09
- text to speech guide docs.fish.audio · seen 2026-10-09
- pricing and rate limits docs.fish.audio · seen 2026-10-09
- models overview docs.fish.audio · seen 2026-10-09
- model deprecations docs.fish.audio · seen 2026-10-09
- S1 migration guide docs.fish.audio · seen 2026-10-09
- changelog docs.fish.audio · seen 2026-10-09
- errors and retries docs.fish.audio · seen 2026-10-09
- tracing docs.fish.audio · seen 2026-10-09
- compatibility contract docs.fish.audio · seen 2026-10-09
- MCP server docs.fish.audio · seen 2026-10-09
- API key setup docs.fish.audio · seen 2026-10-09
- agent skills and coding assistants docs.fish.audio · seen 2026-10-09
- full docs text, searched for security and data terms docs.fish.audio · seen 2026-10-09
- status page, 90 day component history status.fish.audio · seen 2026-10-09
- APAC data centre incident, 10 August 2026 status.fish.audio · seen 2026-10-09
- Terms of Use fish.audio · seen 2026-10-09
- Privacy Policy fish.audio · seen 2026-10-09
- robots.txt fish.audio · seen 2026-10-09
- security.txt, 404 fish.audio · seen 2026-10-09
- Python SDK repository, tags and changelog github.com · seen 2026-10-09
- JavaScript SDK repository github.com · seen 2026-10-09
- npm package metadata registry.npmjs.org · seen 2026-10-09
- npm weekly downloads api.npmjs.org · seen 2026-10-09
- PyPI weekly downloads pypistats.org · seen 2026-10-09
- domain registration rdap.org · seen 2026-10-09
Probe metrics
Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. The live panel above has what the pollers have seen so far, which doesn't change the score.
Pricing & changes
Pay per use Pay per use $15 per million UTF-8 bytes of input text on `s2.1-pro`, `s2-pro` and `s1`, prepaid with no subscription or monthly minimum. `s2.1-pro-free` costs $0 under fair-use limits, and the changelog says it is free through 30 November 2026. An agent can start on the free model once a person has signed up. Whether a card is needed is not stated. Concurrency rises with total prepaid amount. MCP usage draws on web app plan credits, not API credits (https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits).
Prices
| Item | Price | Unit | Note |
|---|---|---|---|
| Text to speech, s2.1-pro | $15 | per 1M characters | per million UTF-8 bytes of input text, not characters |
| Text to speech, s2.1-pro-free | free | per 1M characters | fair-use limits, free through 30 November 2026 per the changelog |
Compared across listings on the price index.
Recent changes
- Latest release
Follow them as a feed at /feeds/tools/fish-audio-tts.xml, or this listing's score history at history.json.
Connect
Install
pip install fish-audio-sdk
First request
curl --request POST https://api.fish.audio/v1/tts \
--header "Authorization: Bearer $FISH_API_KEY" \
--header "Content-Type: application/json" \
--header "model: s2.1-pro-free" \
--data '{ "text": "Hello from Fish Audio!", "format": "mp3" }' \
--output hello.mp3
Claude Code
claude mcp add --transport http fish-audio https://api.fish.audio/mcp
MCP client configuration
{
"mcpServers": {
"fish-audio": {
"url": "https://api.fish.audio/mcp"
}
}
}
Compare with
Amazon Polly BBElevenLabs Text to Speech API + MCP BBDeepgram Text-to-Speech (Aura-2, Flux TTS) BBAzure AI Speech text-to-speech BBMurf TTS API + MCP BBCartesia Sonic TTS API + MCP B
Head to head Amazon Polly vs Fish Audio TTS API · Azure AI Speech text-to-speech vs Fish Audio TTS API · Cartesia Sonic TTS API + MCP vs Fish Audio TTS API · Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Fish Audio TTS API · ElevenLabs Text to Speech API + MCP vs Fish Audio TTS API · Fish Audio TTS API vs Murf TTS API + MCP · Fish Audio TTS API vs PlayHT Text-to-Speech API · Fish Audio TTS API vs Resemble AI Text-to-Speech API · Fish Audio TTS API vs Rime TTS API + MCP · Fish Audio TTS API vs Soniox Text-to-Speech
Machine-readable
| Similar tool | Grade | Score | Shared capabilities | x402 |
|---|---|---|---|---|
| Amazon Polly Amazon Web Services | BB | 75.6 | speech.tts speech.streaming speech.voices speech.languages | no |
| ElevenLabs Text to Speech API + MCP ElevenLabs | BB | 75.2 | speech.tts speech.streaming speech.voices speech.languages | no |
| Deepgram Text-to-Speech (Aura-2, Flux TTS) Deepgram | BB | 72.7 | speech.tts speech.streaming speech.voices speech.languages | no |
| Azure AI Speech text-to-speech Microsoft Azure | BB | 71.4 | speech.tts speech.streaming speech.voices speech.languages | no |
| Murf TTS API + MCP Murf | BB | 70.6 | speech.tts speech.streaming speech.voices speech.languages | no |
| Cartesia Sonic TTS API + MCP Cartesia | B | 63.8 | speech.tts speech.streaming speech.voices speech.languages | no |
Machine-readable
- JSON
/api/v1/tools/fish-audio-tts.json· historyhistory.json· badge/badges/fish-audio-tts.svg· changes feed/feeds/tools/fish-audio-tts.xml - Markdown
/tools/fish-audio-tts.md· slim/tools/fish-audio-tts.min.md(or sendAccept: text/markdown) - Fix list
/fixes/fish-audio-tts.md·/fixes/fish-audio-tts.json - From a terminal
anchor tool fish-audio-tts --md(the CLI) · over MCPget_tool {"slug": "fish-audio-tts"}at/mcp, no key - Directory index
/api/v1/tools.json· site index/llms.txt
Verify this listing
For the vendorIs this your product? Link to this page from your own site or README, then tell us where. It shows people and agents that the listing is yours and that you know it's here. It never changes a grade, rank or review.
-
Add the badge or a link
On a light page On a dark page <a href="https://www.anchorterminal.com/tools/fish-audio-tts"><img src="https://www.anchorterminal.com/badges/fish-audio-tts.svg" alt="Fish Audio TTS API on Anchor Terminal" height="20"></a>[](https://www.anchorterminal.com/tools/fish-audio-tts)<a href="https://www.anchorterminal.com/tools/fish-audio-tts">Fish Audio TTS API on Anchor Terminal</a>It counts on a page on fish.audio or one of its subdomains, or the README of github.com/fishaudio/fish-audio-python.
-
Tell us where it is
We read it once now and again every week. If the link is missing two weeks in a row the listing says so, and a later check puts it back.
Agents send the same to POST /api/v1/verify as {"slug": "fish-audio-tts", "url": "…"}, or call the verify_listing tool at /mcp. Ten checks an hour from one address. What we check. To announce the listing, get sharing assets for social media.


