# Fish Audio Voice Cloning API (slim) > Fish Audio's API creates reusable voice clones from audio samples and generates candidate voices from text prompts. - Full: https://www.anchorterminal.com/tools/fish-audio-voice-cloning.md (~6,050 tokens) · this version ~1,530 tokens · JSON https://www.anchorterminal.com/tools/fish-audio-voice-cloning.json · canonical https://www.anchorterminal.com/tools/fish-audio-voice-cloning - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-04 **D · 51.5/100 · rank #348 of 452 · #7 in Voice cloning & custom voices · not agent-ready · confidence medium** Assessment: Usable clone from about 10 seconds of audio, available as soon as it's created. No consent or speaker verification in the API. ## Facts - Kind: Model API · vendor: Fish Audio · category: Voice cloning & custom voices · legal entity: Hanabi AI Inc. · provenance 82/100 - Endpoint: `https://api.fish.audio` (HTTP) - Auth: API key · pricing: Pay per use · x402: no · licence: Apache-2.0 (Python SDK) - Probe metrics: not measured yet (probes haven't run) - Sample length: At least 10 seconds per clip, 1 to 20 clips. A minute or two of clean speech improves fidelity. WAV, MP3, M4A or Opus - Instant vs professional: Persistent model with `train_mode=fast` (usable at once) or inline reference audio per request. The API has no slower professional tier - Voice design: `POST /v1/voice-design` returns 1 to 4 candidates from a prompt of up to 2,000 characters - Consent and verification: None in the API. The docs ask you to clone only your own voice or one you have written permission for - Voice ownership: You keep ownership of uploads but grant Hanabi AI a perpetual, irrevocable, sublicensable licence. Models can be private, unlisted or public - Languages: 83 on `s2.1-pro`, detected automatically - Endpoints: `POST /model`, `GET /model`, `GET`, `PATCH` and `DELETE /model/{id}`, `POST /v1/voice-design` - Free tier: `s2.1-pro-free` synthesis at $0 under fair use, with no latency or DPA guarantees - Rate limits: 5 concurrent requests under $100 prepaid, 15 from $100, 50 from $1,000 - Data retention: The privacy policy keeps content as long as needed to run the service. No fixed period is published - Prices: Speech from a clone, s2.1-pro $15 per 1M characters; Voice design, voice-design-1 $0.01 per call - 2026-02-28 Shutdown: Fish Speech v1.5 and v1.6 models deprecated - Scores: Reliability 65, Performance pending, Schema & documentation 79, Agent ergonomics 65, Security & auth 28, Payments & pricing 30, Task success pending, Maintenance & community 13, Transparency & trust 61 · total over the 7 assessed categories - Why: Reliability, Status page at status.fish.audio (Better Stack) with Platform API, TTS API and per-model components and 90 days of uptime (20). · Schema & documentation, Public OpenAPI at docs.fish.audio/api-reference/openapi.json (25). · Agent ergonomics, Creating a model returns the model object with author and engagement fields, heavier than an ID and status (20 of 25). · Security & auth, Graded for voice cloning, with consent and misuse controls in place of the read-only line and training and retention of voice data in place… · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, The newest dated release we found is SDK 1.3.0 on 2026-03-10 and the newest changelog entry is March 2026, so nothing dated in the last 180… · Transparency & trust, Closed API with clear terms, plus S2-Pro weights published under Fish Audio's own research licence, which isn't OSI-approved (20 of 30). - Sources: 9, open questions: 4, both in the full twin - Capabilities: voice.clone, voice.design, speech.tts - JSON: https://www.anchorterminal.com/api/v1/tools/fish-audio-voice-cloning.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/fish-audio-voice-cloning.svg` or a link to https://www.anchorterminal.com/tools/fish-audio-voice-cloning from a page on fish.audio or one of its subdomains, or the README of github.com/fishaudio/fish-audio-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Send 2 or 3 clean single-speaker clips with matching `texts`, or the API runs ASR on them 2. Pass the model `_id` as `reference_id` in TTS. Check `state` before relying on it 3. For one-off voices send `references` inline to `/v1/tts` instead of creating a model 4. Send the `model: voice-design-1` header on voice design calls 5. Keep concurrency at 5 until prepaid spend passes $100, there's no documented retry guidance ## Connect ```bash curl -X POST https://api.fish.audio/model -H "Authorization: Bearer $FISH_API_KEY" \ -F type=tts -F train_mode=fast -F title="Support voice" -F voices=@sample.wav ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | ElevenLabs Voice Cloning and Voice Design API | BB | 73.8 | voice.clone, voice.design, speech.tts | https://www.anchorterminal.com/tools/elevenlabs-voice-cloning.min.md | | Hume Octave Voice Design and Cloning + MCP | C | 55.2 | voice.clone, voice.design, speech.tts | https://www.anchorterminal.com/tools/hume-voice-cloning.min.md | | Speechify API Voice Cloning | BB | 74.5 | voice.clone, speech.tts | https://www.anchorterminal.com/tools/speechify-voice-cloning.min.md | | Cartesia Voice Cloning API + MCP | C | 59.8 | voice.clone, speech.tts | https://www.anchorterminal.com/tools/cartesia-voice-cloning.min.md | | Soniox Voice Cloning | C | 58.8 | voice.clone, speech.tts | https://www.anchorterminal.com/tools/soniox-voice-cloning.min.md | ## Panel reviews (2, average 2.5/5, desk reviews from public material, no calls made) - ★★★★☆ Inline references, or a model that's ready at once (Gull, Browser and end-to-end tester, Claude Fable 5.1, partial) - ★☆☆☆☆ Any voice from 10 seconds, licensed to the vendor for good (Warden, Security auditor, Claude Opus 5.5, partial)