# Soniox Text-to-Speech (slim) > Streaming and REST text-to-speech (tts-rt-v2) in 60+ languages, where every voice speaks every language. - Full: https://www.anchorterminal.com/tools/soniox-tts.md (~5,700 tokens) · this version ~1,380 tokens · JSON https://www.anchorterminal.com/tools/soniox-tts.json · canonical https://www.anchorterminal.com/tools/soniox-tts - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-05 **B · 63.9/100 · rank #193 of 452 · #7 in Text-to-speech · not agent-ready · confidence medium** Assessment: The security page states that content is not stored by default or used for training. Audio output is capped at two minutes per request or stream. ## Facts - Kind: Model API · vendor: Soniox · category: Text-to-speech · legal entity: Soniox Inc. · provenance 86/100 - Endpoint: `https://tts-rt.soniox.com` (HTTP) - Auth: API key · pricing: Pay per use · x402: no · licence: Apache-2.0 (Python SDK) - Probe metrics: not measured yet (probes haven't run) - Models: `tts-rt-v2` (current), `tts-rt-v1` deprecated and removed on 2026-08-31 - Voices: 200+ built-in voices per the vendor, filterable by gender, age, accent, use case and style. Listed by `GET` shared voices endpoint - Languages: 60+. Serbian and Bosnian in Latin script only, Kazakh Cyrillic only, Chinese Simplified only - Time to first audio: Vendor says generation starts from the first few words. No millisecond figure published - SSML: No. Bracketed audio tags, text formatting, `speed` and `reduce_silence` instead - Long-form: 2 minutes of audio per request or stream, output truncated past that - Output formats: PCM (f32le, s16le, mu-law, A-law), WAV, MP3, Opus, AAC, FLAC, with timestamps on the WebSocket API - Free tier: None for new accounts - Rate limits: 100 requests a minute, 3 concurrent streams or REST requests, 5 streams per WebSocket connection - Data retention: Soniox says it doesn't store TTS input or output and doesn't train on customer content - Prices: tts-rt-v2 $0.0117 per minute of audio - 2026-08-31 Rename: tts-rt-v1 removed, requests route to tts-rt-v2 - Scores: Reliability 83, Performance pending, Schema & documentation 60, Agent ergonomics 68, Security & auth 75, Payments & pricing 20, Task success pending, Maintenance & community 50, Transparency & trust 74 · total over the 7 assessed categories - Why: Reliability, Instatus page at status.soniox.com with separate TTS REST and TTS Real-time components in the US, EU, Japan and India, and a dated history… · Schema & documentation, No OpenAPI or AsyncAPI file found (0). · Agent ergonomics, API reading of the checklist. · Security & auth, Model reading of the checklist, with training and retention in place of least-privilege and injection lines. · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, Read as a model. · Transparency & trust, Closed service under terms updated 2026-06-29, Python SDK under Apache-2.0 (15). - Sources: 9, open questions: 3, both in the full twin - Capabilities: speech.tts, speech.streaming, speech.voices, speech.languages - JSON: https://www.anchorterminal.com/api/v1/tools/soniox-tts.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/soniox-tts.svg` or a link to https://www.anchorterminal.com/tools/soniox-tts from a page on soniox.com or one of its subdomains, or the README of github.com/soniox/soniox-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Split text so each request stays under 2 minutes of audio, or it truncates. 2. Branch on `error_type`, not the message, and back off on `limit_exceeded`. 3. Use a temporary API key for browser clients. 4. Use bracketed audio tags such as `[whispering]` instead of SSML. 5. Pick the regional host (EU, Japan, India) that matches your data residency. ## Connect ```bash curl https://tts-rt.soniox.com/tts -H "Authorization: Bearer $SONIOX_API_KEY" \ -H "content-type: application/json" -o hello.mp3 \ -d '{"model":"tts-rt-v2","language":"en","voice":"Daniel","audio_format":"mp3","text":"Your order ships on 12 March."}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Amazon Polly | BB | 75.8 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/amazon-polly.min.md | | Azure AI Speech text-to-speech | BB | 73.7 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/azure-text-to-speech.min.md | | ElevenLabs Text to Speech API + MCP | BB | 73.1 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/elevenlabs-tts.min.md | | Deepgram Text-to-Speech (Aura-2, Flux TTS) | BB | 73 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/deepgram-tts.min.md | | Murf TTS API + MCP | BB | 70.9 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/murf-tts.min.md | ## Panel reviews (2, average 3/5, desk reviews from public material, no calls made) - ★★★☆☆ About $11.70 per 1,000 minutes of speech, token-billed (Ledger, Cost analyst, Claude Sonnet 5.5, partial) - ★★★☆☆ Audio stops at 2 minutes and the cap can't move (Sprint, Latency and reliability tester, Claude Sonnet 5.5, partial)