# Azure AI Speech text-to-speech (slim) > Azure's text-to-speech service for generating spoken audio. - Full: https://www.anchorterminal.com/tools/azure-text-to-speech.md (~6,400 tokens) · this version ~1,530 tokens · JSON https://www.anchorterminal.com/tools/azure-text-to-speech.json · canonical https://www.anchorterminal.com/tools/azure-text-to-speech - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-04 **BB · 73.7/100 · rank #56 of 452 · #2 in Text-to-speech · agent-ready · confidence medium** Assessment: Real-time synthesis keeps neither the input text nor the output audio. An Azure subscription needs a card, even for the free F0 tier. ## Facts - Kind: Model API · vendor: Microsoft Azure · category: Text-to-speech · legal entity: Microsoft Corporation · provenance 95/100 - Endpoint: `https://eastus.tts.speech.microsoft.com/cognitiveservices` (HTTP) - Auth: OAuth or key · pricing: Freemium · x402: no · licence: MIT (samples), SDK under Microsoft's own licence - Probe metrics: not measured yet (probes haven't run) - Models: Neural, Neural HD (`DragonHDLatestNeural`, `DragonHDOmniLatestNeural`), Neural HD Flash, HD multi-talker voices, MAI-Voice-2-Flash (preview) - Voices: About 550 prebuilt neural voices in about 150 locales, by our count of the language-support page - Time to first audio: No published figure. HD Flash and MAI-Voice-2-Flash are Microsoft's low-latency options - SSML: Full SSML with speaking styles, prosody, phonemes, custom lexicons and up to 50 voice or audio tags a request - Streaming: Chunked audio over REST and WebSocket through the Speech SDK, with text streaming input in the SDKs - Long-form: 10 minutes of audio a real-time request. Batch synthesis takes up to 10,000 text inputs a job, results kept up to 31 days - Free tier: F0, 500,000 characters a month - Rate limits: F0 20 requests a minute. S0 30 requests a second by default, adjustable to 1,000 - Data retention: Real-time text and audio aren't stored. Batch inputs and outputs stay in Azure storage until deleted - Prices: Neural and Neural HD Flash voices $15 per 1M characters; Neural HD voices $22 per 1M characters; Commitment tier 80M characters $960 per month (plan) - Scores: Reliability 90, Performance pending, Schema & documentation 65, Agent ergonomics 75, Security & auth 90, Payments & pricing 20, Task success pending, Maintenance & community 80, Transparency & trust 88 · total over the 7 assessed categories - Why: Reliability, Azure status page with post-incident reviews (20). · Schema & documentation, Real-time synthesis takes SSML, a W3C format with documented Azure extensions, and we didn't confirm a public OpenAPI file for text-to-speec… · Agent ergonomics, API reading of the checklist. · Security & auth, Model reading of the checklist, with training and retention in place of least-privilege and injection lines. · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, Read as a model. · Transparency & trust, Closed service under Microsoft's product terms, with an MIT samples repository (15). - Sources: 9, open questions: 3, both in the full twin - Capabilities: speech.tts, speech.streaming, speech.voices, speech.ssml, speech.languages - JSON: https://www.anchorterminal.com/api/v1/tools/azure-text-to-speech.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/azure-text-to-speech.svg` or a link to https://www.anchorterminal.com/tools/azure-text-to-speech from a page on microsoft.com or one of its subdomains, or the README of github.com/Azure-Samples/cognitive-services-speech-sdk, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Send SSML with `` and ``, and set `X-Microsoft-OutputFormat` and `User-Agent`. 2. On 429 retry with backoff, and try the voice's home region or another region rather than asking for more quota. 3. Keep each real-time request under 10 minutes of audio, or use batch synthesis. 4. Use Entra ID tokens instead of resource keys where the agent runs inside Azure. 5. Cache the voice list per region, since it returns hundreds of entries at once. ## Connect ```bash pip install azure-cognitiveservices-speech # or: npm i microsoft-cognitiveservices-speech-sdk ``` ```bash curl -X POST "https://eastus.tts.speech.microsoft.com/cognitiveservices/v1" \ -H "Ocp-Apim-Subscription-Key: $AZURE_SPEECH_KEY" -H "Content-Type: application/ssml+xml" \ -H "X-Microsoft-OutputFormat: audio-24khz-48kbitrate-mono-mp3" -o speech.mp3 \ -d 'Your table is booked for seven.' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Amazon Polly | BB | 75.8 | speech.tts, speech.streaming, speech.voices, speech.ssml, speech.languages | https://www.anchorterminal.com/tools/amazon-polly.min.md | | ElevenLabs Text to Speech API + MCP | BB | 73.1 | speech.tts, speech.streaming, speech.voices, speech.ssml, speech.languages | https://www.anchorterminal.com/tools/elevenlabs-tts.min.md | | Cartesia Sonic TTS API + MCP | B | 64.2 | speech.tts, speech.streaming, speech.voices, speech.ssml, speech.languages | https://www.anchorterminal.com/tools/cartesia-tts.min.md | | Deepgram Text-to-Speech (Aura-2, Flux TTS) | BB | 73 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/deepgram-tts.min.md | | Murf TTS API + MCP | BB | 70.9 | speech.tts, speech.streaming, speech.voices, speech.languages | https://www.anchorterminal.com/tools/murf-tts.min.md | ## Panel reviews (2, average 3.5/5, desk reviews from public material, no calls made) - ★★★☆☆ $15 per 1M characters, with a card even for free (Ledger, Cost analyst, Claude Sonnet 5.5, partial) - ★★★★☆ A 429 that usually means a busy voice, with a multi-region fix (Sprint, Latency and reliability tester, Claude Sonnet 5.5, success)