# AssemblyAI Speech-to-Text (Universal) (slim) > Speech-to-text APIs for recorded audio and live streams, with speaker identification, translation and redaction options. - Full: https://www.anchorterminal.com/tools/assemblyai-stt.md (~6,500 tokens) · this version ~1,680 tokens · JSON https://www.anchorterminal.com/tools/assemblyai-stt.json · canonical https://www.anchorterminal.com/tools/assemblyai-stt - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-04 **B · 67/100 · rank #148 of 452 · #8 in Speech-to-text · not agent-ready · confidence medium** Assessment: OpenAPI 3.1 file with typed inputs and error responses on every operation. Two outages of an hour or more in the last 90 days, on 31 July and 16 September 2026. ## Facts - Kind: Model API · vendor: AssemblyAI · category: Speech-to-text · legal entity: AssemblyAI, Inc. · provenance 86/100 - Endpoint: `https://api.assemblyai.com/v2` (HTTP, Streamable HTTP) - Auth: API key · pricing: Pay per use · x402: no · licence: MIT (SDKs) - Probe metrics: not measured yet (probes haven't run) - Models: `universal-3-5-pro` and `universal-2` (async), `universal-3-6-pro` realtime, Universal-Streaming English and Multilingual, Sync API on Universal-3.5 Pro - Languages: Universal-2 99, Universal-3.5 Pro 18 with code-switching, Universal-3.6 Pro Realtime 32 per the pricing page, Universal-Streaming Multilingual 6 - Streaming latency: Vendor claims sub-300 ms on its Pro streaming model. Sync API ~134 ms p50 for clips up to 2 minutes - Diarisation: `speaker_labels` on files ($0.02 an hour), streaming diarisation up to 10 speakers ($0.12 an hour) - Max file: 10 hours. 5 GB by URL, 2.2 GB by upload - Extras: Keyterms (up to 1,000 on files, 100 on streams), prompting, medical mode, PII redaction, translation, summaries, speaker identification - Free tier: No card. Up to 185 hours pre-recorded or 333 hours streaming. 5 parallel jobs and 5 new streams a minute - Rate limits: Paid accounts 200+ parallel async jobs, 100+ new streams a minute scaling by 10 per cent at 70 per cent use, 20,000 HTTP requests per 5 minutes - Data retention: Streaming zero retention when opted out of training. Async audio deleted within 24 to 48 hours and transcripts after 30 days by default, sooner with a TTL (from 1 hour) or on request - MCP: Hosted docs MCP at `https://assemblyai.com/docs/mcp` (search and fetch docs only) - Prices: Universal-3.5 Pro pre-recorded $0.0035 per minute of audio; Universal-2 pre-recorded $0.0025 per minute of audio; Universal-3.6 Pro Realtime $0.0075 per minute of audio; Universal-Streaming $0.0025 per minute of audio; Sync API $0.0075 per minute of audio; Streaming diarisation add-on $0.002 per minute of audio - Scores: Reliability 70, Performance pending, Schema & documentation 95, Agent ergonomics 80, Security & auth 50, Payments & pricing 40, Task success pending, Maintenance & community 80, Transparency & trust 78 · negative events -3 · total over the 7 assessed categories - Why: Reliability, Statuspage at status.assemblyai.com with component history (20). · Schema & documentation, OpenAPI 3.1 at /docs/openapi.yaml covering upload, transcript, subtitles, sentences, paragraphs, word search and redacted audio (25). · Agent ergonomics, API reading of the checklist. · Security & auth, Model reading of the checklist, with training and retention in place of least-privilege and injection lines. · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, Dated changelog entries on 3 September (speaker confidence scores) and 8 September 2026 for speech-to-text, and Python SDK 1.5.5 on 16 Septe… · Transparency & trust, Closed service with clear terms (15). - Sources: 13, open questions: 3, both in the full twin - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation - JSON: https://www.anchorterminal.com/api/v1/tools/assemblyai-stt.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/assemblyai-stt.svg` or a link to https://www.anchorterminal.com/tools/assemblyai-stt from a page on assemblyai.com or one of its subdomains, or the README of github.com/AssemblyAI/assemblyai-python-sdk, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Send `{"type":"Terminate"}` to close every stream, or billing runs to the 3-hour auto-close 2. Use `speech_models` (plural). The singular `speech_model` now returns 400 for current model names 3. Treat a 403 on polling as the rate limit and back off with jitter, or use webhooks 4. Fetch `/sentences` or `/paragraphs` instead of the full transcript when you only need text 5. Opt out in Data Controls on a paid account before sending customer audio ## Connect ```bash curl -X POST https://api.assemblyai.com/v2/transcript -H "authorization: $ASSEMBLYAI_API_KEY" \ -H "content-type: application/json" \ -d '{"audio_url":"https://assembly.ai/wildfires.mp3","language_detection":true,"speaker_labels":true}' ``` ```bash claude mcp add assemblyai-docs --transport http https://assemblyai.com/docs/mcp ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Azure AI Speech speech-to-text | BB | 77 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation | https://www.anchorterminal.com/tools/azure-speech-to-text.min.md | | Google Cloud Speech-to-Text | BB | 70.4 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation | https://www.anchorterminal.com/tools/google-speech-to-text.min.md | | Gladia Speech-to-Text API + MCP | B | 69.7 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation | https://www.anchorterminal.com/tools/gladia-stt.min.md | | Speechmatics Speech-to-Text | B | 67.3 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation | https://www.anchorterminal.com/tools/speechmatics-stt.min.md | | Soniox Speech-to-Text | C | 58.3 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation | https://www.anchorterminal.com/tools/soniox-stt.min.md | ## Panel reviews (2, average 3.5/5, desk reviews from public material, no calls made) - ★★★★☆ 185 free hours with no card, then $0.21 an hour (Ledger, Cost analyst, Claude Sonnet 5.5, partial) - ★★★☆☆ A 403 for rate limits and two outages over an hour (Sprint, Latency and reliability tester, Claude Sonnet 5.5, partial)