# ElevenLabs Scribe Speech to Text API (slim) > ElevenLabs' speech-to-text service for audio transcription. - Full: https://www.anchorterminal.com/tools/elevenlabs-scribe.md (~6,050 tokens) · this version ~1,430 tokens · JSON https://www.anchorterminal.com/tools/elevenlabs-scribe.json · canonical https://www.anchorterminal.com/tools/elevenlabs-scribe - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-04 **B · 69/100 · rank #117 of 452 · #6 in Speech-to-text · not agent-ready · confidence medium** Assessment: $0.22 an hour for batch with 90+ languages and diarisation to 32 speakers. Audio may be used for training unless the account opts out, and the opt-out isn't retroactive. ## Facts - Kind: Model API · vendor: ElevenLabs · category: Speech-to-text · legal entity: Eleven Labs Inc. · provenance 92/100 - Endpoint: `https://api.elevenlabs.io/v1` (HTTP, stdio) - Auth: API key · pricing: Freemium · x402: no · licence: MIT (SDKs) - Probe metrics: not measured yet (probes haven't run) - Models: `scribe_v2` (batch), `scribe_v2_realtime` (WebSocket), `scribe_v2_medical` (batch, clinical). `scribe_v1` is deprecated - Languages: 90+, with automatic language detection - Streaming latency: Vendor claims ~150 ms to partial transcripts on Scribe v2 Realtime, excluding network latency - Diarisation: Up to 32 speakers on batch, plus an agent and customer role mode - Max file: 3 GB upload, 10 hours (combined across channels in multichannel mode) - Extras: Word timestamps, audio event tags, keyterm prompting up to 1,000 terms, 65 entity types, no-verbatim mode, transcript editing - Free tier: About 4.5 hours of batch or 2.5 hours of realtime a month - Rate limits: Batch concurrency 8 on Free to 60 on Scale, realtime 6 to 45. Files over 8 minutes are split into up to 4 parallel chunks - Data retention: Transcripts stored by default. `enable_logging=false` gives zero retention on Enterprise only - MCP: Not on the hosted MCP server. The deprecated local `elevenlabs-mcp` has `speech_to_text` - Prices: Scribe v2 batch $0.0037 per minute of audio; Scribe v2 Realtime $0.0065 per minute of audio; Entity detection add-on $0.0012 per minute of audio; Keyterm prompting add-on $0.0008 per minute of audio - Scores: Reliability 70, Performance pending, Schema & documentation 95, Agent ergonomics 75, Security & auth 55, Payments & pricing 40, Task success pending, Maintenance & community 75, Transparency & trust 71 · total over the 7 assessed categories - Why: Reliability, incident.io status page with a Speech to Text component and a dated history (20). · Schema & documentation, OpenAPI file at api.elevenlabs.io/openapi.json (25). · Agent ergonomics, API reading of the checklist. · Security & auth, Model reading of the checklist, with training and retention in place of least-privilege and injection lines. · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, Dated changelog entries on 28, 23 and 21 September 2026, though we didn't check which touch speech-to-text (30). · Transparency & trust, Closed service with clear terms and MIT SDKs (15). - Sources: 9, open questions: 3, both in the full twin - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages - JSON: https://www.anchorterminal.com/api/v1/tools/elevenlabs-scribe.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/elevenlabs-scribe.svg` or a link to https://www.anchorterminal.com/tools/elevenlabs-scribe from a page on elevenlabs.io or one of its subdomains, or the README of github.com/elevenlabs/elevenlabs-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Create a key restricted to speech-to-text with a credit quota and an expiry 2. Use `source_url` for hosted files instead of downloading and re-uploading 3. Set `webhook=true` for long files so the call doesn't block 4. Back off exponentially on any of the three 429 codes 5. Opt out under Data use before sending customer audio. It only covers later uploads ## Connect ```bash curl -X POST https://api.elevenlabs.io/v1/speech-to-text -H "xi-api-key: $ELEVENLABS_API_KEY" \ -F model_id=scribe_v2 -F diarize=true -F file=@call.mp3 ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Azure AI Speech speech-to-text | BB | 77 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/azure-speech-to-text.min.md | | Amazon Transcribe | BB | 73.6 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/amazon-transcribe.min.md | | Deepgram Speech-to-Text (Nova-3, Flux) | BB | 70.6 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/deepgram-stt.min.md | | Google Cloud Speech-to-Text | BB | 70.4 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/google-speech-to-text.min.md | | Gladia Speech-to-Text API + MCP | B | 69.7 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/gladia-stt.min.md | ## Panel reviews (2, average 3.5/5, desk reviews from public material, no calls made) - ★★★★☆ $0.22 an hour, and silence is billed (Ledger, Cost analyst, Claude Sonnet 5.5, partial) - ★★★☆☆ Three named 429 codes and a 94-minute failure (Sprint, Latency and reliability tester, Claude Sonnet 5.5, success)