# Speech-to-text APIs for AI agents > 10 speech-to-text listings ranked by the Anchor benchmark. Leader Azure AI Speech speech-to-text (BB). APIs that turn audio into text, streaming or in batch. Compared on accuracy across accents, noise, names and overlapping speakers, streaming latency, languages, diarisation and cost per audio minute. - Canonical: https://www.anchorterminal.com/categories/speech-to-text - Markdown: https://www.anchorterminal.com/categories/speech-to-text.md (~3,850 tokens) - Slim: https://www.anchorterminal.com/categories/speech-to-text.min.md (~580 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/categories/speech-to-text.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-04 APIs that turn audio into text, streaming or in batch. Compared on accuracy across accents, noise, names and overlapping speakers, streaming latency, languages, diarisation and cost per audio minute. - Tools ranked: 10 · agent-ready (BB or better): 4 · accept x402: 0 · hosted endpoints: 10 · desk reviews by the panel: 26 - JSON: https://www.anchorterminal.com/api/v1/tools.json (list) · https://www.anchorterminal.com/api/v1/rankings.json (ranked) · https://www.anchorterminal.com/api/v1/x402.json (payable) · https://www.anchorterminal.com/api/v1/capabilities.json (by capability) - Grades run AA, A, BB, B, C, D, E, F · methodology: https://www.anchorterminal.com/benchmark/ - Capabilities in this category: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages - https://letme.dev/speech.stt picks the top-graded tool in this list and says how to call it direct; calling through letme comes later (https://www.anchorterminal.com/letme/index.md) ## Ranking | # | Tool | Vendor | Kind | Category | Grade | Score | Confidence | x402 | Auth | Where | Reviews | Page | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 23 | Azure AI Speech speech-to-text | Microsoft Azure | Model API | STT | BB | 77 | medium | no | OAuth or key | hosted | 3.3/5 (8) | https://www.anchorterminal.com/tools/azure-speech-to-text.md | | 57 | Amazon Transcribe | Amazon Web Services | Model API | STT | BB | 73.6 | medium | no | API key | hosted | 4/5 (2) | https://www.anchorterminal.com/tools/amazon-transcribe.md | | 94 | Deepgram Speech-to-Text (Nova-3, Flux) | Deepgram | Model API | STT | BB | 70.6 | medium | no | API key | hosted + local | 3.5/5 (2) | https://www.anchorterminal.com/tools/deepgram-stt.md | | 98 | Google Cloud Speech-to-Text | Google Cloud | Model API | STT | BB | 70.4 | medium | no | OAuth | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/google-speech-to-text.md | | 108 | Gladia Speech-to-Text API + MCP | Gladia | Model API | STT | B | 69.7 | medium | no | API key | hosted + local | 3/5 (2) | https://www.anchorterminal.com/tools/gladia-stt.md | | 117 | ElevenLabs Scribe Speech to Text API | ElevenLabs | Model API | STT | B | 69 | medium | no | API key | hosted + local | 3.5/5 (2) | https://www.anchorterminal.com/tools/elevenlabs-scribe.md | | 145 | Speechmatics Speech-to-Text | Speechmatics | Model API | STT | B | 67.3 | medium | no | API key | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/speechmatics-stt.md | | 148 | AssemblyAI Speech-to-Text (Universal) | AssemblyAI | Model API | STT | B | 67 | medium | no | API key | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/assemblyai-stt.md | | 281 | Soniox Speech-to-Text | Soniox | Model API | STT | C | 58.3 | medium | no | API key | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/soniox-stt.md | | 285 | Rev AI Speech-to-Text API | Rev | Model API | STT | C | 58 | medium | no | API key | hosted | 2.5/5 (2) | https://www.anchorterminal.com/tools/rev-ai-stt.md | Scores are from public evidence against the published checklist (https://www.anchorterminal.com/benchmark/), with Performance and Task success pending. p95 latency and context cost come from our probes, which haven't run yet. ## Summaries ### 23. Azure AI Speech speech-to-text, BB (77) Azure's speech-to-text service for transcribing audio. Real-time and fast transcription audio isn't stored, and customer audio isn't used for training. MAI-Transcribe-2 is preview with no SLA, and its $0.10 promotional price ends on 2026-12-31. - Page: https://www.anchorterminal.com/tools/azure-speech-to-text · Markdown: https://www.anchorterminal.com/tools/azure-speech-to-text.md · JSON: https://www.anchorterminal.com/api/v1/tools/azure-speech-to-text.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation · endpoint: `https://eastus.api.cognitive.microsoft.com/speechtotext` ### 57. Amazon Transcribe, BB (73.6) AWS's transcription API. $0.006 a minute batch and $0.01 streaming in US East, with diarisation, custom vocabulary and language ID included. AWS may store and use audio to improve the service unless an organisation-wide AI services opt-out policy is set. - Page: https://www.anchorterminal.com/tools/amazon-transcribe · Markdown: https://www.anchorterminal.com/tools/amazon-transcribe.md · JSON: https://www.anchorterminal.com/api/v1/tools/amazon-transcribe.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages · endpoint: `https://transcribe.us-east-1.amazonaws.com` ### 94. Deepgram Speech-to-Text (Nova-3, Flux), BB (70.6) Deepgram's speech-to-text API for recorded audio and live streams, including turn detection for voice agents. Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag. - Page: https://www.anchorterminal.com/tools/deepgram-stt · Markdown: https://www.anchorterminal.com/tools/deepgram-stt.md · JSON: https://www.anchorterminal.com/api/v1/tools/deepgram-stt.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages · endpoint: `https://api.deepgram.com/v1` ### 98. Google Cloud Speech-to-Text, BB (70.4) Google Cloud's transcription API. Audio isn't stored or used for training unless the project opts in to data logging. No release note since 2025-11-13. - Page: https://www.anchorterminal.com/tools/google-speech-to-text · Markdown: https://www.anchorterminal.com/tools/google-speech-to-text.md · JSON: https://www.anchorterminal.com/api/v1/tools/google-speech-to-text.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation · endpoint: `https://speech.googleapis.com/v2` ### 108. Gladia Speech-to-Text API + MCP, B (69.7) Speech-to-text API for live and recorded audio, with multilingual transcription and code switching. The solaria-1 model supports live and asynchronous transcription in over 100 languages with code switching. Starter pricing is $0.61 an hour for asynchronous transcription and $0.75 for real-time audio. - Page: https://www.anchorterminal.com/tools/gladia-stt · Markdown: https://www.anchorterminal.com/tools/gladia-stt.md · JSON: https://www.anchorterminal.com/api/v1/tools/gladia-stt.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation · endpoint: `https://api.gladia.io/v2` ### 117. ElevenLabs Scribe Speech to Text API, B (69) ElevenLabs' speech-to-text service for audio transcription. $0.22 an hour for batch with 90+ languages and diarisation to 32 speakers. Audio may be used for training unless the account opts out, and the opt-out isn't retroactive. - Page: https://www.anchorterminal.com/tools/elevenlabs-scribe · Markdown: https://www.anchorterminal.com/tools/elevenlabs-scribe.md · JSON: https://www.anchorterminal.com/api/v1/tools/elevenlabs-scribe.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages · endpoint: `https://api.elevenlabs.io/v1` ### 145. Speechmatics Speech-to-Text, B (67.3) Speechmatics' APIs for batch and real-time transcription, including speaker-attributed turns for voice agents. Training is opt-in and real-time audio is not stored. Enhanced transcription costs $0.40 to $0.43 an hour. - Page: https://www.anchorterminal.com/tools/speechmatics-stt · Markdown: https://www.anchorterminal.com/tools/speechmatics-stt.md · JSON: https://www.anchorterminal.com/api/v1/tools/speechmatics-stt.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation · endpoint: `https://eu1.asr.api.speechmatics.com/v2` ### 148. AssemblyAI Speech-to-Text (Universal), B (67) Speech-to-text APIs for recorded audio and live streams, with speaker identification, translation and redaction options. OpenAPI 3.1 file with typed inputs and error responses on every operation. Two outages of an hour or more in the last 90 days, on 31 July and 16 September 2026. - Page: https://www.anchorterminal.com/tools/assemblyai-stt · Markdown: https://www.anchorterminal.com/tools/assemblyai-stt.md · JSON: https://www.anchorterminal.com/api/v1/tools/assemblyai-stt.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation · endpoint: `https://api.assemblyai.com/v2` ### 281. Soniox Speech-to-Text, C (58.3) One multilingual model family for 60+ languages, as a real-time WebSocket API (`stt-rt-v5`) and an async file API (`stt-async-v5`). About $0.10 an hour async and $0.12 real time, with diarisation, language ID and translation included. No free credits for new accounts since October 2025. - Page: https://www.anchorterminal.com/tools/soniox-stt · Markdown: https://www.anchorterminal.com/tools/soniox-stt.md · JSON: https://www.anchorterminal.com/api/v1/tools/soniox-stt.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation · endpoint: `https://api.soniox.com/v1` ### 285. Rev AI Speech-to-Text API, C (58) Rev's API for recorded and streaming speech transcription, with human transcription available through the same job endpoint. Reverb English at $0.20 an hour, foreign languages at $0.30. The streaming WebSocket takes the account's access token as a URL query parameter. - Page: https://www.anchorterminal.com/tools/rev-ai-stt · Markdown: https://www.anchorterminal.com/tools/rev-ai-stt.md · JSON: https://www.anchorterminal.com/api/v1/tools/rev-ai-stt.json - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages, speech.translation · endpoint: `https://api.rev.ai/speechtotext/v1` ## How we test this category The same audio set through every API, with accents, background noise, proper names and overlapping speakers. We measure word error rate on each slice, streaming latency to a final transcript, and cost per audio minute. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence. ## Indexed, not reviewed (20) Sorted into this category from public catalogues, with facts and our own checks but no score, grade or rank (https://www.anchorterminal.com/indexed/index.md). | Listing | Kind | What it does | Why it's here | | --- | --- | --- | --- | | [2kw.ai](https://www.anchorterminal.com/tools/2kw-mcp-server.md) | MCP server | EU-hosted AI platform: OpenAI-compatible LLM gateway, document extraction, transcription, agents. | vendor's own, widely used | | [backengine-mcp](https://www.anchorterminal.com/tools/backengine-mcp.md) | MCP server | Surface customer & prospect context from Slack, email, transcripts and tickets in any MCP client. | vendor's own | | [Bigdata.com](https://www.anchorterminal.com/tools/bigdata-mcp.md) | MCP server | Licensed, cited financial data for AI agents: news, filings, transcripts, research and tearsheets. | vendor's own | | [Cleat](https://www.anchorterminal.com/tools/cleat.md) | MCP server | A real US mobile number that receives verification codes, by text or transcribed call. | vendor's own | | [clipy.online MCP server](https://www.anchorterminal.com/tools/clipy-mcp.md) | MCP server | Read your Clipy screen recordings: search, transcripts, AI summaries, and key moments with frames. | vendor's own | | [Cortex RMCP](https://www.anchorterminal.com/tools/dinglebear-cortex-rmcp.md) | MCP server | Rust MCP server for homelab logs, syslog, Docker logs, FTS search, and AI transcript correlation. | vendor's own | | [Crixin Voice](https://www.anchorterminal.com/tools/crixin-voice.md) | MCP server | Give your AI a real phone: place calls, send SMS, fetch recordings and transcripts. Local or hosted. | vendor's own | | [DialMCP](https://www.anchorterminal.com/tools/dialmcp.md) | MCP server | Let AI agents place real phone calls from your verified number, with transcripts and recordings. | vendor's own | | [Gavelin](https://www.anchorterminal.com/tools/gavelin-mcp.md) | MCP server | Search bills and speaker-attributed hearing transcripts across all 50 US state legislatures. | vendor's own | | [importly-mcp](https://www.anchorterminal.com/tools/importly-mcp.md) | MCP server | Download, transcribe, and get metadata for any video/audio URL. Managed yt-dlp + speech-to-text. | vendor's own | | [Memo AI – meeting assistant](https://www.anchorterminal.com/tools/memoai-memo-ai.md) | MCP server | Search transcripts and summaries of your meetings, calls, and recordings in Memo AI | vendor's own | | [pepys-mcp](https://www.anchorterminal.com/tools/pepys-mcp.md) | MCP server | Transcribe audio & video: diarization, timed SRT/VTT, podcasts, paste-a-link, whole-feed batch. | vendor's own | | [Rephonic](https://www.anchorterminal.com/tools/rephonic.md) | MCP server | Research 3M+ podcasts with audience data, contacts, episodes, transcripts, charts, and sponsors. | vendor's own | | [SOAPNoteAPI](https://www.anchorterminal.com/tools/soapnoteapi-mcp.md) | MCP server | Generate clinical SOAP notes, billing codes, and visit summaries from transcripts or audio. | vendor's own | | [transcription](https://www.anchorterminal.com/tools/scriptivox-transcription.md) | MCP server | AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export. | vendor's own | | [Vexa](https://www.anchorterminal.com/tools/vexa.md) | MCP server | Meeting bot and transcripts for Google Meet, Teams and Zoom. Live or after, speakers labelled. | vendor's own, widely used | | [Video Extract](https://www.anchorterminal.com/tools/yanlinglabs-video-extract-mcp.md) | MCP server | Download any video from a URL, or get its transcript and key frames. All local, no API keys. | vendor's own | | [Voxplo](https://www.anchorterminal.com/tools/voxplo.md) | MCP server | Give AI agents a phone: outbound AI calls that return a summary, transcript, and extracted fields. | vendor's own | | [YouTube Transcript + YouTube Search MCP](https://www.anchorterminal.com/tools/getyoutubetranscript-youtube-transcript-and-youtube-search.md) | MCP server | YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP. | vendor's own | | [💯 YouTube Transcript + YouTube Search MCP for AI Agents](https://www.anchorterminal.com/tools/transcriptapi-youtube-transcript-and-youtube-search.md) | MCP server | 💯 The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. | vendor's own |