# Best speech-to-text APIs for AI agents (slim) > Amazon Transcribe (BB), Azure AI Speech speech-to-text (BB) and OpenAI Speech to Text (BB) lead the 14 ranked speech-to-text APIs. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Full: https://www.anchorterminal.com/best/speech-to-text/index.md (~5,750 tokens) · this version ~1,530 tokens · JSON https://www.anchorterminal.com/best/speech-to-text/index.json · canonical https://www.anchorterminal.com/best/speech-to-text/ - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 The 10 highest-scoring of 14 speech-to-text APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 14 · agent-ready (BB or better): 6 · accept x402: 0 · hosted endpoints: 14 - Full ranked table: https://www.anchorterminal.com/categories/speech-to-text.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/speech-to-text/index.md (91) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Amazon Transcribe](https://www.anchorterminal.com/tools/amazon-transcribe.md) | BB | 73.4 | Teams already on AWS with audio in S3 and IAM in place, especially batch jobs at volume. | Pay per use | hosted | | 2 | [Azure AI Speech speech-to-text](https://www.anchorterminal.com/tools/azure-speech-to-text.md) | BB | 73 | Teams on Azure who need several modes (real time, synchronous files, cheap batch, custom models) under one resource, or strict default data handling. | Freemium | hosted | | 3 | [OpenAI Speech to Text](https://www.anchorterminal.com/tools/openai-speech-to-text.md) | BB | 72.4 | Suited to plain transcription of recorded files at a low price a minute and to teams already holding an OpenAI key. | Pay per use | hosted | | 4 | [Groq Speech-to-Text](https://www.anchorterminal.com/tools/groq-speech-to-text.md) | BB | 71.8 | Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client. | Freemium | hosted | | 5 | [Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/tools/deepgram-stt.md) | BB | 70.3 | Live voice agents that want turn detection in the STT model, and for cheap English batch. | Pay per use | hosted and local | | 6 | [Google Cloud Speech-to-Text](https://www.anchorterminal.com/tools/google-speech-to-text.md) | BB | 70.2 | Google Cloud teams with audio already in Cloud Storage who want no training and no retention by default. | Freemium | hosted | | 7 | [Gladia Speech-to-Text API + MCP](https://www.anchorterminal.com/tools/gladia-stt.md) | B | 69.5 | Multilingual meetings and calls where translation, summaries and NER are wanted in one price, and for teams who want an MCP server. | Freemium | hosted and local | | 8 | [ElevenLabs Scribe Speech to Text API](https://www.anchorterminal.com/tools/elevenlabs-scribe.md) | B | 68.9 | Batch transcription where diarisation, entity detection and keyterms matter, and for teams already on ElevenLabs for speech. | Freemium | hosted and local | | 9 | [Cartesia Ink](https://www.anchorterminal.com/tools/cartesia-ink-stt.md) | B | 67.9 | Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech. | Freemium | hosted | | 10 | [Speechmatics Speech-to-Text](https://www.anchorterminal.com/tools/speechmatics-stt.md) | B | 67.1 | Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns. | Pay per use | hosted | ## Picks by need - Highest score overall: [Amazon Transcribe](https://www.anchorterminal.com/tools/amazon-transcribe.md), BB, 73.4/100 on the benchmark. Also [Azure AI Speech speech-to-text](https://www.anchorterminal.com/tools/azure-speech-to-text.md), BB, 73/100. - Schema & documentation: [Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/tools/deepgram-stt.md), 95/100 on schema & documentation, against 90 for the overall leader. - Security & auth: [Google Cloud Speech-to-Text](https://www.anchorterminal.com/tools/google-speech-to-text.md), 95/100 on security & auth, against 80 for the overall leader. - Maintenance & community: [Cartesia Ink](https://www.anchorterminal.com/tools/cartesia-ink-stt.md), 85/100 on maintenance & community, against 35 for the overall leader. - Transparency & trust: [Google Cloud Speech-to-Text](https://www.anchorterminal.com/tools/google-speech-to-text.md), 86/100 on transparency & trust, against 82 for the overall leader. - A hosted MCP endpoint: [Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/tools/deepgram-stt.md), remote MCP server, nothing to install. - The review panel's favourite: [Azure AI Speech speech-to-text](https://www.anchorterminal.com/tools/azure-speech-to-text.md), 3.3/5 from 8 panel reviews. ## How to choose - Accuracy on hard audio: Check word error rate on accented speech, noisy audio and proper names, because a wrong name can send an agent to the wrong record. - Latency to final transcript: Check the delay from end of speech to the final transcript in streaming mode, because a live agent waits on that delay before replying to the caller. - Speaker separation: Check whether speaker labels are returned and how overlapping speech is handled, because an agent summarising a meeting must attribute each action to the right person. - Cost per audio minute: Check the cost per audio minute, and whether batch and streaming rates differ, because an agent transcribing long recordings pays for every minute it sends. - How the benchmark tests this category: The same audio set through every API, with accents, background noise, proper names and overlapping speakers. We measure word error rate on each slice, streaming latency to a final transcript, and cost per audio minute. Each listing's verdict, strengths and weaknesses: https://www.anchorterminal.com/best/speech-to-text/index.md