# Best speech-to-text APIs for AI agents > Amazon Transcribe (BB), Azure AI Speech speech-to-text (BB) and OpenAI Speech to Text (BB) lead the 14 ranked speech-to-text APIs. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Canonical: https://www.anchorterminal.com/best/speech-to-text/ - Markdown: https://www.anchorterminal.com/best/speech-to-text/index.md (~5,750 tokens) - Slim: https://www.anchorterminal.com/best/speech-to-text/index.min.md (~1,530 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/best/speech-to-text/index.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 The 10 highest-scoring of 14 speech-to-text APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 14 · agent-ready (BB or better): 6 · accept x402: 0 · hosted endpoints: 14 - Full ranked table: https://www.anchorterminal.com/categories/speech-to-text.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/speech-to-text/index.md (91) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Amazon Transcribe](https://www.anchorterminal.com/tools/amazon-transcribe.md) | BB | 73.4 | Teams already on AWS with audio in S3 and IAM in place, especially batch jobs at volume. | Pay per use | hosted | | 2 | [Azure AI Speech speech-to-text](https://www.anchorterminal.com/tools/azure-speech-to-text.md) | BB | 73 | Teams on Azure who need several modes (real time, synchronous files, cheap batch, custom models) under one resource, or strict default data handling. | Freemium | hosted | | 3 | [OpenAI Speech to Text](https://www.anchorterminal.com/tools/openai-speech-to-text.md) | BB | 72.4 | Suited to plain transcription of recorded files at a low price a minute and to teams already holding an OpenAI key. | Pay per use | hosted | | 4 | [Groq Speech-to-Text](https://www.anchorterminal.com/tools/groq-speech-to-text.md) | BB | 71.8 | Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client. | Freemium | hosted | | 5 | [Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/tools/deepgram-stt.md) | BB | 70.3 | Live voice agents that want turn detection in the STT model, and for cheap English batch. | Pay per use | hosted and local | | 6 | [Google Cloud Speech-to-Text](https://www.anchorterminal.com/tools/google-speech-to-text.md) | BB | 70.2 | Google Cloud teams with audio already in Cloud Storage who want no training and no retention by default. | Freemium | hosted | | 7 | [Gladia Speech-to-Text API + MCP](https://www.anchorterminal.com/tools/gladia-stt.md) | B | 69.5 | Multilingual meetings and calls where translation, summaries and NER are wanted in one price, and for teams who want an MCP server. | Freemium | hosted and local | | 8 | [ElevenLabs Scribe Speech to Text API](https://www.anchorterminal.com/tools/elevenlabs-scribe.md) | B | 68.9 | Batch transcription where diarisation, entity detection and keyterms matter, and for teams already on ElevenLabs for speech. | Freemium | hosted and local | | 9 | [Cartesia Ink](https://www.anchorterminal.com/tools/cartesia-ink-stt.md) | B | 67.9 | Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech. | Freemium | hosted | | 10 | [Speechmatics Speech-to-Text](https://www.anchorterminal.com/tools/speechmatics-stt.md) | B | 67.1 | Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns. | Pay per use | hosted | ## Picks by need - Highest score overall: [Amazon Transcribe](https://www.anchorterminal.com/tools/amazon-transcribe.md), BB, 73.4/100 on the benchmark. Also [Azure AI Speech speech-to-text](https://www.anchorterminal.com/tools/azure-speech-to-text.md), BB, 73/100. - Schema & documentation: [Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/tools/deepgram-stt.md), 95/100 on schema & documentation, against 90 for the overall leader. - Security & auth: [Google Cloud Speech-to-Text](https://www.anchorterminal.com/tools/google-speech-to-text.md), 95/100 on security & auth, against 80 for the overall leader. - Maintenance & community: [Cartesia Ink](https://www.anchorterminal.com/tools/cartesia-ink-stt.md), 85/100 on maintenance & community, against 35 for the overall leader. - Transparency & trust: [Google Cloud Speech-to-Text](https://www.anchorterminal.com/tools/google-speech-to-text.md), 86/100 on transparency & trust, against 82 for the overall leader. - A hosted MCP endpoint: [Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/tools/deepgram-stt.md), remote MCP server, nothing to install. - The review panel's favourite: [Azure AI Speech speech-to-text](https://www.anchorterminal.com/tools/azure-speech-to-text.md), 3.3/5 from 8 panel reviews. ## How to choose - Accuracy on hard audio: Check word error rate on accented speech, noisy audio and proper names, because a wrong name can send an agent to the wrong record. - Latency to final transcript: Check the delay from end of speech to the final transcript in streaming mode, because a live agent waits on that delay before replying to the caller. - Speaker separation: Check whether speaker labels are returned and how overlapping speech is handled, because an agent summarising a meeting must attribute each action to the right person. - Cost per audio minute: Check the cost per audio minute, and whether batch and streaming rates differ, because an agent transcribing long recordings pays for every minute it sends. - How the benchmark tests this category: The same audio set through every API, with accents, background noise, proper names and overlapping speakers. We measure word error rate on each slice, streaming latency to a final transcript, and cost per audio minute. ## Each one in detail ### 1. Amazon Transcribe, BB 73.4/100 AWS's transcription API. - Verdict: $0.006 a minute batch and $0.01 streaming in US East, with diarisation, custom vocabulary and language ID included. AWS may store and use audio to improve the service unless an organisation-wide AI services opt-out policy is set. - Choose it for: Teams already on AWS with audio in S3 and IAM in place, especially batch jobs at volume. - Strength: $0.006 a minute batch and $0.01 streaming in US East, with diarisation, custom vocabulary and language ID included - Strength: A reused `TranscriptionJobName` fails with `ConflictException`, so a retried submission can't create a second job - Strength: IAM policies can limit a credential to single actions and resources - Weakness: AWS may store and use audio to improve the service unless an organisation-wide AI services opt-out policy is set - Weakness: Batch input must sit in S3, and the WebSocket stream needs a presigned SigV4 URL - Weakness: No Transcribe document history entry since 2026-07-01 - Price: Pay per use · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/amazon-transcribe.md ### 2. Azure AI Speech speech-to-text, BB 73/100 Azure's speech-to-text service for transcribing audio. - Verdict: Real-time and fast transcription audio isn't stored, and customer audio isn't used for training. MAI-Transcribe-2 and the new MAI-Transcribe-2-Streaming are preview with no SLA, and the streaming model's WebSocket route accepts the resource key in the URL query string. - Choose it for: Teams on Azure who need several modes (real time, synchronous files, cheap batch, custom models) under one resource, or strict default data handling. - Strength: Real-time and fast transcription audio isn't stored, and customer audio isn't used for training - Strength: Fast transcription returns files up to 5 hours and 500 MB in one synchronous call - Strength: Batch at $0.18 an hour, and a free F0 tier with 5 real-time hours a month - Weakness: MAI-Transcribe-2 and MAI-Transcribe-2-Streaming are preview with no SLA, and both introductory prices end with 2026 - Weakness: The MAI-Transcribe-2-Streaming WebSocket docs allow the resource key as an `api-key` query parameter - Weakness: REST API v3.0 and the v3.2 previews were retired on 2026-03-31, and older samples still target them - Price: Freemium · Auth: OAuth or key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/azure-speech-to-text.md - Against #1: https://www.anchorterminal.com/compare/amazon-transcribe-vs-azure-speech-to-text.md ### 3. OpenAI Speech to Text, BB 72.4/100 OpenAI's speech-to-text API. It transcribes uploaded audio files through `/v1/audio/transcriptions`, translates recordings into English through `/v1/audio/translations`, and transcribes live audio in Realtime transcription sessions over WebSocket or WebRTC. - Verdict: `gpt-transcribe` costs $0.0045 an audio minute, and the audio endpoints keep no abuse-monitoring logs or application state. Speaker labels, timestamps, subtitles and translation exist only on `whisper-1` and `gpt-4o-transcribe-diarize`, which shut down on 26 February 2027 with no named replacement for those functions. - Choose it for: Suited to plain transcription of recorded files at a low price a minute and to teams already holding an OpenAI key. - Strength: `gpt-transcribe` is priced at $0.0045 an audio minute on a public page, with per-tier request limits of 5,000, 10,000 and 30,000 a minute - Strength: The data controls page lists `/v1/audio/transcriptions` and `/v1/audio/translations` with no training, no abuse-monitoring retention and no stored application state - Strength: A public OpenAPI 3.1 document, llms.txt and a Markdown twin of every docs page cover the audio endpoints - Weakness: `whisper-1`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe` and `gpt-4o-transcribe-diarize` were deprecated on 26 August 2026 and shut down on 26 February 2027 - Weakness: The two named replacements return no speaker labels, word timestamps, `srt` or `vtt` output or English translation in the reviewed documentation - Weakness: Uploads stop at 25 MB, the caller splits longer recordings, and the `gpt-transcribe` model page marks the Batch API as not supported - Price: Pay per use · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/openai-speech-to-text.md - Against #1: https://www.anchorterminal.com/compare/amazon-transcribe-vs-openai-speech-to-text.md ### 4. Groq Speech-to-Text, BB 71.8/100 Groq's hosted speech-to-text API. It runs OpenAI's Whisper Large v3 and Whisper Large v3 Turbo on OpenAI-compatible transcription and translation endpoints, for uploaded files or audio URLs, with a half-price batch mode. - Verdict: Whisper Large v3 Turbo costs $0.04 an audio hour and Whisper Large v3 $0.111, with a no-card free plan and zero data retention as a self-serve setting. There is no streaming endpoint, no diarisation and no subtitle output, and uploads stop at 25 MB on the free plan and 100 MB on the Developer plan. - Choose it for: Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client. - Strength: Published prices of $0.04 an audio hour for Whisper Large v3 Turbo and $0.111 for Whisper Large v3, with 50% off through the Batch API - Strength: Free plan with no card at 20 requests a minute, 2,000 a day and 7,200 audio seconds an hour on both models - Strength: Inputs and outputs are not retained by default, and zero data retention is a console setting that covers both audio endpoints - Weakness: No streaming or realtime endpoint and no diarisation in the reviewed documentation - Weakness: Uploads are capped at 25 MB on the free plan and 100 MB on the Developer plan, so long recordings need client-side chunking - Weakness: `srt` and `vtt` response formats are not supported, and Whisper Large v3 Turbo cannot translate - Price: Freemium · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/groq-speech-to-text.md - Against #1: https://www.anchorterminal.com/compare/amazon-transcribe-vs-groq-speech-to-text.md ### 5. Deepgram Speech-to-Text (Nova-3, Flux), BB 70.3/100 Deepgram's speech-to-text API for recorded audio and live streams, including turn detection for voice agents. - Verdict: Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag. - Choose it for: Live voice agents that want turn detection in the STT model, and for cheap English batch. - Strength: Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD - Strength: Nova-3 pre-recorded at $0.0043 a minute with diarisation included - Strength: OpenAPI and AsyncAPI files, an llms.txt and SDKs in six languages - Weakness: Training on audio is the default and the opt-out is a per-request flag - Weakness: Two incidents over 2 hours in July 2026, on Flux streaming and batch - Weakness: No SLA published for self-serve plans - Price: Pay per use · Auth: API key · x402: no · Where: hosted and local - Full assessment: https://www.anchorterminal.com/tools/deepgram-stt.md - Against #1: https://www.anchorterminal.com/compare/amazon-transcribe-vs-deepgram-stt.md ### 6. Google Cloud Speech-to-Text, BB 70.2/100 Google Cloud's transcription API. - Verdict: Audio isn't stored or used for training unless the project opts in to data logging. No release note since 2025-11-13. - Choose it for: Google Cloud teams with audio already in Cloud Storage who want no training and no retention by default. - Strength: Audio isn't stored or used for training unless the project opts in to data logging - Strength: No Speech-to-Text incident on the Google Cloud status page since 12 June 2025 - Strength: Dynamic batch at $0.003 a minute, and standard recognition tiers down to $0.004 past 2M minutes - Weakness: No release note since 2025-11-13 - Weakness: 82 of the 111 Chirp 3 locales are preview - Weakness: Sync requests stop at 1 minute and streams at 5 minutes, and batch reads only from Cloud Storage - Price: Freemium · Auth: OAuth · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/google-speech-to-text.md - Against #1: https://www.anchorterminal.com/compare/amazon-transcribe-vs-google-speech-to-text.md ### 7. Gladia Speech-to-Text API + MCP, B 69.5/100 Speech-to-text API for live and recorded audio, with multilingual transcription and code switching. - Verdict: The solaria-1 model supports live and asynchronous transcription in over 100 languages with code switching. Starter pricing is $0.61 an hour for asynchronous transcription and $0.75 for real-time audio. - Choose it for: Multilingual meetings and calls where translation, summaries and NER are wanted in one price, and for teams who want an MCP server. - Strength: 100+ languages on `solaria-1` with code switching, live and async - Strength: Translation, summaries, NER and PII redaction included in the hourly price - Strength: OpenAPI file, llms.txt, JavaScript and Python SDKs at 2.0.0 and an official MCP server - Weakness: Starter costs $0.61 an hour async and $0.75 real time, several times the cheapest rivals - Weakness: Free-plan audio may be used for training - Weakness: The security page and the retention page give different retention defaults - Price: Freemium · Auth: API key · x402: no · Where: hosted and local - Full assessment: https://www.anchorterminal.com/tools/gladia-stt.md - Against #1: https://www.anchorterminal.com/compare/amazon-transcribe-vs-gladia-stt.md ### 8. ElevenLabs Scribe Speech to Text API, B 68.9/100 ElevenLabs' speech-to-text service for audio transcription. - Verdict: $0.22 an hour for batch with 90+ languages and diarisation to 32 speakers. Audio may be used for training unless the account opts out, and the opt-out isn't retroactive. - Choose it for: Batch transcription where diarisation, entity detection and keyterms matter, and for teams already on ElevenLabs for speech. - Strength: $0.22 an hour for batch with 90+ languages and diarisation to 32 speakers - Strength: Keys restricted by endpoint, capped by credits and set to expire - Strength: 429 codes named in the error reference with exponential backoff guidance - Weakness: Audio may be used for training unless the account opts out, and the opt-out isn't retroactive - Weakness: Speech-to-text data is retained by default, and zero retention needs Enterprise - Weakness: STT request failures for 94 minutes on 29 September 2026, plus latency incidents on 26 August, 4 September and 28 September - Price: Freemium · Auth: API key · x402: no · Where: hosted and local - Full assessment: https://www.anchorterminal.com/tools/elevenlabs-scribe.md - Against #1: https://www.anchorterminal.com/compare/amazon-transcribe-vs-elevenlabs-scribe.md ### 9. Cartesia Ink, B 67.9/100 Cartesia's hosted speech-to-text API. Ink 2 transcribes live audio in five languages over a WebSocket with built-in turn detection, and the older Ink Whisper model transcribes uploaded files in about 100 languages. - Verdict: Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation. - Choose it for: Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech. - Strength: `/stt/turns/websocket` emits `turn.start`, `turn.update`, `turn.eager_end`, `turn.resume` and `turn.end`, so no separate voice activity detector is needed - Strength: OpenAPI and AsyncAPI files, llms.txt and Markdown twins of every docs page, with enums and numeric ranges on the WebSocket parameters - Strength: Structured errors on `Cartesia-Version` 2026-03-01 and later, with `error_code`, `title`, `message`, `request_id` and an optional `doc_url` - Weakness: The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls - Weakness: Zero data retention is an Enterprise plan setting, and no retention period for other plans was found in the terms, privacy policy or docs - Weakness: `ink-2` supports English, French, Hindi, Japanese and Spanish only and has no batch endpoint. `POST /stt` accepts `ink-whisper` only - Price: Freemium · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/cartesia-ink-stt.md - Against #1: https://www.anchorterminal.com/compare/amazon-transcribe-vs-cartesia-ink-stt.md ### 10. Speechmatics Speech-to-Text, B 67.1/100 Speechmatics' APIs for batch and real-time transcription, including speaker-attributed turns for voice agents. - Verdict: Training is opt-in and real-time audio is not stored. Enhanced transcription costs $0.40 to $0.43 an hour. - Choose it for: Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns. - Strength: Training is opt-in only, and realtime audio isn't stored - Strength: ISO/IEC 27001:2022 and SOC 2 Type II - Strength: Transcripts fetched as plain text, JSON or SRT - Weakness: Enhanced costs $0.40 to $0.43 an hour, above most rivals - Weakness: No OpenAPI or AsyncAPI file linked from the docs - Weakness: 429s carry a reason but no Retry-After or backoff guidance - Price: Pay per use · Auth: API key · x402: no · Where: hosted - Full assessment: https://www.anchorterminal.com/tools/speechmatics-stt.md - Against #1: https://www.anchorterminal.com/compare/amazon-transcribe-vs-speechmatics-stt.md 4 more are ranked in the full table: https://www.anchorterminal.com/categories/speech-to-text.md ## Head to head - [Amazon Transcribe vs Azure AI Speech speech-to-text](https://www.anchorterminal.com/compare/amazon-transcribe-vs-azure-speech-to-text.md) - [Amazon Transcribe vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/amazon-transcribe-vs-openai-speech-to-text.md) - [Amazon Transcribe vs Groq Speech-to-Text](https://www.anchorterminal.com/compare/amazon-transcribe-vs-groq-speech-to-text.md) - [Amazon Transcribe vs Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/compare/amazon-transcribe-vs-deepgram-stt.md) - [Azure AI Speech speech-to-text vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-openai-speech-to-text.md) - [Azure AI Speech speech-to-text vs Groq Speech-to-Text](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-groq-speech-to-text.md) - [Azure AI Speech speech-to-text vs Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-deepgram-stt.md) - [Groq Speech-to-Text vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/groq-speech-to-text-vs-openai-speech-to-text.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/deepgram-stt-vs-openai-speech-to-text.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs Groq Speech-to-Text](https://www.anchorterminal.com/compare/deepgram-stt-vs-groq-speech-to-text.md) ## Questions ### What are the highest-rated speech-to-text APIs for AI agents? Amazon Transcribe has the highest benchmark score of the 14 ranked speech-to-text APIs, 73.4 (BB). Azure AI Speech speech-to-text is second with 73 (BB). ### How many speech-to-text APIs are agent-ready? 6 of the 14 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark. ### Which speech-to-text APIs accept x402 payments? None of the ranked listings here accepts x402 for its main call yet. ### How is this list ranked? By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026. ## How this list is made The order is the Anchor benchmark score, the same number as on each listing. Each listing is graded from public evidence against the benchmark checklist, and the picks are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.