Best of · Voice & speech
Best speech-to-text APIs for AI agents
The 10 highest-scoring of 14 speech-to-text APIs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 14 ranked
- 6 agent-ready
- 14 hosted endpoints
- Updated 9 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Highest score overall
BB, 73.4/100 on the benchmark.
Also Azure AI Speech speech-to-text, BB, 73/100.
Schema & documentation
Deepgram Speech-to-Text (Nova-3, Flux) BB
95/100 on schema & documentation, against 90 for the overall leader.
Security & auth
Google Cloud Speech-to-Text BB
95/100 on security & auth, against 80 for the overall leader.
Maintenance & community
85/100 on maintenance & community, against 35 for the overall leader.
Transparency & trust
Google Cloud Speech-to-Text BB
86/100 on transparency & trust, against 82 for the overall leader.
A hosted MCP endpoint
Deepgram Speech-to-Text (Nova-3, Flux) BB
remote MCP server, nothing to install.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | Amazon Transcribe Amazon Web Services |
BB 73.4 | Teams already on AWS with audio in S3 and IAM in place, especially batch jobs at volume. | Pay per use | hosted |
| 2 | Azure AI Speech speech-to-text Microsoft Azure |
BB 73 | Teams on Azure who need several modes (real time, synchronous files, cheap batch, custom models) under one resource, or strict default data handling. | Freemium | hosted |
| 3 | OpenAI Speech to Text OpenAI |
BB 72.4 | Suited to plain transcription of recorded files at a low price a minute and to teams already holding an OpenAI key. | Pay per use | hosted |
| 4 | Groq Speech-to-Text Groq |
BB 71.8 | Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client. | Freemium | hosted |
| 5 | Deepgram Speech-to-Text (Nova-3, Flux) Deepgram |
BB 70.3 | Live voice agents that want turn detection in the STT model, and for cheap English batch. | Pay per use | hosted and local |
| 6 | Google Cloud Speech-to-Text Google Cloud |
BB 70.2 | Google Cloud teams with audio already in Cloud Storage who want no training and no retention by default. | Freemium | hosted |
| 7 | Gladia Speech-to-Text API + MCP Gladia |
B 69.5 | Multilingual meetings and calls where translation, summaries and NER are wanted in one price, and for teams who want an MCP server. | Freemium | hosted and local |
| 8 | ElevenLabs Scribe Speech to Text API ElevenLabs |
B 68.9 | Batch transcription where diarisation, entity detection and keyterms matter, and for teams already on ElevenLabs for speech. | Freemium | hosted and local |
| 9 | Cartesia Ink Cartesia |
B 67.9 | Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech. | Freemium | hosted |
| 10 | Speechmatics Speech-to-Text Speechmatics |
B 67.1 | Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns. | Pay per use | hosted |
4 more are ranked in the full table.
How to choose
- Accuracy on hard audioCheck word error rate on accented speech, noisy audio and proper names, because a wrong name can send an agent to the wrong record.
- Latency to final transcriptCheck the delay from end of speech to the final transcript in streaming mode, because a live agent waits on that delay before replying to the caller.
- Speaker separationCheck whether speaker labels are returned and how overlapping speech is handled, because an agent summarising a meeting must attribute each action to the right person.
- Cost per audio minuteCheck the cost per audio minute, and whether batch and streaming rates differ, because an agent transcribing long recordings pays for every minute it sends.
How the benchmark tests this category. The same audio set through every API, with accents, background noise, proper names and overlapping speakers. We measure word error rate on each slice, streaming latency to a final transcript, and cost per audio minute.
Each one in detail
Amazon Transcribe
BB 73.4/100AWS's transcription API.
Verdict $0.006 a minute batch and $0.01 streaming in US East, with diarisation, custom vocabulary and language ID included. AWS may store and use audio to improve the service unless an organisation-wide AI services opt-out policy is set.
Choose it for Teams already on AWS with audio in S3 and IAM in place, especially batch jobs at volume.
Strengths
- $0.006 a minute batch and $0.01 streaming in US East, with diarisation, custom vocabulary and language ID included
- A reused
TranscriptionJobNamefails withConflictException, so a retried submission can't create a second job - IAM policies can limit a credential to single actions and resources
Weaknesses
- AWS may store and use audio to improve the service unless an organisation-wide AI services opt-out policy is set
- Batch input must sit in S3, and the WebSocket stream needs a presigned SigV4 URL
- No Transcribe document history entry since 2026-07-01
Price Pay per useAuth API keyx402 nohosted
Azure AI Speech speech-to-text
BB 73/100Azure's speech-to-text service for transcribing audio.
Verdict Real-time and fast transcription audio isn't stored, and customer audio isn't used for training. MAI-Transcribe-2 and the new MAI-Transcribe-2-Streaming are preview with no SLA, and the streaming model's WebSocket route accepts the resource key in the URL query string.
Choose it for Teams on Azure who need several modes (real time, synchronous files, cheap batch, custom models) under one resource, or strict default data handling.
Strengths
- Real-time and fast transcription audio isn't stored, and customer audio isn't used for training
- Fast transcription returns files up to 5 hours and 500 MB in one synchronous call
- Batch at $0.18 an hour, and a free F0 tier with 5 real-time hours a month
Weaknesses
- MAI-Transcribe-2 and MAI-Transcribe-2-Streaming are preview with no SLA, and both introductory prices end with 2026
- The MAI-Transcribe-2-Streaming WebSocket docs allow the resource key as an
api-keyquery parameter - REST API v3.0 and the v3.2 previews were retired on 2026-03-31, and older samples still target them
Price FreemiumAuth OAuth or keyx402 nohosted
OpenAI Speech to Text
BB 72.4/100OpenAI's speech-to-text API. It transcribes uploaded audio files through /v1/audio/transcriptions, translates recordings into English through /v1/audio/translations, and transcribes live audio in Realtime transcription sessions over WebSocket or WebRTC.
Verdict gpt-transcribe costs $0.0045 an audio minute, and the audio endpoints keep no abuse-monitoring logs or application state. Speaker labels, timestamps, subtitles and translation exist only on whisper-1 and gpt-4o-transcribe-diarize, which shut down on 26 February 2027 with no named replacement for those functions.
Choose it for Suited to plain transcription of recorded files at a low price a minute and to teams already holding an OpenAI key.
Strengths
gpt-transcribeis priced at $0.0045 an audio minute on a public page, with per-tier request limits of 5,000, 10,000 and 30,000 a minute- The data controls page lists
/v1/audio/transcriptionsand/v1/audio/translationswith no training, no abuse-monitoring retention and no stored application state - A public OpenAPI 3.1 document, llms.txt and a Markdown twin of every docs page cover the audio endpoints
Weaknesses
whisper-1,gpt-4o-transcribe,gpt-4o-mini-transcribeandgpt-4o-transcribe-diarizewere deprecated on 26 August 2026 and shut down on 26 February 2027- The two named replacements return no speaker labels, word timestamps,
srtorvttoutput or English translation in the reviewed documentation - Uploads stop at 25 MB, the caller splits longer recordings, and the
gpt-transcribemodel page marks the Batch API as not supported
Price Pay per useAuth API keyx402 nohosted
Groq Speech-to-Text
BB 71.8/100Groq's hosted speech-to-text API. It runs OpenAI's Whisper Large v3 and Whisper Large v3 Turbo on OpenAI-compatible transcription and translation endpoints, for uploaded files or audio URLs, with a half-price batch mode.
Verdict Whisper Large v3 Turbo costs $0.04 an audio hour and Whisper Large v3 $0.111, with a no-card free plan and zero data retention as a self-serve setting. There is no streaming endpoint, no diarisation and no subtitle output, and uploads stop at 25 MB on the free plan and 100 MB on the Developer plan.
Choose it for Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client.
Strengths
- Published prices of $0.04 an audio hour for Whisper Large v3 Turbo and $0.111 for Whisper Large v3, with 50% off through the Batch API
- Free plan with no card at 20 requests a minute, 2,000 a day and 7,200 audio seconds an hour on both models
- Inputs and outputs are not retained by default, and zero data retention is a console setting that covers both audio endpoints
Weaknesses
- No streaming or realtime endpoint and no diarisation in the reviewed documentation
- Uploads are capped at 25 MB on the free plan and 100 MB on the Developer plan, so long recordings need client-side chunking
srtandvttresponse formats are not supported, and Whisper Large v3 Turbo cannot translate
Price FreemiumAuth API keyx402 nohosted
Deepgram Speech-to-Text (Nova-3, Flux)
BB 70.3/100Deepgram's speech-to-text API for recorded audio and live streams, including turn detection for voice agents.
Verdict Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag.
Choose it for Live voice agents that want turn detection in the STT model, and for cheap English batch.
Strengths
- Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD
- Nova-3 pre-recorded at $0.0043 a minute with diarisation included
- OpenAPI and AsyncAPI files, an llms.txt and SDKs in six languages
Weaknesses
- Training on audio is the default and the opt-out is a per-request flag
- Two incidents over 2 hours in July 2026, on Flux streaming and batch
- No SLA published for self-serve plans
Price Pay per useAuth API keyx402 nohosted and local
Google Cloud Speech-to-Text
BB 70.2/100Google Cloud's transcription API.
Verdict Audio isn't stored or used for training unless the project opts in to data logging. No release note since 2025-11-13.
Choose it for Google Cloud teams with audio already in Cloud Storage who want no training and no retention by default.
Strengths
- Audio isn't stored or used for training unless the project opts in to data logging
- No Speech-to-Text incident on the Google Cloud status page since 12 June 2025
- Dynamic batch at $0.003 a minute, and standard recognition tiers down to $0.004 past 2M minutes
Weaknesses
- No release note since 2025-11-13
- 82 of the 111 Chirp 3 locales are preview
- Sync requests stop at 1 minute and streams at 5 minutes, and batch reads only from Cloud Storage
Price FreemiumAuth OAuthx402 nohosted
Gladia Speech-to-Text API + MCP
B 69.5/100Speech-to-text API for live and recorded audio, with multilingual transcription and code switching.
Verdict The solaria-1 model supports live and asynchronous transcription in over 100 languages with code switching. Starter pricing is $0.61 an hour for asynchronous transcription and $0.75 for real-time audio.
Choose it for Multilingual meetings and calls where translation, summaries and NER are wanted in one price, and for teams who want an MCP server.
Strengths
- 100+ languages on
solaria-1with code switching, live and async - Translation, summaries, NER and PII redaction included in the hourly price
- OpenAPI file, llms.txt, JavaScript and Python SDKs at 2.0.0 and an official MCP server
Weaknesses
- Starter costs $0.61 an hour async and $0.75 real time, several times the cheapest rivals
- Free-plan audio may be used for training
- The security page and the retention page give different retention defaults
Price FreemiumAuth API keyx402 nohosted and local
ElevenLabs Scribe Speech to Text API
B 68.9/100ElevenLabs' speech-to-text service for audio transcription.
Verdict $0.22 an hour for batch with 90+ languages and diarisation to 32 speakers. Audio may be used for training unless the account opts out, and the opt-out isn't retroactive.
Choose it for Batch transcription where diarisation, entity detection and keyterms matter, and for teams already on ElevenLabs for speech.
Strengths
- $0.22 an hour for batch with 90+ languages and diarisation to 32 speakers
- Keys restricted by endpoint, capped by credits and set to expire
- 429 codes named in the error reference with exponential backoff guidance
Weaknesses
- Audio may be used for training unless the account opts out, and the opt-out isn't retroactive
- Speech-to-text data is retained by default, and zero retention needs Enterprise
- STT request failures for 94 minutes on 29 September 2026, plus latency incidents on 26 August, 4 September and 28 September
Price FreemiumAuth API keyx402 nohosted and local
Cartesia Ink
B 67.9/100Cartesia's hosted speech-to-text API. Ink 2 transcribes live audio in five languages over a WebSocket with built-in turn detection, and the older Ink Whisper model transcribes uploaded files in about 100 languages.
Verdict Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation.
Choose it for Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech.
Strengths
/stt/turns/websocketemitsturn.start,turn.update,turn.eager_end,turn.resumeandturn.end, so no separate voice activity detector is needed- OpenAPI and AsyncAPI files, llms.txt and Markdown twins of every docs page, with enums and numeric ranges on the WebSocket parameters
- Structured errors on
Cartesia-Version2026-03-01 and later, witherror_code,title,message,request_idand an optionaldoc_url
Weaknesses
- The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls
- Zero data retention is an Enterprise plan setting, and no retention period for other plans was found in the terms, privacy policy or docs
ink-2supports English, French, Hindi, Japanese and Spanish only and has no batch endpoint.POST /sttacceptsink-whisperonly
Price FreemiumAuth API keyx402 nohosted
Speechmatics Speech-to-Text
B 67.1/100Speechmatics' APIs for batch and real-time transcription, including speaker-attributed turns for voice agents.
Verdict Training is opt-in and real-time audio is not stored. Enhanced transcription costs $0.40 to $0.43 an hour.
Choose it for Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns.
Strengths
- Training is opt-in only, and realtime audio isn't stored
- ISO/IEC 27001:2022 and SOC 2 Type II
- Transcripts fetched as plain text, JSON or SRT
Weaknesses
- Enhanced costs $0.40 to $0.43 an hour, above most rivals
- No OpenAPI or AsyncAPI file linked from the docs
- 429s carry a reason but no Retry-After or backoff guidance
Price Pay per useAuth API keyx402 nohosted
Head to head
- Amazon Transcribe vs Azure AI Speech speech-to-text BB 73.4 vs BB 73
- Amazon Transcribe vs OpenAI Speech to Text BB 73.4 vs BB 72.4
- Amazon Transcribe vs Groq Speech-to-Text BB 73.4 vs BB 71.8
- Amazon Transcribe vs Deepgram Speech-to-Text (Nova-3, Flux) BB 73.4 vs BB 70.3
- Azure AI Speech speech-to-text vs OpenAI Speech to Text BB 73 vs BB 72.4
- Azure AI Speech speech-to-text vs Groq Speech-to-Text BB 73 vs BB 71.8
- Azure AI Speech speech-to-text vs Deepgram Speech-to-Text (Nova-3, Flux) BB 73 vs BB 70.3
- Groq Speech-to-Text vs OpenAI Speech to Text BB 71.8 vs BB 72.4
- Deepgram Speech-to-Text (Nova-3, Flux) vs OpenAI Speech to Text BB 70.3 vs BB 72.4
- Deepgram Speech-to-Text (Nova-3, Flux) vs Groq Speech-to-Text BB 70.3 vs BB 71.8
Questions
What are the highest-rated speech-to-text APIs for AI agents?
Amazon Transcribe has the highest benchmark score of the 14 ranked speech-to-text APIs, 73.4 (BB). Azure AI Speech speech-to-text is second with 73 (BB).
How many speech-to-text APIs are agent-ready?
6 of the 14 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which speech-to-text APIs accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 91 head-to-head comparisons · Best tools in every category