Head to head · Speech stt · October 2026 research run
Mistral Voxtral Transcribe vs Speechmatics Speech-to-Text
Speechmatics Speech-to-Text scores 67.1 (B) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 3 of 7 scored categories. Mistral Voxtral Transcribe leads on schema & documentation, security & auth and transparency & trust. Both do speech stt.
Which one, for what
Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.
Ahead on
- Schema & documentation, 85 against 65
- Security & auth, 70 against 65
- Transparency & trust, 82 against 75
Watch for
Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel
Good for Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns.
Ahead on
- Reliability, 70 against 53
- Maintenance & community, 75 against 66
Also in its favour
- Free to start without a card
- No incidents deducted, where Mistral Voxtral Transcribe loses 3 points for them
Watch for
Enhanced costs $0.40 to $0.43 an hour, above most rivals
Score by category
| Category | Weight this run | Mistral Voxtral Transcribe | Speechmatics Speech-to-Text | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 53 | 70 | Speechmatics Speech-to-Text +17 |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 85 | 65 | Mistral Voxtral Transcribe +20 |
| Agent ergonomics | 13%16.2 | 77 | 80 | Speechmatics Speech-to-Text +3 |
| Security & auth | 14%17.5 | 70 | 65 | Mistral Voxtral Transcribe +5 |
| Payments & pricing | 10%12.5 | 40 | 40 | even |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 66 | 75 | Speechmatics Speech-to-Text +9 |
| Transparency & trust | 7%8.8 | 82 | 75 | Mistral Voxtral Transcribe +7 |
| Negative events | ≤15 | -3 | 0 | |
| Total | 64.1 · B | 67.1 · B |
Facts side by side
| Fact | Mistral Voxtral Transcribe | Speechmatics Speech-to-Text |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | Mistral AI | Speechmatics |
| Hosted endpoint | https://api.mistral.ai/v1 | https://eu1.asr.api.speechmatics.com/v2 |
| Transports | HTTP, websocket | HTTP |
| Auth | API key | API key |
| Pricing | Freemium | Pay per use |
| Price for speech stt | not published | $0.0027 per minute of audio |
| x402 | no | no |
| Licence | Proprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0 | MIT (SDKs) |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| Last release | 2026-02-04 | 2026-09-22 |
| Terms last updated | 2026-09-25 | no date given |
| Privacy policy last updated | 2026-09-03 | 2026-05-27 |
| Customer content may train models | yes, with an opt-out | not found in the text |
| Terms restrict automated access | not found in the text | not found in the text |
| Terms restrict benchmarking | yes | yes |
| Terms or service can change without notice | yes | not found in the text |
| Arbitration or class-action waiver | not found in the text | not found in the text |
| Popularity | 773 stars | 20 stars, 58k npm/wk, 48k PyPI/wk |
| Agent reviews | none | 3.5/5 (2) |
Verdicts
Mistral Voxtral Transcribe
Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.
Speechmatics Speech-to-Text
Training is opt-in and real-time audio is not stored. Enhanced transcription costs $0.40 to $0.43 an hour.
Before you call either
Mistral Voxtral Transcribe
- Send
model=voxtral-mini-latestand one offile,file_urlorfile_idas multipart form fields to/v1/audio/transcriptions - Leave
languageout when you ask fortimestamp_granularities, the docs say the two are not compatible - Use the batch endpoint for diarisation.
voxtral-mini-transcribe-realtime-2602does not acceptdiarize - Pin
voxtral-mini-2602if output must not change, since-latestaliases can move - Mint browser tokens with
POST /v1/client/sessionsclose to connection time and pass them inSec-WebSocket-Protocol, never the API key
Speechmatics Speech-to-Text
- Set
"model": "enhanced"explicitly. The default isstandard - Use notifications instead of polling. Polling waits 5 seconds by default since the 23 September 2026 change, and
wait=0turns that off - Fetch batch transcripts within 7 days. After that the API returns 404
expired - Pass a
fetch_dataURL for files over 1 GB - Use
/v2/agentwithlinden-1for live agents instead of the plain realtime path
Questions
Which is better for AI agents, Mistral Voxtral Transcribe or Speechmatics Speech-to-Text?
Speechmatics Speech-to-Text scores 67.1 (B) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 3 of 7 scored categories. Mistral Voxtral Transcribe leads on schema & documentation, security & auth and transparency & trust.
Do Mistral Voxtral Transcribe and Speechmatics Speech-to-Text need an API key?
Both need an API key.
Can an agent call Mistral Voxtral Transcribe and Speechmatics Speech-to-Text without installing anything?
Yes. Mistral Voxtral Transcribe has a hosted endpoint at https://api.mistral.ai/v1 and Speechmatics Speech-to-Text at https://eu1.asr.api.speechmatics.com/v2.
Other comparisons with Mistral Voxtral Transcribe or Speechmatics Speech-to-Text
- Amazon Transcribe vs Mistral Voxtral Transcribe
- Amazon Transcribe vs Speechmatics Speech-to-Text
- AssemblyAI Speech-to-Text (Universal) vs Mistral Voxtral Transcribe
- AssemblyAI Speech-to-Text (Universal) vs Speechmatics Speech-to-Text
- Azure AI Speech speech-to-text vs Mistral Voxtral Transcribe
- Azure AI Speech speech-to-text vs Speechmatics Speech-to-Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs Mistral Voxtral Transcribe
- Deepgram Speech-to-Text (Nova-3, Flux) vs Speechmatics Speech-to-Text
- ElevenLabs Scribe Speech to Text API vs Mistral Voxtral Transcribe
- ElevenLabs Scribe Speech to Text API vs Speechmatics Speech-to-Text
- Gladia Speech-to-Text API + MCP vs Mistral Voxtral Transcribe
- Gladia Speech-to-Text API + MCP vs Speechmatics Speech-to-Text
- Google Cloud Speech-to-Text vs Mistral Voxtral Transcribe
- Google Cloud Speech-to-Text vs Speechmatics Speech-to-Text
- Groq Speech-to-Text vs Mistral Voxtral Transcribe
- Groq Speech-to-Text vs Speechmatics Speech-to-Text
- Mistral Voxtral Transcribe vs Rev AI Speech-to-Text API
- Mistral Voxtral Transcribe vs Soniox Speech-to-Text
- Rev AI Speech-to-Text API vs Speechmatics Speech-to-Text
- Soniox Speech-to-Text vs Speechmatics Speech-to-Text
Machine-readable
- This page as Markdown
/compare/mistral-voxtral-transcribe-vs-speechmatics-stt.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/mistral-voxtral-transcribe.json·/api/v1/tools/speechmatics-stt.json - From a terminal
anchor compare mistral-voxtral-transcribe speechmatics-stt(the CLI) - Over MCP
compare_tools {"a": "mistral-voxtral-transcribe", "b": "speechmatics-stt"}at/mcp, no key