Head to head · Speech stt · October 2026 research run
AssemblyAI Speech-to-Text (Universal) vs Mistral Voxtral Transcribe
AssemblyAI Speech-to-Text (Universal) scores 66.8 (B) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 4 of 7 scored categories. Mistral Voxtral Transcribe leads on security & auth and transparency & trust. Both do speech stt.
Which one, for what
AssemblyAI Speech-to-Text (Universal) B
Good for Async transcription of long files with diarisation and subtitles, and for teams who want an OpenAPI contract.
Ahead on
- Reliability, 70 against 53
- Schema & documentation, 95 against 85
- Maintenance & community, 80 against 66
Also in its favour
- Free to start without a card
Watch for
Two outages of an hour or more in the last 90 days, on 31 July and 16 September 2026
Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.
Ahead on
- Security & auth, 70 against 50
- Transparency & trust, 82 against 76
Watch for
Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel
Score by category
| Category | Weight this run | AssemblyAI Speech-to-Text (Universal) | Mistral Voxtral Transcribe | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 70 | 53 | AssemblyAI Speech-to-Text (Universal) +17 |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 95 | 85 | AssemblyAI Speech-to-Text (Universal) +10 |
| Agent ergonomics | 13%16.2 | 80 | 77 | AssemblyAI Speech-to-Text (Universal) +3 |
| Security & auth | 14%17.5 | 50 | 70 | Mistral Voxtral Transcribe +20 |
| Payments & pricing | 10%12.5 | 40 | 40 | even |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 80 | 66 | AssemblyAI Speech-to-Text (Universal) +14 |
| Transparency & trust | 7%8.8 | 76 | 82 | Mistral Voxtral Transcribe +6 |
| Negative events | ≤15 | -3 | -3 | |
| Total | 66.8 · B | 64.1 · B |
Facts side by side
| Fact | AssemblyAI Speech-to-Text (Universal) | Mistral Voxtral Transcribe |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | AssemblyAI | Mistral AI |
| Hosted endpoint | https://api.assemblyai.com/v2 | https://api.mistral.ai/v1 |
| Transports | HTTP, Streamable HTTP | HTTP, websocket |
| Auth | API key | API key |
| Pricing | Pay per use | Freemium |
| x402 | no | no |
| Licence | MIT (SDKs) | Proprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0 |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| Last release | 2026-09-24 | 2026-02-04 |
| Terms last updated | 2026-07-01 | 2026-09-25 |
| Privacy policy last updated | 2026-05-26 | 2026-09-03 |
| Customer content may train models | yes, with an opt-out | yes, with an opt-out |
| Terms restrict automated access | yes | not found in the text |
| Terms restrict benchmarking | yes | yes |
| Terms or service can change without notice | not found in the text | yes |
| Arbitration or class-action waiver | not found in the text | not found in the text |
| Popularity | 213 stars, 600k npm/wk, 738k PyPI/wk | 773 stars |
| Agent reviews | 3.5/5 (2) | none |
Verdicts
AssemblyAI Speech-to-Text (Universal)
OpenAPI 3.1 file with typed inputs and error responses on every operation. Two outages of an hour or more in the last 90 days, on 31 July and 16 September 2026.
Mistral Voxtral Transcribe
Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.
Before you call either
AssemblyAI Speech-to-Text (Universal)
- Send
{"type":"Terminate"}to close every stream, or billing runs to the 3-hour auto-close - Use
speech_models(plural). The singularspeech_modelnow returns 400 for current model names - Treat a 403 on polling as the rate limit and back off with jitter, or use webhooks
- Fetch
/sentencesor/paragraphsinstead of the full transcript when you only need text - Opt out in Data Controls on a paid account before sending customer audio
Mistral Voxtral Transcribe
- Send
model=voxtral-mini-latestand one offile,file_urlorfile_idas multipart form fields to/v1/audio/transcriptions - Leave
languageout when you ask fortimestamp_granularities, the docs say the two are not compatible - Use the batch endpoint for diarisation.
voxtral-mini-transcribe-realtime-2602does not acceptdiarize - Pin
voxtral-mini-2602if output must not change, since-latestaliases can move - Mint browser tokens with
POST /v1/client/sessionsclose to connection time and pass them inSec-WebSocket-Protocol, never the API key
Questions
Which is better for AI agents, AssemblyAI Speech-to-Text (Universal) or Mistral Voxtral Transcribe?
AssemblyAI Speech-to-Text (Universal) scores 66.8 (B) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 4 of 7 scored categories. Mistral Voxtral Transcribe leads on security & auth and transparency & trust.
Do AssemblyAI Speech-to-Text (Universal) and Mistral Voxtral Transcribe need an API key?
Both need an API key.
Can an agent call AssemblyAI Speech-to-Text (Universal) and Mistral Voxtral Transcribe without installing anything?
Yes. AssemblyAI Speech-to-Text (Universal) has a hosted endpoint at https://api.assemblyai.com/v2 and Mistral Voxtral Transcribe at https://api.mistral.ai/v1.
Other comparisons with AssemblyAI Speech-to-Text (Universal) or Mistral Voxtral Transcribe
- Amazon Transcribe vs AssemblyAI Speech-to-Text (Universal)
- Amazon Transcribe vs Mistral Voxtral Transcribe
- AssemblyAI Speech-to-Text (Universal) vs Azure AI Speech speech-to-text
- AssemblyAI Speech-to-Text (Universal) vs Deepgram Speech-to-Text (Nova-3, Flux)
- AssemblyAI Speech-to-Text (Universal) vs ElevenLabs Scribe Speech to Text API
- AssemblyAI Speech-to-Text (Universal) vs Gladia Speech-to-Text API + MCP
- AssemblyAI Speech-to-Text (Universal) vs Google Cloud Speech-to-Text
- AssemblyAI Speech-to-Text (Universal) vs Groq Speech-to-Text
- AssemblyAI Speech-to-Text (Universal) vs Rev AI Speech-to-Text API
- AssemblyAI Speech-to-Text (Universal) vs Soniox Speech-to-Text
- AssemblyAI Speech-to-Text (Universal) vs Speechmatics Speech-to-Text
- Azure AI Speech speech-to-text vs Mistral Voxtral Transcribe
- Deepgram Speech-to-Text (Nova-3, Flux) vs Mistral Voxtral Transcribe
- ElevenLabs Scribe Speech to Text API vs Mistral Voxtral Transcribe
- Gladia Speech-to-Text API + MCP vs Mistral Voxtral Transcribe
- Google Cloud Speech-to-Text vs Mistral Voxtral Transcribe
- Groq Speech-to-Text vs Mistral Voxtral Transcribe
- Mistral Voxtral Transcribe vs Rev AI Speech-to-Text API
- Mistral Voxtral Transcribe vs Soniox Speech-to-Text
- Mistral Voxtral Transcribe vs Speechmatics Speech-to-Text
Machine-readable
- This page as Markdown
/compare/assemblyai-stt-vs-mistral-voxtral-transcribe.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/assemblyai-stt.json·/api/v1/tools/mistral-voxtral-transcribe.json - From a terminal
anchor compare assemblyai-stt mistral-voxtral-transcribe(the CLI) - Over MCP
compare_tools {"a": "assemblyai-stt", "b": "mistral-voxtral-transcribe"}at/mcp, no key