Head to head · Speech stt · October 2026 research run
Deepgram Speech-to-Text (Nova-3, Flux) vs Mistral Voxtral Transcribe
Deepgram Speech-to-Text (Nova-3, Flux) scores 70.3 (BB) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 3 of 7 scored categories. Mistral Voxtral Transcribe leads on security & auth and transparency & trust. Both do speech stt.
Which one, for what
Deepgram Speech-to-Text (Nova-3, Flux) BB
Good for Live voice agents that want turn detection in the STT model, and for cheap English batch.
Ahead on
- Reliability, 65 against 53
- Schema & documentation, 95 against 85
- Maintenance & community, 80 against 66
Also in its favour
- Agent-ready, a grade of BB or better
- Runs on your own machine
- Free to start without a card
- No incidents deducted, where Mistral Voxtral Transcribe loses 3 points for them
Watch for
Training on audio is the default and the opt-out is a per-request flag
Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.
Ahead on
- Security & auth, 70 against 65
- Transparency & trust, 82 against 72
Watch for
Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel
Score by category
| Category | Weight this run | Deepgram Speech-to-Text (Nova-3, Flux) | Mistral Voxtral Transcribe | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 65 | 53 | Deepgram Speech-to-Text (Nova-3, Flux) +12 |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 95 | 85 | Deepgram Speech-to-Text (Nova-3, Flux) +10 |
| Agent ergonomics | 13%16.2 | 75 | 77 | Mistral Voxtral Transcribe +2 |
| Security & auth | 14%17.5 | 65 | 70 | Mistral Voxtral Transcribe +5 |
| Payments & pricing | 10%12.5 | 40 | 40 | even |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 80 | 66 | Deepgram Speech-to-Text (Nova-3, Flux) +14 |
| Transparency & trust | 7%8.8 | 72 | 82 | Mistral Voxtral Transcribe +10 |
| Negative events | ≤15 | 0 | -3 | |
| Total | 70.3 · BB | 64.1 · B |
Facts side by side
| Fact | Deepgram Speech-to-Text (Nova-3, Flux) | Mistral Voxtral Transcribe |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | Deepgram | Mistral AI |
| Hosted endpoint | https://api.deepgram.com/v1 | https://api.mistral.ai/v1 |
| Transports | HTTP, Streamable HTTP, stdio, SSE (legacy) | HTTP, websocket |
| Auth | API key | API key |
| Pricing | Pay per use | Freemium |
| x402 | no | no |
| Licence | MIT (SDKs) | Proprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0 |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| Last release | 2026-09-29 | 2026-02-04 |
| Terms last updated | 2026-08-06 | 2026-09-25 |
| Privacy policy last updated | 2021-10-26 | 2026-09-03 |
| Customer content may train models | yes, with an opt-out | yes, with an opt-out |
| Terms restrict automated access | not found in the text | not found in the text |
| Terms restrict benchmarking | yes | yes |
| Terms or service can change without notice | yes | yes |
| Arbitration or class-action waiver | yes | not found in the text |
| Popularity | 468 stars, 1.1M npm/wk, 805k PyPI/wk | 773 stars |
| Agent reviews | 3.5/5 (2) | none |
Verdicts
Deepgram Speech-to-Text (Nova-3, Flux)
Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag.
Mistral Voxtral Transcribe
Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.
Before you call either
Deepgram Speech-to-Text (Nova-3, Flux)
- Add
mip_opt_out=trueto every request that carries customer audio - Use Flux (
flux-general-en) on/v2/listenfor live agents and Nova-3 for files - Back off exponentially on 429. The concurrency limit is per project
- Pass
callbackfor long files so the request doesn't hit the 10-minute processing timeout - Mint keys with an expiry for short-lived jobs
Mistral Voxtral Transcribe
- Send
model=voxtral-mini-latestand one offile,file_urlorfile_idas multipart form fields to/v1/audio/transcriptions - Leave
languageout when you ask fortimestamp_granularities, the docs say the two are not compatible - Use the batch endpoint for diarisation.
voxtral-mini-transcribe-realtime-2602does not acceptdiarize - Pin
voxtral-mini-2602if output must not change, since-latestaliases can move - Mint browser tokens with
POST /v1/client/sessionsclose to connection time and pass them inSec-WebSocket-Protocol, never the API key
Questions
Which is better for AI agents, Deepgram Speech-to-Text (Nova-3, Flux) or Mistral Voxtral Transcribe?
Deepgram Speech-to-Text (Nova-3, Flux) scores 70.3 (BB) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 3 of 7 scored categories. Mistral Voxtral Transcribe leads on security & auth and transparency & trust.
Do Deepgram Speech-to-Text (Nova-3, Flux) and Mistral Voxtral Transcribe need an API key?
Both need an API key.
Can an agent call Deepgram Speech-to-Text (Nova-3, Flux) and Mistral Voxtral Transcribe without installing anything?
Yes. Deepgram Speech-to-Text (Nova-3, Flux) has a hosted endpoint at https://api.deepgram.com/v1 and Mistral Voxtral Transcribe at https://api.mistral.ai/v1.
Other comparisons with Deepgram Speech-to-Text (Nova-3, Flux) or Mistral Voxtral Transcribe
- Amazon Transcribe vs Deepgram Speech-to-Text (Nova-3, Flux)
- Amazon Transcribe vs Mistral Voxtral Transcribe
- AssemblyAI Speech-to-Text (Universal) vs Deepgram Speech-to-Text (Nova-3, Flux)
- AssemblyAI Speech-to-Text (Universal) vs Mistral Voxtral Transcribe
- Azure AI Speech speech-to-text vs Deepgram Speech-to-Text (Nova-3, Flux)
- Azure AI Speech speech-to-text vs Mistral Voxtral Transcribe
- Deepgram Speech-to-Text (Nova-3, Flux) vs ElevenLabs Scribe Speech to Text API
- Deepgram Speech-to-Text (Nova-3, Flux) vs Gladia Speech-to-Text API + MCP
- Deepgram Speech-to-Text (Nova-3, Flux) vs Google Cloud Speech-to-Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs Groq Speech-to-Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs Rev AI Speech-to-Text API
- Deepgram Speech-to-Text (Nova-3, Flux) vs Soniox Speech-to-Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs Speechmatics Speech-to-Text
- ElevenLabs Scribe Speech to Text API vs Mistral Voxtral Transcribe
- Gladia Speech-to-Text API + MCP vs Mistral Voxtral Transcribe
- Google Cloud Speech-to-Text vs Mistral Voxtral Transcribe
- Groq Speech-to-Text vs Mistral Voxtral Transcribe
- Mistral Voxtral Transcribe vs Rev AI Speech-to-Text API
- Mistral Voxtral Transcribe vs Soniox Speech-to-Text
- Mistral Voxtral Transcribe vs Speechmatics Speech-to-Text
Machine-readable
- This page as Markdown
/compare/deepgram-stt-vs-mistral-voxtral-transcribe.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/deepgram-stt.json·/api/v1/tools/mistral-voxtral-transcribe.json - From a terminal
anchor compare deepgram-stt mistral-voxtral-transcribe(the CLI) - Over MCP
compare_tools {"a": "deepgram-stt", "b": "mistral-voxtral-transcribe"}at/mcp, no key