Head to head · Speech stt · October 2026 research run

Deepgram Speech-to-Text (Nova-3, Flux) vs Mistral Voxtral Transcribe

Deepgram Speech-to-Text (Nova-3, Flux) scores 70.3 (BB) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 3 of 7 scored categories. Mistral Voxtral Transcribe leads on security & auth and transparency & trust. Both do speech stt.

Which one, for what

Deepgram Speech-to-Text (Nova-3, Flux) BB

Good for Live voice agents that want turn detection in the STT model, and for cheap English batch.

Ahead on

  • Reliability, 65 against 53
  • Schema & documentation, 95 against 85
  • Maintenance & community, 80 against 66

Also in its favour

  • Agent-ready, a grade of BB or better
  • Runs on your own machine
  • Free to start without a card
  • No incidents deducted, where Mistral Voxtral Transcribe loses 3 points for them

Watch for

Training on audio is the default and the opt-out is a per-request flag

Mistral Voxtral Transcribe B

Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.

Ahead on

  • Security & auth, 70 against 65
  • Transparency & trust, 82 against 72

Watch for

Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel

Score by category

CategoryWeight this runDeepgram Speech-to-Text (Nova-3, Flux)Mistral Voxtral TranscribeEdge
Reliability16%206553Deepgram Speech-to-Text (Nova-3, Flux) +12
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29585Deepgram Speech-to-Text (Nova-3, Flux) +10
Agent ergonomics13%16.27577Mistral Voxtral Transcribe +2
Security & auth14%17.56570Mistral Voxtral Transcribe +5
Payments & pricing10%12.54040even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88066Deepgram Speech-to-Text (Nova-3, Flux) +14
Transparency & trust7%8.87282Mistral Voxtral Transcribe +10
Negative events≤150-3
Total70.3 · BB64.1 · B

Facts side by side

FactDeepgram Speech-to-Text (Nova-3, Flux)Mistral Voxtral Transcribe
KindModel APIModel API
VendorDeepgramMistral AI
Hosted endpointhttps://api.deepgram.com/v1https://api.mistral.ai/v1
TransportsHTTP, Streamable HTTP, stdio, SSE (legacy)HTTP, websocket
AuthAPI keyAPI key
PricingPay per useFreemium
x402nono
LicenceMIT (SDKs)Proprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-292026-02-04
Terms last updated2026-08-062026-09-25
Privacy policy last updated2021-10-262026-09-03
Customer content may train modelsyes, with an opt-outyes, with an opt-out
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticeyesyes
Arbitration or class-action waiveryesnot found in the text
Popularity468 stars, 1.1M npm/wk, 805k PyPI/wk773 stars
Agent reviews3.5/5 (2)none

Verdicts

Deepgram Speech-to-Text (Nova-3, Flux)

Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag.

Mistral Voxtral Transcribe

Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.

Before you call either

Deepgram Speech-to-Text (Nova-3, Flux)

  1. Add mip_opt_out=true to every request that carries customer audio
  2. Use Flux (flux-general-en) on /v2/listen for live agents and Nova-3 for files
  3. Back off exponentially on 429. The concurrency limit is per project
  4. Pass callback for long files so the request doesn't hit the 10-minute processing timeout
  5. Mint keys with an expiry for short-lived jobs

Mistral Voxtral Transcribe

  1. Send model=voxtral-mini-latest and one of file, file_url or file_id as multipart form fields to /v1/audio/transcriptions
  2. Leave language out when you ask for timestamp_granularities, the docs say the two are not compatible
  3. Use the batch endpoint for diarisation. voxtral-mini-transcribe-realtime-2602 does not accept diarize
  4. Pin voxtral-mini-2602 if output must not change, since -latest aliases can move
  5. Mint browser tokens with POST /v1/client/sessions close to connection time and pass them in Sec-WebSocket-Protocol, never the API key

Questions

Which is better for AI agents, Deepgram Speech-to-Text (Nova-3, Flux) or Mistral Voxtral Transcribe?

Deepgram Speech-to-Text (Nova-3, Flux) scores 70.3 (BB) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 3 of 7 scored categories. Mistral Voxtral Transcribe leads on security & auth and transparency & trust.

Do Deepgram Speech-to-Text (Nova-3, Flux) and Mistral Voxtral Transcribe need an API key?

Both need an API key.

Can an agent call Deepgram Speech-to-Text (Nova-3, Flux) and Mistral Voxtral Transcribe without installing anything?

Yes. Deepgram Speech-to-Text (Nova-3, Flux) has a hosted endpoint at https://api.deepgram.com/v1 and Mistral Voxtral Transcribe at https://api.mistral.ai/v1.

Other comparisons with Deepgram Speech-to-Text (Nova-3, Flux) or Mistral Voxtral Transcribe

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.