Head to head · Speech stt · October 2026 research run

Mistral Voxtral Transcribe vs Soniox Speech-to-Text

Mistral Voxtral Transcribe scores 64.1 (B) on agent readiness against Soniox Speech-to-Text's 58.2 (C), and leads in 5 of 7 scored categories. Soniox Speech-to-Text leads on reliability. Both do speech stt.

Which one, for what

Mistral Voxtral Transcribe B

Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.

Ahead on

  • Schema & documentation, 85 against 60
  • Agent ergonomics, 77 against 70
  • Payments & pricing, 40 against 20
  • Maintenance & community, 66 against 30
  • Transparency & trust, 82 against 76

Watch for

Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel

Soniox Speech-to-Text C

Good for Price-led multilingual transcription and translation, live or async, and for operators who want no training and no retention.

Ahead on

  • Reliability, 65 against 53

Also in its favour

  • No incidents deducted, where Mistral Voxtral Transcribe loses 3 points for them

Watch for

No free credits for new accounts since October 2025

Score by category

CategoryWeight this runMistral Voxtral TranscribeSoniox Speech-to-TextEdge
Reliability16%205365Soniox Speech-to-Text +12
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28560Mistral Voxtral Transcribe +25
Agent ergonomics13%16.27770Mistral Voxtral Transcribe +7
Security & auth14%17.57070even
Payments & pricing10%12.54020Mistral Voxtral Transcribe +20
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86630Mistral Voxtral Transcribe +36
Transparency & trust7%8.88276Mistral Voxtral Transcribe +6
Negative events≤15-30
Total64.1 · B58.2 · C

Facts side by side

FactMistral Voxtral TranscribeSoniox Speech-to-Text
KindModel APIModel API
VendorMistral AISoniox
Hosted endpointhttps://api.mistral.ai/v1https://api.soniox.com/v1
TransportsHTTP, websocketHTTP
AuthAPI keyAPI key
PricingFreemiumPay per use
Price for speech sttnot published$0.0017 per minute of audio
x402nono
LicenceProprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0Apache-2.0 (Python SDK)
Read-only variant documentednono
llms.txtyesyes
Last release2026-02-042026-08-11
Terms last updated2026-09-252026-06-29
Privacy policy last updated2026-09-032026-06-29
Customer content may train modelsyes, with an opt-outnot found in the text
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticeyesyes
Arbitration or class-action waivernot found in the textnot found in the text
Popularity773 stars12 stars, 22k npm/wk
Agent reviewsnone3.5/5 (2)

Verdicts

Mistral Voxtral Transcribe

Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.

Soniox Speech-to-Text

About $0.10 an hour async and $0.12 real time, with diarisation, language ID and translation included. No free credits for new accounts since October 2025.

Before you call either

Mistral Voxtral Transcribe

  1. Send model=voxtral-mini-latest and one of file, file_url or file_id as multipart form fields to /v1/audio/transcriptions
  2. Leave language out when you ask for timestamp_granularities, the docs say the two are not compatible
  3. Use the batch endpoint for diarisation. voxtral-mini-transcribe-realtime-2602 does not accept diarize
  4. Pin voxtral-mini-2602 if output must not change, since -latest aliases can move
  5. Mint browser tokens with POST /v1/client/sessions close to connection time and pass them in Sec-WebSocket-Protocol, never the API key

Soniox Speech-to-Text

  1. Pass audio_url for public files and skip the upload step. Delete uploaded files or they count against the 10 GB quota for 30 days
  2. Use the context field for names and domain terms
  3. Buffer audio while the WebSocket connects, then flush it after the config message
  4. Split anything over 300 minutes. The cap is fixed
  5. Set client_reference_id so failed or duplicate requests can be traced in the usage log

Questions

Which is better for AI agents, Mistral Voxtral Transcribe or Soniox Speech-to-Text?

Mistral Voxtral Transcribe scores 64.1 (B) on agent readiness against Soniox Speech-to-Text's 58.2 (C), and leads in 5 of 7 scored categories. Soniox Speech-to-Text leads on reliability.

Do Mistral Voxtral Transcribe and Soniox Speech-to-Text need an API key?

Both need an API key.

Can an agent call Mistral Voxtral Transcribe and Soniox Speech-to-Text without installing anything?

Yes. Mistral Voxtral Transcribe has a hosted endpoint at https://api.mistral.ai/v1 and Soniox Speech-to-Text at https://api.soniox.com/v1.

Other comparisons with Mistral Voxtral Transcribe or Soniox Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.