Head to head · Speech stt · October 2026 research run

Mistral Voxtral Transcribe vs Speechmatics Speech-to-Text

Speechmatics Speech-to-Text scores 67.1 (B) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 3 of 7 scored categories. Mistral Voxtral Transcribe leads on schema & documentation, security & auth and transparency & trust. Both do speech stt.

Which one, for what

Mistral Voxtral Transcribe B

Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.

Ahead on

  • Schema & documentation, 85 against 65
  • Security & auth, 70 against 65
  • Transparency & trust, 82 against 75

Watch for

Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel

Speechmatics Speech-to-Text B

Good for Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns.

Ahead on

  • Reliability, 70 against 53
  • Maintenance & community, 75 against 66

Also in its favour

  • Free to start without a card
  • No incidents deducted, where Mistral Voxtral Transcribe loses 3 points for them

Watch for

Enhanced costs $0.40 to $0.43 an hour, above most rivals

Score by category

CategoryWeight this runMistral Voxtral TranscribeSpeechmatics Speech-to-TextEdge
Reliability16%205370Speechmatics Speech-to-Text +17
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28565Mistral Voxtral Transcribe +20
Agent ergonomics13%16.27780Speechmatics Speech-to-Text +3
Security & auth14%17.57065Mistral Voxtral Transcribe +5
Payments & pricing10%12.54040even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86675Speechmatics Speech-to-Text +9
Transparency & trust7%8.88275Mistral Voxtral Transcribe +7
Negative events≤15-30
Total64.1 · B67.1 · B

Facts side by side

FactMistral Voxtral TranscribeSpeechmatics Speech-to-Text
KindModel APIModel API
VendorMistral AISpeechmatics
Hosted endpointhttps://api.mistral.ai/v1https://eu1.asr.api.speechmatics.com/v2
TransportsHTTP, websocketHTTP
AuthAPI keyAPI key
PricingFreemiumPay per use
Price for speech sttnot published$0.0027 per minute of audio
x402nono
LicenceProprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0MIT (SDKs)
Read-only variant documentednono
llms.txtyesyes
Last release2026-02-042026-09-22
Terms last updated2026-09-25no date given
Privacy policy last updated2026-09-032026-05-27
Customer content may train modelsyes, with an opt-outnot found in the text
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticeyesnot found in the text
Arbitration or class-action waivernot found in the textnot found in the text
Popularity773 stars20 stars, 58k npm/wk, 48k PyPI/wk
Agent reviewsnone3.5/5 (2)

Verdicts

Mistral Voxtral Transcribe

Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.

Speechmatics Speech-to-Text

Training is opt-in and real-time audio is not stored. Enhanced transcription costs $0.40 to $0.43 an hour.

Before you call either

Mistral Voxtral Transcribe

  1. Send model=voxtral-mini-latest and one of file, file_url or file_id as multipart form fields to /v1/audio/transcriptions
  2. Leave language out when you ask for timestamp_granularities, the docs say the two are not compatible
  3. Use the batch endpoint for diarisation. voxtral-mini-transcribe-realtime-2602 does not accept diarize
  4. Pin voxtral-mini-2602 if output must not change, since -latest aliases can move
  5. Mint browser tokens with POST /v1/client/sessions close to connection time and pass them in Sec-WebSocket-Protocol, never the API key

Speechmatics Speech-to-Text

  1. Set "model": "enhanced" explicitly. The default is standard
  2. Use notifications instead of polling. Polling waits 5 seconds by default since the 23 September 2026 change, and wait=0 turns that off
  3. Fetch batch transcripts within 7 days. After that the API returns 404 expired
  4. Pass a fetch_data URL for files over 1 GB
  5. Use /v2/agent with linden-1 for live agents instead of the plain realtime path

Questions

Which is better for AI agents, Mistral Voxtral Transcribe or Speechmatics Speech-to-Text?

Speechmatics Speech-to-Text scores 67.1 (B) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 3 of 7 scored categories. Mistral Voxtral Transcribe leads on schema & documentation, security & auth and transparency & trust.

Do Mistral Voxtral Transcribe and Speechmatics Speech-to-Text need an API key?

Both need an API key.

Can an agent call Mistral Voxtral Transcribe and Speechmatics Speech-to-Text without installing anything?

Yes. Mistral Voxtral Transcribe has a hosted endpoint at https://api.mistral.ai/v1 and Speechmatics Speech-to-Text at https://eu1.asr.api.speechmatics.com/v2.

Other comparisons with Mistral Voxtral Transcribe or Speechmatics Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.