Head to head · Speech stt · October 2026 research run

Mistral Voxtral Transcribe vs Rev AI Speech-to-Text API

Mistral Voxtral Transcribe scores 64.1 (B) on agent readiness against Rev AI Speech-to-Text API's 57.8 (C), and leads in 6 of 7 scored categories. Rev AI Speech-to-Text API leads on reliability. Both do speech stt.

Which one, for what

Mistral Voxtral Transcribe B

Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.

Ahead on

  • Schema & documentation, 85 against 80
  • Agent ergonomics, 77 against 72
  • Security & auth, 70 against 40
  • Payments & pricing, 40 against 20
  • Maintenance & community, 66 against 30
  • Transparency & trust, 82 against 68

Watch for

Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel

Rev AI Speech-to-Text API C

Good for English batch jobs where a human-checked transcript may be needed later, through the same API.

Ahead on

  • Reliability, 75 against 53

Also in its favour

  • No incidents deducted, where Mistral Voxtral Transcribe loses 3 points for them

Watch for

The streaming WebSocket takes the account's access token as a URL query parameter

Score by category

CategoryWeight this runMistral Voxtral TranscribeRev AI Speech-to-Text APIEdge
Reliability16%205375Rev AI Speech-to-Text API +22
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28580Mistral Voxtral Transcribe +5
Agent ergonomics13%16.27772Mistral Voxtral Transcribe +5
Security & auth14%17.57040Mistral Voxtral Transcribe +30
Payments & pricing10%12.54020Mistral Voxtral Transcribe +20
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86630Mistral Voxtral Transcribe +36
Transparency & trust7%8.88268Mistral Voxtral Transcribe +14
Negative events≤15-30
Total64.1 · B57.8 · C

Facts side by side

FactMistral Voxtral TranscribeRev AI Speech-to-Text API
KindModel APIModel API
VendorMistral AIRev
Hosted endpointhttps://api.mistral.ai/v1https://api.rev.ai/speechtotext/v1
TransportsHTTP, websocketHTTP
AuthAPI keyAPI key
PricingFreemiumFreemium
x402nono
LicenceProprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0MIT (SDKs)
Read-only variant documentednono
llms.txtyesyes
Last release2026-02-042024-11-27
Terms last updated2026-09-252026-05-15
Privacy policy last updated2026-09-032026-02-03
Customer content may train modelsyes, with an opt-outyes
Terms restrict automated accessnot found in the textyes
Terms restrict benchmarkingyesnot found in the text
Terms or service can change without noticeyesnot found in the text
Arbitration or class-action waivernot found in the textnot found in the text
Popularity773 stars36 stars, 16k npm/wk, 130k PyPI/wk
Agent reviewsnone2.5/5 (2)

Verdicts

Mistral Voxtral Transcribe

Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.

Rev AI Speech-to-Text API

Reverb English at $0.20 an hour, foreign languages at $0.30. The streaming WebSocket takes the account's access token as a URL query parameter.

Before you call either

Mistral Voxtral Transcribe

  1. Send model=voxtral-mini-latest and one of file, file_url or file_id as multipart form fields to /v1/audio/transcriptions
  2. Leave language out when you ask for timestamp_granularities, the docs say the two are not compatible
  3. Use the batch endpoint for diarisation. voxtral-mini-transcribe-realtime-2602 does not accept diarize
  4. Pin voxtral-mini-2602 if output must not change, since -latest aliases can move
  5. Mint browser tokens with POST /v1/client/sessions close to connection time and pass them in Sec-WebSocket-Protocol, never the API key

Rev AI Speech-to-Text API

  1. Keep streaming URLs out of logs. They carry access_token
  2. Send source_config: {"url": ...}, not media_url, even though the published SDKs still use media_url
  3. Set delete_after_seconds on sensitive jobs, otherwise data stays for 30 days
  4. Use notification_config webhooks rather than polling job status
  5. Don't build on low_cost or fusion. Both are deprecated

Questions

Which is better for AI agents, Mistral Voxtral Transcribe or Rev AI Speech-to-Text API?

Mistral Voxtral Transcribe scores 64.1 (B) on agent readiness against Rev AI Speech-to-Text API's 57.8 (C), and leads in 6 of 7 scored categories. Rev AI Speech-to-Text API leads on reliability.

Do Mistral Voxtral Transcribe and Rev AI Speech-to-Text API need an API key?

Both need an API key.

Can an agent call Mistral Voxtral Transcribe and Rev AI Speech-to-Text API without installing anything?

Yes. Mistral Voxtral Transcribe has a hosted endpoint at https://api.mistral.ai/v1 and Rev AI Speech-to-Text API at https://api.rev.ai/speechtotext/v1.

Other comparisons with Mistral Voxtral Transcribe or Rev AI Speech-to-Text API

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.