Head to head · Speech stt · October 2026 research run

AssemblyAI Speech-to-Text (Universal) vs Mistral Voxtral Transcribe

AssemblyAI Speech-to-Text (Universal) scores 66.8 (B) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 4 of 7 scored categories. Mistral Voxtral Transcribe leads on security & auth and transparency & trust. Both do speech stt.

Which one, for what

AssemblyAI Speech-to-Text (Universal) B

Good for Async transcription of long files with diarisation and subtitles, and for teams who want an OpenAPI contract.

Ahead on

  • Reliability, 70 against 53
  • Schema & documentation, 95 against 85
  • Maintenance & community, 80 against 66

Also in its favour

  • Free to start without a card

Watch for

Two outages of an hour or more in the last 90 days, on 31 July and 16 September 2026

Mistral Voxtral Transcribe B

Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.

Ahead on

  • Security & auth, 70 against 50
  • Transparency & trust, 82 against 76

Watch for

Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel

Score by category

CategoryWeight this runAssemblyAI Speech-to-Text (Universal)Mistral Voxtral TranscribeEdge
Reliability16%207053AssemblyAI Speech-to-Text (Universal) +17
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29585AssemblyAI Speech-to-Text (Universal) +10
Agent ergonomics13%16.28077AssemblyAI Speech-to-Text (Universal) +3
Security & auth14%17.55070Mistral Voxtral Transcribe +20
Payments & pricing10%12.54040even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88066AssemblyAI Speech-to-Text (Universal) +14
Transparency & trust7%8.87682Mistral Voxtral Transcribe +6
Negative events≤15-3-3
Total66.8 · B64.1 · B

Facts side by side

FactAssemblyAI Speech-to-Text (Universal)Mistral Voxtral Transcribe
KindModel APIModel API
VendorAssemblyAIMistral AI
Hosted endpointhttps://api.assemblyai.com/v2https://api.mistral.ai/v1
TransportsHTTP, Streamable HTTPHTTP, websocket
AuthAPI keyAPI key
PricingPay per useFreemium
x402nono
LicenceMIT (SDKs)Proprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-242026-02-04
Terms last updated2026-07-012026-09-25
Privacy policy last updated2026-05-262026-09-03
Customer content may train modelsyes, with an opt-outyes, with an opt-out
Terms restrict automated accessyesnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textyes
Arbitration or class-action waivernot found in the textnot found in the text
Popularity213 stars, 600k npm/wk, 738k PyPI/wk773 stars
Agent reviews3.5/5 (2)none

Verdicts

AssemblyAI Speech-to-Text (Universal)

OpenAPI 3.1 file with typed inputs and error responses on every operation. Two outages of an hour or more in the last 90 days, on 31 July and 16 September 2026.

Mistral Voxtral Transcribe

Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.

Before you call either

AssemblyAI Speech-to-Text (Universal)

  1. Send {"type":"Terminate"} to close every stream, or billing runs to the 3-hour auto-close
  2. Use speech_models (plural). The singular speech_model now returns 400 for current model names
  3. Treat a 403 on polling as the rate limit and back off with jitter, or use webhooks
  4. Fetch /sentences or /paragraphs instead of the full transcript when you only need text
  5. Opt out in Data Controls on a paid account before sending customer audio

Mistral Voxtral Transcribe

  1. Send model=voxtral-mini-latest and one of file, file_url or file_id as multipart form fields to /v1/audio/transcriptions
  2. Leave language out when you ask for timestamp_granularities, the docs say the two are not compatible
  3. Use the batch endpoint for diarisation. voxtral-mini-transcribe-realtime-2602 does not accept diarize
  4. Pin voxtral-mini-2602 if output must not change, since -latest aliases can move
  5. Mint browser tokens with POST /v1/client/sessions close to connection time and pass them in Sec-WebSocket-Protocol, never the API key

Questions

Which is better for AI agents, AssemblyAI Speech-to-Text (Universal) or Mistral Voxtral Transcribe?

AssemblyAI Speech-to-Text (Universal) scores 66.8 (B) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 4 of 7 scored categories. Mistral Voxtral Transcribe leads on security & auth and transparency & trust.

Do AssemblyAI Speech-to-Text (Universal) and Mistral Voxtral Transcribe need an API key?

Both need an API key.

Can an agent call AssemblyAI Speech-to-Text (Universal) and Mistral Voxtral Transcribe without installing anything?

Yes. AssemblyAI Speech-to-Text (Universal) has a hosted endpoint at https://api.assemblyai.com/v2 and Mistral Voxtral Transcribe at https://api.mistral.ai/v1.

Other comparisons with AssemblyAI Speech-to-Text (Universal) or Mistral Voxtral Transcribe

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.