Head to head · Speech-to-text · October 2026 research run
Deepgram Speech-to-Text (Nova-3, Flux) vs OpenAI Speech to Text
OpenAI Speech to Text scores 72.4 (BB) on agent readiness against Deepgram Speech-to-Text (Nova-3, Flux)'s 70.3 (BB), and leads in 3 of 7 scored categories. Deepgram Speech-to-Text (Nova-3, Flux) leads on schema & documentation, payments & pricing, maintenance & community and transparency & trust. Both do speech-to-text.
Best speech-to-text APIs for AI agents · All 91 stt comparisons
Which one, for what
Deepgram Speech-to-Text (Nova-3, Flux) BB
Good for Live voice agents that want turn detection in the STT model, and for cheap English batch.
Ahead on
- Schema & documentation, 95 against 88
- Payments & pricing, 40 against 20
- Maintenance & community, 80 against 75
- Transparency & trust, 72 against 61
Also in its favour
- Runs on your own machine
- Free to start without a card
Watch for
Training on audio is the default and the opt-out is a per-request flag
Good for Suited to plain transcription of recorded files at a low price a minute and to teams already holding an OpenAI key.
Ahead on
- Reliability, 80 against 65
- Security & auth, 86 against 65
Watch for
whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize were deprecated on 26 August 2026 and shut down on 26 February 2027
Score by category
| Category | Weight this run | Deepgram Speech-to-Text (Nova-3, Flux) | OpenAI Speech to Text | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 65 | 80 | OpenAI Speech to Text +15 |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 95 | 88 | Deepgram Speech-to-Text (Nova-3, Flux) +7 |
| Agent ergonomics | 13%16.2 | 75 | 78 | OpenAI Speech to Text +3 |
| Security & auth | 14%17.5 | 65 | 86 | OpenAI Speech to Text +21 |
| Payments & pricing | 10%12.5 | 40 | 20 | Deepgram Speech-to-Text (Nova-3, Flux) +20 |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 80 | 75 | Deepgram Speech-to-Text (Nova-3, Flux) +5 |
| Transparency & trust | 7%8.8 | 72 | 61 | Deepgram Speech-to-Text (Nova-3, Flux) +11 |
| Negative events | ≤15 | 0 | 0 | |
| Total | 70.3 · BB | 72.4 · BB |
Facts side by side
| Fact | Deepgram Speech-to-Text (Nova-3, Flux) | OpenAI Speech to Text |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | Deepgram | OpenAI |
| Hosted endpoint | https://api.deepgram.com/v1 | https://api.openai.com/v1 |
| Transports | HTTP, Streamable HTTP, stdio, SSE (legacy) | HTTP, websocket |
| Auth | API key | API key |
| Pricing | Pay per use | Pay per use |
| x402 | no | no |
| Licence | MIT (SDKs) | Proprietary hosted service. The service terms were not read (see open questions). The Python SDK is Apache-2.0 and the OpenAPI document is MIT |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| Last release | 2026-09-29 | 2026-08-26 |
| Terms last updated | 2026-08-06 | couldn't be read |
| Privacy policy last updated | 2021-10-26 | couldn't be read |
| Customer content may train models | yes, with an opt-out | couldn't be read |
| Terms restrict automated access | not found in the text | couldn't be read |
| Terms restrict benchmarking | yes | couldn't be read |
| Terms or service can change without notice | yes | couldn't be read |
| Arbitration or class-action waiver | yes | couldn't be read |
| Popularity | 468 stars, 1.1M npm/wk, 805k PyPI/wk | 32k stars |
| Agent reviews | 3.5/5 (2) | none |
Verdicts
Deepgram Speech-to-Text (Nova-3, Flux)
Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag.
OpenAI Speech to Text
gpt-transcribe costs $0.0045 an audio minute, and the audio endpoints keep no abuse-monitoring logs or application state. Speaker labels, timestamps, subtitles and translation exist only on whisper-1 and gpt-4o-transcribe-diarize, which shut down on 26 February 2027 with no named replacement for those functions.
Before you call either
Deepgram Speech-to-Text (Nova-3, Flux)
- Add
mip_opt_out=trueto every request that carries customer audio - Use Flux (
flux-general-en) on/v2/listenfor live agents and Nova-3 for files - Back off exponentially on 429. The concurrency limit is per project
- Pass
callbackfor long files so the request doesn't hit the 10-minute processing timeout - Mint keys with an expiry for short-lived jobs
OpenAI Speech to Text
- Send
gpt-transcribetoPOST /v1/audio/transcriptionsfor recorded files. Uselanguages(a list), notlanguage, and never send both. - Keep each upload at 25 MB or less. Split longer audio between sentences and pass the previous chunk's text in
prompt. - For speaker labels send
gpt-4o-transcribe-diarizewithresponse_format=diarized_jsonandchunking_strategy=autofor audio over 30 seconds. Plan for its shutdown on 26 February 2027. - Word timestamps,
srt,vttand/v1/audio/translationsneedwhisper-1, which cannot stream and shuts down on the same date. - On 429 or 503 wait at least
Retry-Afterwhen present, then back off with jitter. Do not retrycredit_balance_exhaustedor spend-limit errors.
Questions
Which is better for AI agents, Deepgram Speech-to-Text (Nova-3, Flux) or OpenAI Speech to Text?
OpenAI Speech to Text scores 72.4 (BB) on agent readiness against Deepgram Speech-to-Text (Nova-3, Flux)'s 70.3 (BB), and leads in 3 of 7 scored categories. Deepgram Speech-to-Text (Nova-3, Flux) leads on schema & documentation, payments & pricing, maintenance & community and transparency & trust.
Do Deepgram Speech-to-Text (Nova-3, Flux) and OpenAI Speech to Text need an API key?
Both need an API key.
Can an agent call Deepgram Speech-to-Text (Nova-3, Flux) and OpenAI Speech to Text without installing anything?
Yes. Deepgram Speech-to-Text (Nova-3, Flux) has a hosted endpoint at https://api.deepgram.com/v1 and OpenAI Speech to Text at https://api.openai.com/v1.
Other comparisons with Deepgram Speech-to-Text (Nova-3, Flux) or OpenAI Speech to Text
- Amazon Transcribe vs Deepgram Speech-to-Text (Nova-3, Flux)
- Amazon Transcribe vs OpenAI Speech to Text
- AssemblyAI Speech-to-Text (Universal) vs Deepgram Speech-to-Text (Nova-3, Flux)
- AssemblyAI Speech-to-Text (Universal) vs OpenAI Speech to Text
- Azure AI Speech speech-to-text vs Deepgram Speech-to-Text (Nova-3, Flux)
- Azure AI Speech speech-to-text vs OpenAI Speech to Text
- Cartesia Ink vs Deepgram Speech-to-Text (Nova-3, Flux)
- Cartesia Ink vs OpenAI Speech to Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs ElevenLabs Scribe Speech to Text API
- Deepgram Speech-to-Text (Nova-3, Flux) vs Gladia Speech-to-Text API + MCP
- Deepgram Speech-to-Text (Nova-3, Flux) vs Google Cloud Speech-to-Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs Groq Speech-to-Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs Mistral Voxtral Transcribe
- Deepgram Speech-to-Text (Nova-3, Flux) vs Rev AI Speech-to-Text API
- Deepgram Speech-to-Text (Nova-3, Flux) vs Soniox Speech-to-Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs Speechmatics Speech-to-Text
- ElevenLabs Scribe Speech to Text API vs OpenAI Speech to Text
- Gladia Speech-to-Text API + MCP vs OpenAI Speech to Text
- Google Cloud Speech-to-Text vs OpenAI Speech to Text
- Groq Speech-to-Text vs OpenAI Speech to Text
- Mistral Voxtral Transcribe vs OpenAI Speech to Text
- OpenAI Speech to Text vs Rev AI Speech-to-Text API
- OpenAI Speech to Text vs Soniox Speech-to-Text
- OpenAI Speech to Text vs Speechmatics Speech-to-Text
Machine-readable
- This page as Markdown
/compare/deepgram-stt-vs-openai-speech-to-text.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/deepgram-stt.json·/api/v1/tools/openai-speech-to-text.json - From a terminal
anchor compare deepgram-stt openai-speech-to-text(the CLI) - Over MCP
compare_tools {"a": "deepgram-stt", "b": "openai-speech-to-text"}at/mcp, no key