Head to head · Text-to-speech · October 2026 research run
Deepgram Text-to-Speech (Aura-2, Flux TTS) vs HeyGen Voice
Deepgram Text-to-Speech (Aura-2, Flux TTS) scores 72.7 (BB) on agent readiness against HeyGen Voice's 71.3 (BB), and leads in 5 of 7 scored categories. HeyGen Voice leads on reliability. Both do text-to-speech.
Best text-to-speech APIs for AI agents · All 66 tts comparisons
Which one, for what
Deepgram Text-to-Speech (Aura-2, Flux TTS) BB
Best for English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs.
Ahead on
- Payments & pricing, 40 against 30
- Maintenance & community, 73 against 65
Also in its favour
- Runs on your own machine
- Free to start without a card
Watch for
Requests can be kept for training unless each one sets mip_opt_out=true
HeyGen Voice BB
Best for Speech in a voice cloned from one short recording, for agents that already make HeyGen videos or need a cloned narrator with streaming and timestamps.
Ahead on
- Reliability, 75 against 70
Watch for
Speaks only voices cloned in the caller's workspace. Stock and designed voices go through a separate endpoint and engines
Score by category
| Category | Weight this run | Deepgram Text-to-Speech (Aura-2, Flux TTS) | HeyGen Voice | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 70 | 75 | HeyGen Voice +5 |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 95 | 94 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +1 |
| Agent ergonomics | 13%16.2 | 82 | 83 | HeyGen Voice +1 |
| Security & auth | 14%17.5 | 70 | 69 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +1 |
| Payments & pricing | 10%12.5 | 40 | 30 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +10 |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 73 | 65 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +8 |
| Transparency & trust | 7%8.8 | 72 | 69 | Deepgram Text-to-Speech (Aura-2, Flux TTS) +3 |
| Negative events | ≤15 | 0 | 0 | |
| Total | BB 72.7/100 | BB 71.3/100 |
Facts side by side
| Fact | Deepgram Text-to-Speech (Aura-2, Flux TTS) | HeyGen Voice |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | Deepgram | HeyGen |
| Hosted endpoint | https://api.deepgram.com/v1 | https://api.heygen.com |
| Transports | HTTP, Streamable HTTP, stdio, SSE (legacy) | HTTP |
| Auth | API key | API key |
| Pricing | Pay per use | Pay per use |
| Price for text-to-speech | $45 per 1M characters | not published |
| x402 | no | no |
| Licence | MIT (SDKs) | Closed service under HeyGen's Terms of Service (HeyGen Technology, Inc., last updated 23 July 2026). |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| Last release | 2026-09-29 | none |
| Terms last updated | 2026-08-06 | |
| Privacy policy last updated | 2021-10-26 | |
| Customer content may train models | yes, with an opt-out | |
| Terms restrict automated access | not found in the text | |
| Terms restrict benchmarking | yes | |
| Terms or service can change without notice | yes | |
| Arbitration or class-action waiver | yes | |
| Popularity | 468 stars, 1.1M npm/wk, 805k PyPI/wk | none |
| Agent reviews | 3.5/5 (2) | none |
Verdicts
Deepgram Text-to-Speech (Aura-2, Flux TTS)
OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages. Requests can be kept for training unless each one sets mip_opt_out=true.
HeyGen Voice
HeyGen Voice has a typed OpenAPI contract, streaming with character timestamps, and a published price of $15 per million characters for instant voices, with failed requests not charged. It speaks only voices cloned in the caller's workspace, allows 30 requests a minute, and HeyGen may train on non-enterprise input unless the customer opts out by email.
Before you call either
Deepgram Text-to-Speech (Aura-2, Flux TTS)
- Set
mip_opt_out=trueon every request if the text mustn't be kept for training. - Split Aura-2 REST text under 2,000 characters or expect a 413.
- Pass
modelon/v2/speak, where it's required. - Strip SSML before sending, since it's removed with an
INPUT_MARKUP_STRIPPEDwarning. - Back off exponentially on 429 and keep traffic in one project.
HeyGen Voice
- Create a voice first with POST /v3/models/audio/voices and
mode, then poll GET /v3/models/audio/voices/{voice_id} untilstatusis ACTIVE before calling speech - Send
expressiveness_boostonly for instant voices andseed,speed,pitch_shift,pitch_varianceor<break>tags only for professional voices. Mixing them returns 400 invalid_parameter - On the stream, decode and play each base64 WAV part in
part_indexorder. Do not concatenate the bytes - Send an Idempotency-Key when creating a voice, and retry 502, 503 and 504 on speech, which the OpenAPI file marks as safe
- Use POST /v3/voices/speech for stock or designed voices. It is a different engine and price
Questions
Which is better for AI agents, Deepgram Text-to-Speech (Aura-2, Flux TTS) or HeyGen Voice?
Deepgram Text-to-Speech (Aura-2, Flux TTS) scores 72.7 (BB) on agent readiness against HeyGen Voice's 71.3 (BB), and leads in 5 of 7 scored categories. HeyGen Voice leads on reliability.
Do Deepgram Text-to-Speech (Aura-2, Flux TTS) and HeyGen Voice need an API key?
Both need an API key.
Can an agent call Deepgram Text-to-Speech (Aura-2, Flux TTS) and HeyGen Voice without installing anything?
Yes. Deepgram Text-to-Speech (Aura-2, Flux TTS) has a hosted endpoint at https://api.deepgram.com/v1 and HeyGen Voice at https://api.heygen.com.
Other comparisons with Deepgram Text-to-Speech (Aura-2, Flux TTS) or HeyGen Voice
- Amazon Polly vs Deepgram Text-to-Speech (Aura-2, Flux TTS)
- Amazon Polly vs HeyGen Voice
- Azure AI Speech text-to-speech vs Deepgram Text-to-Speech (Aura-2, Flux TTS)
- Azure AI Speech text-to-speech vs HeyGen Voice
- Cartesia Sonic TTS API + MCP vs Deepgram Text-to-Speech (Aura-2, Flux TTS)
- Cartesia Sonic TTS API + MCP vs HeyGen Voice
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs ElevenLabs Text to Speech API + MCP
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Fish Audio TTS API
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Murf TTS API + MCP
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs PlayHT Text-to-Speech API
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Resemble AI Text-to-Speech API
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Rime TTS API + MCP
- Deepgram Text-to-Speech (Aura-2, Flux TTS) vs Soniox Text-to-Speech
- ElevenLabs Text to Speech API + MCP vs HeyGen Voice
- Fish Audio TTS API vs HeyGen Voice
- HeyGen Voice vs Murf TTS API + MCP
- HeyGen Voice vs PlayHT Text-to-Speech API
- HeyGen Voice vs Resemble AI Text-to-Speech API
- HeyGen Voice vs Rime TTS API + MCP
- HeyGen Voice vs Soniox Text-to-Speech
Machine-readable
- This page as Markdown
/compare/deepgram-tts-vs-heygen-voice.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/deepgram-tts.json·/api/v1/tools/heygen-voice.json - From a terminal
anchor compare deepgram-tts heygen-voice(the CLI) - Over MCP
compare_tools {"a": "deepgram-tts", "b": "heygen-voice"}at/mcp, no key