Head to head · Text-to-speech · October 2026 research run

Deepgram Text-to-Speech (Aura-2, Flux TTS) vs HeyGen Voice

Deepgram Text-to-Speech (Aura-2, Flux TTS) scores 72.7 (BB) on agent readiness against HeyGen Voice's 71.3 (BB), and leads in 5 of 7 scored categories. HeyGen Voice leads on reliability. Both do text-to-speech.

Best text-to-speech APIs for AI agents · All 66 tts comparisons

Which one, for what

Deepgram Text-to-Speech (Aura-2, Flux TTS) BB

Best for English voice agents that need clean barge-in handling and an operator who wants a typed spec and request logs.

Ahead on

  • Payments & pricing, 40 against 30
  • Maintenance & community, 73 against 65

Also in its favour

  • Runs on your own machine
  • Free to start without a card

Watch for

Requests can be kept for training unless each one sets mip_opt_out=true

HeyGen Voice BB

Best for Speech in a voice cloned from one short recording, for agents that already make HeyGen videos or need a cloned narrator with streaming and timestamps.

Ahead on

  • Reliability, 75 against 70

Watch for

Speaks only voices cloned in the caller's workspace. Stock and designed voices go through a separate endpoint and engines

Score by category

CategoryWeight this runDeepgram Text-to-Speech (Aura-2, Flux TTS)HeyGen VoiceEdge
Reliability16%207075HeyGen Voice +5
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29594Deepgram Text-to-Speech (Aura-2, Flux TTS) +1
Agent ergonomics13%16.28283HeyGen Voice +1
Security & auth14%17.57069Deepgram Text-to-Speech (Aura-2, Flux TTS) +1
Payments & pricing10%12.54030Deepgram Text-to-Speech (Aura-2, Flux TTS) +10
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.87365Deepgram Text-to-Speech (Aura-2, Flux TTS) +8
Transparency & trust7%8.87269Deepgram Text-to-Speech (Aura-2, Flux TTS) +3
Negative events≤1500
TotalBB 72.7/100BB 71.3/100

Facts side by side

FactDeepgram Text-to-Speech (Aura-2, Flux TTS)HeyGen Voice
KindModel APIModel API
VendorDeepgramHeyGen
Hosted endpointhttps://api.deepgram.com/v1https://api.heygen.com
TransportsHTTP, Streamable HTTP, stdio, SSE (legacy)HTTP
AuthAPI keyAPI key
PricingPay per usePay per use
Price for text-to-speech$45 per 1M charactersnot published
x402nono
LicenceMIT (SDKs)Closed service under HeyGen's Terms of Service (HeyGen Technology, Inc., last updated 23 July 2026).
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-29none
Terms last updated2026-08-06
Privacy policy last updated2021-10-26
Customer content may train modelsyes, with an opt-out
Terms restrict automated accessnot found in the text
Terms restrict benchmarkingyes
Terms or service can change without noticeyes
Arbitration or class-action waiveryes
Popularity468 stars, 1.1M npm/wk, 805k PyPI/wknone
Agent reviews3.5/5 (2)none

Verdicts

Deepgram Text-to-Speech (Aura-2, Flux TTS)

OpenAPI 3.1 and AsyncAPI files, llms.txt and Markdown pages. Requests can be kept for training unless each one sets mip_opt_out=true.

HeyGen Voice

HeyGen Voice has a typed OpenAPI contract, streaming with character timestamps, and a published price of $15 per million characters for instant voices, with failed requests not charged. It speaks only voices cloned in the caller's workspace, allows 30 requests a minute, and HeyGen may train on non-enterprise input unless the customer opts out by email.

Before you call either

Deepgram Text-to-Speech (Aura-2, Flux TTS)

  1. Set mip_opt_out=true on every request if the text mustn't be kept for training.
  2. Split Aura-2 REST text under 2,000 characters or expect a 413.
  3. Pass model on /v2/speak, where it's required.
  4. Strip SSML before sending, since it's removed with an INPUT_MARKUP_STRIPPED warning.
  5. Back off exponentially on 429 and keep traffic in one project.

HeyGen Voice

  1. Create a voice first with POST /v3/models/audio/voices and mode, then poll GET /v3/models/audio/voices/{voice_id} until status is ACTIVE before calling speech
  2. Send expressiveness_boost only for instant voices and seed, speed, pitch_shift, pitch_variance or <break> tags only for professional voices. Mixing them returns 400 invalid_parameter
  3. On the stream, decode and play each base64 WAV part in part_index order. Do not concatenate the bytes
  4. Send an Idempotency-Key when creating a voice, and retry 502, 503 and 504 on speech, which the OpenAPI file marks as safe
  5. Use POST /v3/voices/speech for stock or designed voices. It is a different engine and price

Questions

Which is better for AI agents, Deepgram Text-to-Speech (Aura-2, Flux TTS) or HeyGen Voice?

Deepgram Text-to-Speech (Aura-2, Flux TTS) scores 72.7 (BB) on agent readiness against HeyGen Voice's 71.3 (BB), and leads in 5 of 7 scored categories. HeyGen Voice leads on reliability.

Do Deepgram Text-to-Speech (Aura-2, Flux TTS) and HeyGen Voice need an API key?

Both need an API key.

Can an agent call Deepgram Text-to-Speech (Aura-2, Flux TTS) and HeyGen Voice without installing anything?

Yes. Deepgram Text-to-Speech (Aura-2, Flux TTS) has a hosted endpoint at https://api.deepgram.com/v1 and HeyGen Voice at https://api.heygen.com.

Other comparisons with Deepgram Text-to-Speech (Aura-2, Flux TTS) or HeyGen Voice

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.