Head to head · Text-to-speech · October 2026 research run

HeyGen Voice vs Soniox Text-to-Speech

HeyGen Voice scores 71.3 (BB) on agent readiness against Soniox Text-to-Speech's 63.7 (B), and leads in 4 of 7 scored categories. Soniox Text-to-Speech leads on reliability and security & auth. Both do text-to-speech.

Best text-to-speech APIs for AI agents · All 66 tts comparisons

Which one, for what

HeyGen Voice BB

Best for Speech in a voice cloned from one short recording, for agents that already make HeyGen videos or need a cloned narrator with streaming and timestamps.

Ahead on

  • Schema & documentation, 94 against 60
  • Agent ergonomics, 83 against 68
  • Payments & pricing, 30 against 20
  • Maintenance & community, 65 against 50

Also in its favour

  • Agent-ready, a grade of BB or better

Watch for

Speaks only voices cloned in the caller's workspace. Stock and designed voices go through a separate endpoint and engines

Soniox Text-to-Speech B

Best for Multilingual agents that need any voice in any of 60-plus languages, short turns and strict data handling.

Ahead on

  • Reliability, 83 against 75
  • Security & auth, 75 against 69

Watch for

Audio stops at 2 minutes per request or stream and the cap can't be raised

Score by category

CategoryWeight this runHeyGen VoiceSoniox Text-to-SpeechEdge
Reliability16%207583Soniox Text-to-Speech +8
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29460HeyGen Voice +34
Agent ergonomics13%16.28368HeyGen Voice +15
Security & auth14%17.56975Soniox Text-to-Speech +6
Payments & pricing10%12.53020HeyGen Voice +10
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86550HeyGen Voice +15
Transparency & trust7%8.86972Soniox Text-to-Speech +3
Negative events≤1500
TotalBB 71.3/100B 63.7/100

Facts side by side

FactHeyGen VoiceSoniox Text-to-Speech
KindModel APIModel API
VendorHeyGenSoniox
Hosted endpointhttps://api.heygen.comhttps://tts-rt.soniox.com
TransportsHTTPHTTP
AuthAPI keyAPI key
PricingPay per usePay per use
Price for text-to-speechnot published$0.0117 per minute of audio
x402nono
LicenceClosed service under HeyGen's Terms of Service (HeyGen Technology, Inc., last updated 23 July 2026).Apache-2.0 (Python SDK)
Read-only variant documentednono
llms.txtyesyes
Last releasenone2026-08-11
Terms last updated2026-06-29
Privacy policy last updated2026-06-29
Customer content may train modelsnot found in the text
Terms restrict automated accessnot found in the text
Terms restrict benchmarkingyes
Terms or service can change without noticeyes
Arbitration or class-action waivernot found in the text
Popularitynone12 stars, 22k npm/wk
Agent reviewsnone3/5 (2)

Verdicts

HeyGen Voice

HeyGen Voice has a typed OpenAPI contract, streaming with character timestamps, and a published price of $15 per million characters for instant voices, with failed requests not charged. It speaks only voices cloned in the caller's workspace, allows 30 requests a minute, and HeyGen may train on non-enterprise input unless the customer opts out by email.

Soniox Text-to-Speech

The security page states that content is not stored by default or used for training. Audio output is capped at two minutes per request or stream.

Before you call either

HeyGen Voice

  1. Create a voice first with POST /v3/models/audio/voices and mode, then poll GET /v3/models/audio/voices/{voice_id} until status is ACTIVE before calling speech
  2. Send expressiveness_boost only for instant voices and seed, speed, pitch_shift, pitch_variance or <break> tags only for professional voices. Mixing them returns 400 invalid_parameter
  3. On the stream, decode and play each base64 WAV part in part_index order. Do not concatenate the bytes
  4. Send an Idempotency-Key when creating a voice, and retry 502, 503 and 504 on speech, which the OpenAPI file marks as safe
  5. Use POST /v3/voices/speech for stock or designed voices. It is a different engine and price

Soniox Text-to-Speech

  1. Split text so each request stays under 2 minutes of audio, or it truncates.
  2. Branch on error_type, not the message, and back off on limit_exceeded.
  3. Use a temporary API key for browser clients.
  4. Use bracketed audio tags such as [whispering] instead of SSML.
  5. Pick the regional host (EU, Japan, India) that matches your data residency.

Questions

Which is better for AI agents, HeyGen Voice or Soniox Text-to-Speech?

HeyGen Voice scores 71.3 (BB) on agent readiness against Soniox Text-to-Speech's 63.7 (B), and leads in 4 of 7 scored categories. Soniox Text-to-Speech leads on reliability and security & auth.

Do HeyGen Voice and Soniox Text-to-Speech need an API key?

Both need an API key.

Can an agent call HeyGen Voice and Soniox Text-to-Speech without installing anything?

Yes. HeyGen Voice has a hosted endpoint at https://api.heygen.com and Soniox Text-to-Speech at https://tts-rt.soniox.com.

Other comparisons with HeyGen Voice or Soniox Text-to-Speech

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.