# HeyGen Voice (slim) > HeyGen Voice is HeyGen's in-house text-to-speech model, heygen-voice-1, which speaks text in a voice cloned from one recording (instant) or from 20 minutes or more (professional), through HeyGen's v3 API. - Full: https://www.anchorterminal.com/tools/heygen-voice.md (~6,450 tokens) · this version ~1,780 tokens · JSON https://www.anchorterminal.com/tools/heygen-voice.json · canonical https://www.anchorterminal.com/tools/heygen-voice - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-11 **BB · 71.3/100 · rank #131 of 961 · #5 in Text-to-speech · agent-ready · confidence medium** Assessment: HeyGen Voice has a typed OpenAPI contract, streaming with character timestamps, and a published price of $15 per million characters for instant voices, with failed requests not charged. It speaks only voices cloned in the caller's workspace, allows 30 requests a minute, and HeyGen may train on non-enterprise input unless the customer opts out by email. ## Facts - Kind: Model API · vendor: HeyGen · category: Text-to-speech · legal entity: HeyGen Technology, Inc. · provenance 85/100 - Endpoint: `https://api.heygen.com` (HTTP) - Auth: API key · pricing: Pay per use · x402: no · licence: Closed service under HeyGen's Terms of Service (HeyGen Technology, Inc., last updated 23 July 2026). - Probe metrics: not measured yet (probes haven't run) - Surface graded: HeyGen Voice through HeyGen's v3 API, POST /v3/models/audio/tts and /tts/stream with the voice endpoints under /v3/models/audio/voices. Read from the developer docs and OpenAPI file on 10 October 2026. No authenticated call was made. - Models: heygen-voice-1 (default, instant and professional voices) and heygen-voice-1-turbo (instant voices only, lower latency, per the OpenAPI file). - Voice modes: Instant voices use one recording (the first 3 minutes), are usually ready within seconds and cannot be retrained. Professional voices use 1 to 10 recordings of one speaker, 20 minutes or more in total, are trained and retrainable, and need one purchased slot each with five trainings a month. - Output: Completed speech returns a URL to a 44.1 kHz mono PCM16 WAV and its duration. Streaming returns Server-Sent Events of base64 WAV parts, ending with [DONE], with optional character alignment events. - Controls: Instant voices take expressiveness_boost from 0.0 to 1.0. Professional voices take seed, speed 0.75 to 1.25, pitch_shift -4 to +4 semitones, pitch_variance 0.5 to 2.0, and pauses up to 5 seconds. No other SSML. - Languages: 39 base languages for cloning and speech, including ar, de, en, es, fr, hi, ja, ko, pt, ru and zh. Regional tags normalise to the base language. Detected from the text when omitted. - Limits: 1 to 5,000 characters a request. 30 requests a minute per workspace member on each speech endpoint, on every plan. - Errors: 400 invalid_parameter or voice_expired, 402 insufficient_credit or usage_limit_reached, 403 forbidden, 404 voice_not_found, 409 voice_not_ready or voice_training_failed, 429 rate_limit_exceeded with Retry-After, 502, 503 and 504 safe to retry. - Preview status: Instant voice creation is free during a preview, per the model page. The professional mode carries no preview label. - Data handling: The privacy policy of 11 August 2026 says input may train HeyGen's models, with an opt-out by email, and backups are kept 60 days after deletion. The security page says enterprise data is excluded from training by default and states SOC 2 Type II. - Prices: Instant-voice speech $15 per 1M characters; Professional-voice speech $0.60 per minute of audio - Scores: Reliability 75, Performance pending, Schema & documentation 94, Agent ergonomics 83, Security & auth 69, Payments & pricing 30, Task success pending, Maintenance & community 65, Transparency & trust 69 · total over the 7 assessed categories - Why: Reliability, Status page at status.heygen.com with components for www, app and api.heygen.com and dated incident history (20). · Schema & documentation, OpenAPI 3.1 file at developers.heygen.com/openapi/external-api.json covering create, list, get and delete voice, completed speech and stream… · Agent ergonomics, API reading of the checklist. · Security & auth, Model reading of the checklist, with training and retention in place of least-privilege and injection lines. · Payments & pricing, No x402, MPP or L402 in the llms.txt, the OpenAPI file or the voice pages (0). · Maintenance & community, Read as a model. · Transparency & trust, Closed service with terms dated 23 July 2026. Paid plans own their output, and free-plan output is limited to non-commercial use (15). - Sources: 14, open questions: 6, both in the full twin - Capabilities: speech.tts, speech.streaming, speech.languages, voice.clone - JSON: https://www.anchorterminal.com/api/v1/tools/heygen-voice.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/heygen-voice.svg` or a link to https://www.anchorterminal.com/tools/heygen-voice from a page on heygen.com or one of its subdomains, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Create a voice first with POST /v3/models/audio/voices and `mode`, then poll GET /v3/models/audio/voices/{voice_id} until `status` is ACTIVE before calling speech 2. Send `expressiveness_boost` only for instant voices and `seed`, `speed`, `pitch_shift`, `pitch_variance` or `` tags only for professional voices. Mixing them returns 400 invalid_parameter 3. On the stream, decode and play each base64 WAV part in `part_index` order. Do not concatenate the bytes 4. Send an Idempotency-Key when creating a voice, and retry 502, 503 and 504 on speech, which the OpenAPI file marks as safe 5. Use POST /v3/voices/speech for stock or designed voices. It is a different engine and price ## Connect ```bash curl -X POST "https://api.heygen.com/v3/models/audio/tts" \ -H "X-Api-Key: $HEYGEN_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"heygen-voice-1","voice_id":"","text":"Hello.","language":"en","expressiveness_boost":0.8}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Fish Audio TTS API | C | 60.7 | speech.tts, speech.streaming, speech.languages, voice.clone | https://www.anchorterminal.com/tools/fish-audio-tts.min.md | | Amazon Polly | BB | 75.6 | speech.tts, speech.streaming, speech.languages | https://www.anchorterminal.com/tools/amazon-polly.min.md | | ElevenLabs Text to Speech API + MCP | BB | 75.2 | speech.tts, speech.streaming, speech.languages | https://www.anchorterminal.com/tools/elevenlabs-tts.min.md | | Deepgram Text-to-Speech (Aura-2, Flux TTS) | BB | 72.7 | speech.tts, speech.streaming, speech.languages | https://www.anchorterminal.com/tools/deepgram-tts.min.md | | Azure AI Speech text-to-speech | BB | 71.4 | speech.tts, speech.streaming, speech.languages | https://www.anchorterminal.com/tools/azure-text-to-speech.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)