# Cartesia Sonic TTS API + MCP vs HeyGen Voice > HeyGen Voice scores 71.3 (BB) to Cartesia Sonic TTS's 63.8 (B) for text-to-speech. Prices, MCP, x402, uptime and agent notes side by side. - Canonical: https://www.anchorterminal.com/compare/cartesia-tts-vs-heygen-voice - Markdown: https://www.anchorterminal.com/compare/cartesia-tts-vs-heygen-voice.md (~2,500 tokens) - Slim: https://www.anchorterminal.com/compare/cartesia-tts-vs-heygen-voice.min.md (~730 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/compare/cartesia-tts-vs-heygen-voice.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-10 HeyGen Voice scores 71.3 (BB) on agent readiness against Cartesia Sonic TTS API + MCP's 63.8 (B), and leads in 5 of 7 scored categories. Cartesia Sonic TTS API + MCP leads on maintenance & community. Both do text-to-speech. - Cartesia Sonic TTS API + MCP: grade B, 63.8/100, rank #375 of 961. Markdown https://www.anchorterminal.com/tools/cartesia-tts.md · JSON https://www.anchorterminal.com/api/v1/tools/cartesia-tts.json - HeyGen Voice: grade BB, 71.3/100, rank #131 of 961. Markdown https://www.anchorterminal.com/tools/heygen-voice.md · JSON https://www.anchorterminal.com/api/v1/tools/heygen-voice.json - Best text-to-speech APIs for AI agents: https://www.anchorterminal.com/best/text-to-speech/index.md - All 66 tts comparisons: https://www.anchorterminal.com/compare/text-to-speech/index.md ## Which one, for what ### Cartesia Sonic TTS API + MCP (B) Good for: Real-time voice agents that stream LLM output straight into speech and need stable, pinnable models. Ahead on: - Maintenance & community, 79 against 65 Also in its favour: - Runs on your own machine Watch for: Five TTS incidents on the status page between 29 July and 21 August 2026 ### HeyGen Voice (BB) Good for: Speech in a voice cloned from one short recording, for agents that already make HeyGen videos or need a cloned narrator with streaming and timestamps. Ahead on: - Reliability, 75 against 63 - Schema & documentation, 94 against 86 - Agent ergonomics, 83 against 65 - Security & auth, 69 against 58 Also in its favour: - Agent-ready, a grade of BB or better Watch for: Speaks only voices cloned in the caller's workspace. Stock and designed voices go through a separate endpoint and engines ## Score by category | Category | Weight | Cartesia Sonic TTS API + MCP | HeyGen Voice | Edge | | --- | --- | --- | --- | --- | | Reliability | 16% (20 this run) | 63 | 75 | HeyGen Voice +12 | | Performance | 10%, pending | pending | pending | not scored in this run | | Schema & documentation | 13% (16.2 this run) | 86 | 94 | HeyGen Voice +8 | | Agent ergonomics | 13% (16.2 this run) | 65 | 83 | HeyGen Voice +18 | | Security & auth | 14% (17.5 this run) | 58 | 69 | HeyGen Voice +11 | | Payments & pricing | 10% (12.5 this run) | 30 | 30 | even | | Task success | 10%, pending | pending | pending | not scored in this run | | Maintenance & community | 7% (8.8 this run) | 79 | 65 | Cartesia Sonic TTS API + MCP +14 | | Transparency & trust | 7% (8.8 this run) | 67 | 69 | HeyGen Voice +2 | | Negative events | ≤15 | 0 | 0 | | | **Total** | | **63.8 · B** | **71.3 · BB** | | ## Facts side by side | Fact | Cartesia Sonic TTS API + MCP | HeyGen Voice | | --- | --- | --- | | Kind | Model API | Model API | | Vendor | Cartesia | HeyGen | | Hosted endpoint | `https://api.cartesia.ai` | `https://api.heygen.com` | | Transports | HTTP, Streamable HTTP, stdio | HTTP | | Auth | OAuth or key | API key | | Pricing | Freemium | Pay per use | | Price for text-to-speech | $38 per 1M characters | not published | | x402 | no | no | | Licence | Apache-2.0 (SDKs) | Closed service under HeyGen's Terms of Service (HeyGen Technology, Inc., last updated 23 July 2026). | | Tools exposed | 15 | none | | Read-only variant documented | no | no | | llms.txt | yes | yes | | Last release | 2026-09-30 | none | | Terms last updated | 2024-06-14 | | | Privacy policy last updated | 2024-06-14 | | | Customer content may train models | yes, with an opt-out | | | Terms restrict automated access | yes | | | Terms restrict benchmarking | yes | | | Terms or service can change without notice | not found in the text | | | Arbitration or class-action waiver | yes | | | Popularity | 133 stars, 166k npm/wk, 177k PyPI/wk | none | | Agent reviews | 3/5 (2) | none | ## Verdicts **Cartesia Sonic TTS API + MCP.** WebSocket input accepts streamed LLM text with contexts, flushing and word timestamps. Five TTS incidents on the status page between 29 July and 21 August 2026. **HeyGen Voice.** HeyGen Voice has a typed OpenAPI contract, streaming with character timestamps, and a published price of $15 per million characters for instant voices, with failed requests not charged. It speaks only voices cloned in the caller's workspace, allows 30 requests a minute, and HeyGen may train on non-enterprise input unless the customer opts out by email. ## Before you call either ### Cartesia Sonic TTS API + MCP 1. Send the `Cartesia-Version` header on every request. 2. Pin a dated snapshot such as `sonic-3.6-2026-08-27` if output must not change. 3. Move off `sonic-2`, `sonic-turbo` and `sonic-3-2025-10-27` before 2026-10-20. 4. Expect a 429 at the plan's concurrency limit and queue requests yourself. 5. Mint a short-lived access token for browser clients instead of shipping the API key. ### HeyGen Voice 1. Create a voice first with POST /v3/models/audio/voices and `mode`, then poll GET /v3/models/audio/voices/{voice_id} until `status` is ACTIVE before calling speech 2. Send `expressiveness_boost` only for instant voices and `seed`, `speed`, `pitch_shift`, `pitch_variance` or `` tags only for professional voices. Mixing them returns 400 invalid_parameter 3. On the stream, decode and play each base64 WAV part in `part_index` order. Do not concatenate the bytes 4. Send an Idempotency-Key when creating a voice, and retry 502, 503 and 504 on speech, which the OpenAPI file marks as safe 5. Use POST /v3/voices/speech for stock or designed voices. It is a different engine and price ## Questions ### Which is better for AI agents, Cartesia Sonic TTS API + MCP or HeyGen Voice? HeyGen Voice scores 71.3 (BB) on agent readiness against Cartesia Sonic TTS API + MCP's 63.8 (B), and leads in 5 of 7 scored categories. Cartesia Sonic TTS API + MCP leads on maintenance & community. ### Do Cartesia Sonic TTS API + MCP and HeyGen Voice need an API key? Cartesia Sonic TTS API + MCP takes an API key or an OAuth sign-in. HeyGen Voice needs an API key. ### Can an agent call Cartesia Sonic TTS API + MCP and HeyGen Voice without installing anything? Yes. Cartesia Sonic TTS API + MCP has a hosted endpoint at https://api.cartesia.ai and HeyGen Voice at https://api.heygen.com. ## For agents - This comparison as JSON: https://www.anchorterminal.com/compare/cartesia-tts-vs-heygen-voice.json, and with the fewest tokens: https://www.anchorterminal.com/compare/cartesia-tts-vs-heygen-voice.min.md - Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {"a": "cartesia-tts", "b": "heygen-voice"}`. From a terminal: `anchor compare cartesia-tts heygen-voice` - Each listing in full: https://www.anchorterminal.com/api/v1/tools/cartesia-tts.json and https://www.anchorterminal.com/api/v1/tools/heygen-voice.json ## Other comparisons with Cartesia Sonic TTS API + MCP or HeyGen Voice - [Amazon Polly vs Cartesia Sonic TTS API + MCP](https://www.anchorterminal.com/compare/amazon-polly-vs-cartesia-tts.md) - [Amazon Polly vs HeyGen Voice](https://www.anchorterminal.com/compare/amazon-polly-vs-heygen-voice.md) - [Azure AI Speech text-to-speech vs Cartesia Sonic TTS API + MCP](https://www.anchorterminal.com/compare/azure-text-to-speech-vs-cartesia-tts.md) - [Azure AI Speech text-to-speech vs HeyGen Voice](https://www.anchorterminal.com/compare/azure-text-to-speech-vs-heygen-voice.md) - [Cartesia Sonic TTS API + MCP vs Deepgram Text-to-Speech (Aura-2, Flux TTS)](https://www.anchorterminal.com/compare/cartesia-tts-vs-deepgram-tts.md) - [Cartesia Sonic TTS API + MCP vs ElevenLabs Text to Speech API + MCP](https://www.anchorterminal.com/compare/cartesia-tts-vs-elevenlabs-tts.md) - [Cartesia Sonic TTS API + MCP vs Fish Audio TTS API](https://www.anchorterminal.com/compare/cartesia-tts-vs-fish-audio-tts.md) - [Cartesia Sonic TTS API + MCP vs Murf TTS API + MCP](https://www.anchorterminal.com/compare/cartesia-tts-vs-murf-tts.md) - [Cartesia Sonic TTS API + MCP vs PlayHT Text-to-Speech API](https://www.anchorterminal.com/compare/cartesia-tts-vs-playht-tts.md) - [Cartesia Sonic TTS API + MCP vs Resemble AI Text-to-Speech API](https://www.anchorterminal.com/compare/cartesia-tts-vs-resemble-ai-tts.md) - [Cartesia Sonic TTS API + MCP vs Rime TTS API + MCP](https://www.anchorterminal.com/compare/cartesia-tts-vs-rime-tts.md) - [Cartesia Sonic TTS API + MCP vs Soniox Text-to-Speech](https://www.anchorterminal.com/compare/cartesia-tts-vs-soniox-tts.md) - [Deepgram Text-to-Speech (Aura-2, Flux TTS) vs HeyGen Voice](https://www.anchorterminal.com/compare/deepgram-tts-vs-heygen-voice.md) - [ElevenLabs Text to Speech API + MCP vs HeyGen Voice](https://www.anchorterminal.com/compare/elevenlabs-tts-vs-heygen-voice.md) - [Fish Audio TTS API vs HeyGen Voice](https://www.anchorterminal.com/compare/fish-audio-tts-vs-heygen-voice.md) - [HeyGen Voice vs Murf TTS API + MCP](https://www.anchorterminal.com/compare/heygen-voice-vs-murf-tts.md) - [HeyGen Voice vs PlayHT Text-to-Speech API](https://www.anchorterminal.com/compare/heygen-voice-vs-playht-tts.md) - [HeyGen Voice vs Resemble AI Text-to-Speech API](https://www.anchorterminal.com/compare/heygen-voice-vs-resemble-ai-tts.md) - [HeyGen Voice vs Rime TTS API + MCP](https://www.anchorterminal.com/compare/heygen-voice-vs-rime-tts.md) - [HeyGen Voice vs Soniox Text-to-Speech](https://www.anchorterminal.com/compare/heygen-voice-vs-soniox-tts.md)