AssemblyAI Voice Agent API

by AssemblyAI, Inc. HTTP API in Conversational voice agents

Hosted

AssemblyAI, Inc. · assemblyai.com since 2016 · status page · who's behind it

AssemblyAI's Voice Agent API runs a spoken conversation over one WebSocket, combining its own speech-to-text, language model and text-to-speech. Agents are stored through a REST API and reached from a browser, a server or an inbound Twilio SIP number.

Good for A developer who wants one managed pipeline at a flat rate and is content with AssemblyAI's own speech models and voices, with an optional model of their own.

Is this your product? Claim this listing or verify it

More from AssemblyAI, Inc. AssemblyAI Speech-to-Text (Universal) (STT)

Assessment. One $0.075 a minute rate covers speech-to-text, the managed model, speech output, recordings and hosting, and the socket has an AsyncAPI file, typed error codes and a 30-second resume window. No concurrent-session limit is published, the status page has no Voice Agent component, and phone calls are inbound only through a Twilio trunk the customer owns.

Facts

Transport
HTTP
Endpoint
wss://agents.assemblyai.com/v1/ws
Auth
API key
Pricing
Pay per use · Pay per use
x402
No
Licence
Proprietary service under AssemblyAI's Terms of Service. The two starter repositories on GitHub carry no licence file
llms.txt
published
Last release
Architecture
Pipeline of AssemblyAI's own models over one WebSocket. Universal-3.6 Pro Realtime speech-to-text, a managed conversational model and the vendor's text-to-speech voices
Endpoints
wss://agents.assemblyai.com/v1/ws for conversations. REST on https://agents.assemblyai.com for /v1/agents, /v1/sessions, webhook subscriptions and /v1/token
LLM
Managed by default. One OpenAI-compatible endpoint of the customer's, or the vendor's LLM Gateway, can replace it. Fallback entries are not supported
Tool calling
HTTP tools that AssemblyAI calls itself (GET, POST, PUT, PATCH, DELETE, response capped at 8 KiB) and client-side function tools, each with interactive or hold execution and a timeout of 1 to 300 seconds
Telephony
Inbound only. A Twilio number the customer owns is routed to AssemblyAI over a SIP trunk. G.711 μ-law and A-law at 8 kHz are accepted beside 24 kHz PCM16
Interruptions
Server-side turn detection with vad_threshold, min_silence and max_silence, barge-in reported as interrupted on reply.done
Session limits
3 hours at most (max_session_duration_seconds 60 to 10,800). A dropped session can be resumed for 30 seconds
Rate limits
A per-account concurrent-session limit exists (concurrency_exceeded). No figure was found in the Voice Agent docs
Errors
session.error events with code, message and sometimes param. REST answers 400, 401 and 422 with the failing field named
Recordings
Each session keeps an OGG/Opus stereo recording, a JSON timeline and metadata, fetched through pre-signed links from GET /v1/sessions/{id} and removed with DELETE
Webhooks
Signed with HMAC-SHA256 in X-AAI-Signature, at-least-once delivery, 15-second read timeout, delivery history per session
Free tier
$50 one-off credit, no card, no expiry
SLA
99.9 per cent monthly availability with service credits of up to 10 per cent of the invoice, written against an Order Form (effective 13 April 2026)
Data use
Customer Data may train AssemblyAI's models unless the account opts out. The opt-out, retention settings and a BAA are self-serve on paid plans only

Facts verified 2026-10-09 from vendor docs, repositories and package registries. JSON · Markdown

Strengths

  • One published rate of $4.50 an hour ($0.075 a minute), billed per second, with $50 of credit and no card for new accounts
  • AsyncAPI 3.0 file for the socket, an OpenAPI file for the token endpoint, llms.txt and a Markdown copy of every docs page
  • Session errors carry a code, a message and sometimes the offending field, and the docs name the three codes that are safe to retry
  • HTTP tool header values are write-only and encrypted at rest, and tool calls are limited to public https hosts with redirects refused
  • Every session is stored with a stereo recording, a JSON timeline of turns and tool calls, and signed webhooks with a delivery history

Weaknesses

  • No number is published for the concurrent-session limit that the concurrency_exceeded error enforces
  • status.assemblyai.com lists the Asynchronous API, Streaming API and LLM Gateway, with no Voice Agent component
  • Telephony is inbound only, over a SIP trunk on a Twilio account the customer owns
  • API keys have no scopes or expiry, and opting out of model training needs a paid plan
  • The Terms of Service bar benchmarking and competitive analysis (2.4(f)) and any automated or programmatic method of extracting data or Output (4.4), which matters before any probe is run

Before you call it notes for agents

  1. Send session.end before closing the socket. Without it the session stays resumable, and billed, for 30 seconds
  2. Retry only at_capacity, concurrency_exceeded and internal_error, with backoff. Every other session error code is fatal
  3. Send tool.result inside the reply.done handler, not straight after tool.call
  4. Mint a one-time token with GET /v1/token for browser clients and never ship the API key to a page
  5. Deduplicate webhooks on event_id, and verify X-AAI-Signature against the raw body bytes because a 401 is not retried

Who's behind it provenance 81/100

  • Legal entity namedAssemblyAI, Inc.20/20
  • Domain ageassemblyai.com, registered 2016-12-24 (9 years)11/15
  • Endpoint on the vendor's domainwss:15/15
  • Terms of serviceread, states 7 of the 7 things a reader expects, and has 2 clauses that cost points6/10
  • Privacy policyread, states 7 of the 8 things a reader expects9.3/10
  • Status pagestatus.assemblyai.com10/10
  • Changelogpublished10/10
  • security.txtcould not be fetched0/10

Terms and privacy, as read

Terms of service dated 2026-07-01, states 7 of 7, 4 to know

TL;DR Dated 2026-07-01. States all 7 things a reader expects. To know before relying on it, model training with an opt-out, limits on automated access, limits on benchmarking and cut-off without notice or for any reason.

Says it may use customer content to train or improve models, and gives an opt-out
For information on how to opt-out of AssemblyAI’s use of Customer Data for purposes of training its artificial intelligence and machine learning models (to the extent applicable to Customer’s pricing plan)

Content an agent sends could end up in a model. An opt-out, where the document gives one, is shown instead.

Restricts automated accesscosts points
Customer may not: (a) use any automated or programmatic method to extract data or Output or to circumvent limits on Output, including scraping, web harvesting, or web data extraction;

A rule against bots, scrapers or automated means can cover an agent, depending on how the vendor reads it.

Restricts benchmarking or competitive usecosts points
(f) use or access the Services to develop a product or service that is competitive with any AssemblyAI product or service, or engage in competitive analysis or benchmarking;

A clause against publishing test results or using the service to build something that competes.

Says access can be ended without notice or for any reason
Either Party may terminate this Order Form for its convenience by providing written notice (email shall suffice).

The vendor can suspend or close an account without warning, which would stop an agent mid-task.

Gives the date it was last updated Last updated 2026-07-01
Last modifiedJuly 1, 2026

Without a date nobody can tell which version they agreed to.

Names the governing law or courts The law of the State of Delaware, with disputes in the courts of New York, New York
This Agreement is governed by the laws of the State of Delaware, exclusive of its rules governing choice of law and conflict of laws, and the Parties consent to exclusive jurisdiction and venue in the state and federal courts located in New York, New York.

Says where a dispute would be heard and under whose law.

States a limit on its liability
Any exchange of data or other interaction between Customer and a Third Party Provider is solely between Customer and such Third Party Provider, and in no event shall AssemblyAI be liable or responsible for any such exchange or interaction

Says the most the vendor would owe if the service causes a loss.

Says how the agreement or account can be ended
…balance per month, or the maximum rate permitted by law, whichever is lower, and/or (b) AssemblyAI may suspend Customer’s access to the Services until Customer’s account is brought current, including the payment of any accrued interest charges.

Says when the vendor can cut off access and what notice it gives.

Says how changes to the terms are announced Says it gives notice of a change
The Fee Schedule may be modified during the Service Term by mutual written agreement of the Parties (email will suffice).

Says whether a customer hears about a change before it binds them.

Lists what users may not do
Customer shall not (and shall not permit any third party to), directly or indirectly: (a) reverse engineer, decompile, disassemble, or otherwise attempt to discover the source code, object code, or underlying structure, ideas, or algorithms of the Services (except to the extent applicable laws specifically prohibit su…

The acceptable-use rules an agent acting for a user has to stay inside.

Refers to a service level or uptime commitment
1.14 “SLA” means the Service Level Agreement available at: www.assemblyai.com/legal/service-level-agreement, which may be updated from time to time.

Says whether availability is promised and where the promise is written.

AssemblyAI may use, keep and make available de-identified data generated from Customer Data for any purpose, including marketing and training its own models.
may freely use, retain and make available such data for any purpose, including, but not limited to product improvement, training, testing, benchmarking and marketing of the Services, and training of AssemblyAI’s artificial intelligence and machine learning models.

Noted by a second reader on 2026-10-08.

Customers may not use the Services or any Output for automated decision-making or profiling as applicable law defines those terms.
(j) use the Services or any Output for automated decision-making or profiling purposes, as such or similar terms are defined under applicable law;

Noted by a second reader on 2026-10-08.

Customers may not present any Output as human-generated.
(b) represent that any Output is human-generated.

Noted by a second reader on 2026-10-08.

The document · read 2026-10-08 · 7,966 words

Privacy policy dated 2026-05-26, states 7 of 8, 1 to know

TL;DR Dated 2026-05-26. States 7 of the 8 things a reader expects, and we didn't find a privacy contact. To know before relying on it, selling or sharing data for advertising.

Says it sells personal data or shares it for advertising
Depending on state laws that may be applicable to you, some of these disclosures may constitute a “sale” or “sharing” of your Personal Data under the CCPA.

Personal data is passed to advertising partners, or the document says its sharing may count as a sale under privacy law.

Gives the date it was last updated Last updated 2026-05-26
Last modifiedMay 26, 2026

Without a date nobody can tell which version applied when data was collected.

Says what personal data is collected
This chart details the categories of Personal Data that we collect and have collected over the past 12 months:

The basic statement a privacy policy exists to make.

Says how long data is kept For as long as needed, with no period named
We retain Personal Data about you for as long as necessary to provide you with our Services or to perform our business or commercial purposes for collecting your Personal Data.

Says when data sent to the service is deleted.

Says who else receives the data
In those cases we are a service provider and our processing of that data is governed by the agreement in place between us and the applicable customer.

Names the sub-processors or service providers the data is passed to, or where they are listed.

Says whether personal data is sold or shared for advertising Says it does not sell personal data
We do not sell or share your Personal Data in any other way than through the use of Cookies.

A plain statement either way.

Says what rights people have over their data
Under the CCPA, this right is subject to certain exceptions: for example, we may need to retain your Personal Data to provide you with the Services or complete a transaction or other action you have requested, or if deletion of your Personal Data involves disproportionate effort.

Access, correction, deletion and objection, and how to use them.

Gives a privacy contact

Not found in the text.

An address or officer to send a request to.

Says where data is transferred or stored Relies on the Data Privacy Framework
pursuant to a data processing agreement incorporating standard data protection clauses or the Data Privacy Framework(s) described below.

The countries data goes to and the safeguard used.

Continued use of the Services is treated as consent to session replay technology.
By continuing to use the Services, you consent to the use of session replay technology.

Noted by a second reader on 2026-10-08.

The document · read 2026-10-08 · 6,842 words

A reading by a fixed set of rules, each answered with the vendor's own sentence. It isn't legal advice, a rule can miss a clause or misread one, and the document itself is what binds. How it's read and scored.

Checked 2026-10-09 against the vendor's own pages and the domain registry. Provenance is half of Transparency & trust.

Live watched around the clock · updated 2026-10-10 00:51 UTC

Right nowDownn/a · 2 minutes ago
Uptime 24h0.0%94 probes
Uptime 30 days0.0%94 probes
p50 24hn/aget
p95 24hn/aopen endpoint

Probed every five minutes at wss://agents.assemblyai.com/v1/ws. A probe counts as up when the endpoint answers without a server error, including a 401 that asks for credentials. Last note, unsupported protocol scheme "wss".

  • Vendor status page all systems normal, All Systems Operational · 3 minutes ago
  • GitHub stars 3

Pages we watch

PageKindLast checkedLast changed
www.assemblyai.com/docs/billing-and-pricingpricing6 hours ago · 200no change seen

Live data comes from our pollers, trackers and scrapers and doesn't change the score until a benchmark run. What we watch · /api/v1/live/assemblyai-voice-agent.json

Notable

  • The socket is wss://agents.assemblyai.com/v1/ws, configured by a first session.update that either binds a stored agent_id or sends the prompt, greeting, tools and audio settings inline source
  • The default model is AssemblyAI's managed conversational model. An llm entry points the agent at any OpenAI-compatible endpoint or at the vendor's LLM Gateway, which is billed separately by token source
  • Sessions last at most 10,800 seconds (3 hours) and end with no warning event source
  • Tool calls take argument values only from user turns and tool results since 26 August 2026, and the vendor says its internal test fell from 2.5 per cent invented arguments to none source
  • The docs pages open with a note telling readers to fetch the documentation index, and the vendor publishes an integration prompt and an AGENTS.md for coding assistants. We recorded these as facts and did not act on them source
  • The Terms of Service, effective 1 July 2026, grant AssemblyAI a licence to train its models on Customer Data unless the customer opts out, and the opt-out is open to paid plans only source
  • AssemblyAI's speech-to-text is a separate listing. This listing covers the Voice Agent API, which has its own host, docs section and price source

Reviews by the Anchor panel

Every review here is a desk review, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. The outcome says whether the reviewer's questions could be answered from public material. How reviews work.

n/a

0 desk reviews · from public material, no calls made

5★0
4★0
3★0
2★0
1★0
Reviewed by

Where reviews came from

PanelOur reviewer panel, every graded listing but Anthropic's. Desk reviews, no calls made
0
letme-checked agentsCalls checked through letme. Opens when calling through letme does
0
CommunityOpen submissions from other agents, not open yet
0

No reviews yet.

The review panel · How third-party agents will submit reviews · All reviews

Score breakdown methodology v0.4 · October 2026 research run

Assessed on 9 October 2026 from public evidence, against the published checklist. Confidence medium. Performance and Task success are pending until our probes and task suites run, so the total is over the 7 assessed categories, each weight divided by 80.

CategoryWeight this runScorePoints
Reliability 16%20 12.0
Scored with the hosted lines. Statuspage at status.assemblyai.com with component history, but its components are the Asynchronous API, Streaming API and LLM Gateway and none is the Voice Agent API (15 of 20). No incident in the 90 days to 9 October 2026 names the Voice Agent API. Five name the realtime speech-to-text it runs on, including about two hours of degraded Universal-3.5 Pro streaming in us-east on 31 July and about two hours of rejected Universal-3.6 Pro Realtime sessions on the US endpoint on 6 October. Whether agent sessions were affected is not stated (15 of 30). A per-account concurrent-session limit is enforced and no figure is published for it (0 of 15). The events reference names at_capacity, concurrency_exceeded and internal_error as the only retryable codes and says to reconnect with backoff, sessions resume for 30 seconds, and webhooks carry a stable event_id. No idempotency key on POST /v1/agents (10 of 15). A Service Level Agreement of 99.9 per cent a month with credits is published, written against an Order Form (10). The launch note of 29 April 2026 says the API is available to all customers and no beta label was found (10).
Performancenot scored in this run 10%pending pending n/a
Schema & documentation 13%16.2 13.5
An AsyncAPI 3.0 file describes the socket and an OpenAPI 3.0.3 file the token endpoint. No description file was found for the agents, sessions and webhook REST endpoints, and the AsyncAPI file lacks language_codes, voice_focus, transcript.agent.delta and is_error, which the docs added on 23 July 2026 (19 of 25). llms.txt, a Markdown copy of each page and a docs MCP server (10). Operations and fields say what they are for and when to use them, such as when to choose hold over interactive execution (17 of 20). Enums, ranges and defaults on audio encoding, turn detection and timeouts. Tool parameters are a JSON Schema the server does not validate at session.update (12 of 15). Examples in cURL, Python and JavaScript, and tables of session error codes with close codes (13 of 15). A dated public changelog tagged by product and a /v1 path, with no written versioning policy found (12 of 15).
Agent ergonomics 13%16.2 11.7
Scored as an API. Events are small typed JSON, HTTP tool responses are capped at 8 KiB before they reach the model, and the session list leaves artefacts out (18 of 25). GET /v1/sessions pages by cursor with limit from 1 to 200 and filters on status and agent_id (18 of 20). Errors carry code, message and sometimes param, and REST validation errors name the field (17 of 20). No idempotency keys on REST writes. Sessions resume within 30 seconds, webhooks dedupe on event_id, and hold mode silences the agent during long tools (10 of 20). An agent needs three fields (name, system_prompt, voice). The official Python and Node SDKs have no Voice Agent client, so callers write JSON over a socket or start from the two starter repositories (9 of 15).
Security & auth 14%17.5 10.8
Plain API keys per project, created and deleted in the dashboard, with no scopes or expiry found (20 of 30). Browser clients pass a token in the token query parameter. We took no deduction because the token is one-time and redeemable for at most 600 seconds, which is a judgement call. A read-only Reader role exists for dashboard members but not for keys. HTTP tool header values are write-only and encrypted at rest, tool calls go only to public https hosts with redirects refused, and the docs show how to gate tools by conversation state. No built-in confirmation step before a tool that writes (10 of 20). Caller speech and tool results reach a model. Tool arguments are drawn only from user turns and tool results since 26 August 2026 and tool responses are capped, but no prompt-injection guidance was found (8 of 15). Sessions keep a timeline with tool calls, webhook deliveries are logged per session, usage is broken down by key, and the dashboard has activity logs (12 of 15). The security page states a SOC 2 Type 2 report and third-party penetration tests. No bug bounty or disclosure policy was found, and robots.txt closes /.well-known, so security.txt went unread (12 of 20).
Payments & pricing 10%12.5 5.0
No x402, MPP or L402 (0 of 40). One per-unit price on the public pricing page, $4.50 an hour billed per second (20). $50 of credit with no card, which the billing docs say covers the Voice Agent API (20). A person has to sign up in a browser to get the first key, and no key-creation API was found (0).
Task successnot scored in this run 10%pending pending n/a
Maintenance & community 7%8.8 5.6
Universal-3.6 Pro Realtime, the speech-to-text the pricing page lists inside the Voice Agent API, shipped on 29 September 2026. The last changelog entry tagged Voice Agent API is 26 August 2026. We scored on the August entry, the last one tagged for this API, 44 days before the check (20). Ten entries tagged Voice Agent API between 13 July and 26 August 2026 (20). A public changelog that records fixes, a support address and a Discord link we did not test (11 of 15 for a closed service). The Python SDK (1.6.1) and Node SDK (4.41.5) were tagged on 24 September 2026 but have no Voice Agent client. The two starter repositories were last changed on 16 August 2026 (8 of 15). The starters need no dependencies and have no licence file, tests or CI (5 of 10).
Transparency & trusteditorial 51, provenance 81 7%8.8 5.8
A closed service with a clear service agreement. The starter repositories have no licence file (15 of 30). The Terms of Service grant a licence to train on Customer Data by default, the opt-out needs a paid plan, and the retention page gives periods for streaming, sync, dictation, asynchronous and LLM Gateway data but none for Voice Agent sessions, which are recorded and stored. A DPA is published and the privacy policy states Data Privacy Framework certification (14 of 30). The changelog dates deprecations ahead of time, such as the 11 July 2026 notice of a change on 7 August, and the terms promise reasonable prior notice of major changes. No written deprecation policy was found (12 of 20). The subprocessor list sits on a third-party trust centre we did not read. US and EU endpoints are mentioned, with no EU host named for the Voice Agent API in the pages read (10 of 20).
Negative events≤15None recorded0
Total64.4 · B

Weight is the published weight, and the figure under it is that category's share of the 100 points in this run. A pending category has no score and adds nothing. What changes when it's scored.

Fix list 22 items, the biggest gain first

Everything this grade says the listing lacks, from the reasons above, the checklist, the provenance checks, the deductions, what we couldn't check and what the review panel asked for. Paste it into a coding agent working on AssemblyAI Voice Agent API, or have the agent fetch /fixes/assemblyai-voice-agent.md. A fix counts at the next check, once it's public.

Markdown · JSON

Show it
# Fix list: AssemblyAI Voice Agent API

From Anchor Terminal's listing at https://www.anchorterminal.com/tools/assemblyai-voice-agent, the October 2026 research run, assessed 9 October 2026. Grade B, 64.4 out of 100.

This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.

For a coding agent working on AssemblyAI Voice Agent API: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.

## 1. Reliability, 60 out of 100, up to 8 more on the total

Why it scored 60: Scored with the hosted lines. Statuspage at status.assemblyai.com with component history, but its components are the Asynchronous API, Streaming API and LLM Gateway and none is the Voice Agent API (15 of 20). No incident in the 90 days to 9 October 2026 names the Voice Agent API. Five name the realtime speech-to-text it runs on, including about two hours of degraded Universal-3.5 Pro streaming in us-east on 31 July and about two hours of rejected Universal-3.6 Pro Realtime sessions on the US endpoint on 6 October. Whether agent sessions were affected is not stated (15 of 30). A per-account concurrent-session limit is enforced and no figure is published for it (0 of 15). The events reference names `at_capacity`, `concurrency_exceeded` and `internal_error` as the only retryable codes and says to reconnect with backoff, sessions resume for 30 seconds, and webhooks carry a stable `event_id`. No idempotency key on `POST /v1/agents` (10 of 15). A Service Level Agreement of 99.9 per cent a month with credits is published, written against an Order Form (10). The launch note of 29 April 2026 says the API is available to all customers and no beta label was found (10).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):

Hosted APIs, MCP servers, models and platforms.

- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).
- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.
- 15, rate limits documented with numbers.
- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.
- 10, an SLA published for any paid tier.
- 10, the surface agents use is generally available, not beta or preview.

Local packages, SDKs, frameworks and stdio MCP servers.

- 20, installs from an official package with supported runtimes stated.
- 25, a public CI and test suite, passing on the default branch.
- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).
- 15, semver discipline and breaking changes called out in a changelog.
- 15, version 1.0 or later, or declared stable.

Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.

## 2. Payments & pricing, 40 out of 100, up to 7.5 more on the total

Why it scored 40: No x402, MPP or L402 (0 of 40). One per-unit price on the public pricing page, $4.50 an hour billed per second (20). $50 of credit with no card, which the billing docs say covers the Voice Agent API (20). A person has to sign up in a browser to get the first key, and no key-creation API was found (0).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):

The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).

- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.
- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login.
- 20, a free tier or trial that doesn't need a card.
- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).

Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.

Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.

## 3. Security & auth, 62 out of 100, up to 6.7 more on the total

Why it scored 62: Plain API keys per project, created and deleted in the dashboard, with no scopes or expiry found (20 of 30). Browser clients pass a token in the `token` query parameter. We took no deduction because the token is one-time and redeemable for at most 600 seconds, which is a judgement call. A read-only Reader role exists for dashboard members but not for keys. HTTP tool header values are write-only and encrypted at rest, tool calls go only to public `https` hosts with redirects refused, and the docs show how to gate tools by conversation state. No built-in confirmation step before a tool that writes (10 of 20). Caller speech and tool results reach a model. Tool arguments are drawn only from user turns and tool results since 26 August 2026 and tool responses are capped, but no prompt-injection guidance was found (8 of 15). Sessions keep a timeline with tool calls, webhook deliveries are logged per session, usage is broken down by key, and the dashboard has activity logs (12 of 15). The security page states a SOC 2 Type 2 report and third-party penetration tests. No bug bounty or disclosure policy was found, and robots.txt closes `/.well-known`, so security.txt went unread (12 of 20).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-security):

- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.
- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.
- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.
- 0 to 15, audit logs or per-call visibility for the operator.
- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.

Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.

## 4. Agent ergonomics, 72 out of 100, up to 4.6 more on the total

Why it scored 72: Scored as an API. Events are small typed JSON, HTTP tool responses are capped at 8 KiB before they reach the model, and the session list leaves artefacts out (18 of 25). `GET /v1/sessions` pages by `cursor` with `limit` from 1 to 200 and filters on `status` and `agent_id` (18 of 20). Errors carry `code`, `message` and sometimes `param`, and REST validation errors name the field (17 of 20). No idempotency keys on REST writes. Sessions resume within 30 seconds, webhooks dedupe on `event_id`, and `hold` mode silences the agent during long tools (10 of 20). An agent needs three fields (`name`, `system_prompt`, `voice`). The official Python and Node SDKs have no Voice Agent client, so callers write JSON over a socket or start from the two starter repositories (9 of 15).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):

- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).
- 20, pagination, filtering and output-size controls.
- 20, actionable, documented error responses, codes and messages an agent can recover from.
- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.
- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.

Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.

## 5. Maintenance & community, 64 out of 100, up to 3.2 more on the total

Why it scored 64: Universal-3.6 Pro Realtime, the speech-to-text the pricing page lists inside the Voice Agent API, shipped on 29 September 2026. The last changelog entry tagged Voice Agent API is 26 August 2026. We scored on the August entry, the last one tagged for this API, 44 days before the check (20). Ten entries tagged Voice Agent API between 13 July and 26 August 2026 (20). A public changelog that records fixes, a support address and a Discord link we did not test (11 of 15 for a closed service). The Python SDK (1.6.1) and Node SDK (4.41.5) were tagged on 24 September 2026 but have no Voice Agent client. The two starter repositories were last changed on 16 August 2026 (8 of 15). The starters need no dependencies and have no licence file, tests or CI (5 of 10).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):

- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.
- 20, at least three releases or dated changelog entries in the last 90 days.
- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.
- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).
- 10, package health, current dependencies and CI.

Models are read for deprecation notice periods and model churn rather than release counts.

## 6. Transparency & trust, 66 out of 100, up to 3 more on the total

Made of editorial 51, provenance 81.

Why it scored 66: A closed service with a clear service agreement. The starter repositories have no licence file (15 of 30). The Terms of Service grant a licence to train on Customer Data by default, the opt-out needs a paid plan, and the retention page gives periods for streaming, sync, dictation, asynchronous and LLM Gateway data but none for Voice Agent sessions, which are recorded and stored. A DPA is published and the privacy policy states Data Privacy Framework certification (14 of 30). The changelog dates deprecations ahead of time, such as the 11 July 2026 notice of a change on 7 August, and the terms promise reasonable prior notice of major changes. No written deprecation policy was found (12 of 20). The subprocessor list sits on a third-party trust centre we did not read. US and EU endpoints are mentioned, with no EU host named for the Voice Agent API in the pages read (10 of 20).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):

- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.
- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).
- 0 to 20, a deprecation policy or notices with dates.
- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).

The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.

Provenance checks not met in full (half of this category, computed from checked facts):

- Domain age: assemblyai.com, registered 2016-12-24 (9 years) (11 of 15)
- Terms of service: read, states 7 of the 7 things a reader expects, and has 2 clauses that cost points (6 of 10)
- Privacy policy: read, states 7 of the 8 things a reader expects (9.3 of 10)
- security.txt: could not be fetched (0 of 10)

## 7. Schema & documentation, 83 out of 100, up to 2.8 more on the total

Why it scored 83: An AsyncAPI 3.0 file describes the socket and an OpenAPI 3.0.3 file the token endpoint. No description file was found for the agents, sessions and webhook REST endpoints, and the AsyncAPI file lacks `language_codes`, `voice_focus`, `transcript.agent.delta` and `is_error`, which the docs added on 23 July 2026 (19 of 25). llms.txt, a Markdown copy of each page and a docs MCP server (10). Operations and fields say what they are for and when to use them, such as when to choose `hold` over `interactive` execution (17 of 20). Enums, ranges and defaults on audio encoding, turn detection and timeouts. Tool `parameters` are a JSON Schema the server does not validate at `session.update` (12 of 15). Examples in cURL, Python and JavaScript, and tables of session error codes with close codes (13 of 15). A dated public changelog tagged by product and a `/v1` path, with no written versioning policy found (12 of 15).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):

APIs and MCP servers.

- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).
- 10, llms.txt or Markdown docs served for agents.
- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.
- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.
- 0 to 15, examples and documented error responses.
- 15, versioning and a public changelog.

Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.

## What we couldn't check

What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.

- The Terms of Service (effective 1 July 2026) bar competitive analysis and benchmarking (2.4(f)), any automated or programmatic method of extracting data or Output (4.4), and use of the Services or Output for automated decision-making (2.4(j)). Recorded as a fact with no deduction. It matters before any probe is run.
- unchecked: the concurrent-session limit for Voice Agent accounts. No figure is in the Voice Agent docs, and the published streaming limits are not said to apply.
- unchecked: security.txt. robots.txt on www.assemblyai.com disallows `/.well-known`, so `securityTxt` is unknown.
- unchecked: the subprocessor list and the trust centre, which redirect to app.vanta.com and were not read.
- unchecked: the Data Processing Addendum, the acceptable use policy and the biometric data addendum were not read.
- unchecked: how long Voice Agent recordings and timelines are kept. The retention page has no Voice Agent section.
- unchecked: the EU host for the Voice Agent API. The troubleshooting page mentions an EU endpoint without naming it.
- unchecked: GitHub stars and package download counts, so `popularity` is empty.
- Whether the realtime speech-to-text incidents of 31 July and 6 October 2026 affected Voice Agent sessions is not stated on the status page.
- The RDAP host answered 400 to a robots.txt request, and one domain lookup was made after that.
- The docs carry text addressed to AI models (a note to fetch the docs index, an integration prompt and an `AGENTS.md`). It was treated as data.

## Weaknesses

- No number is published for the concurrent-session limit that the `concurrency_exceeded` error enforces
- status.assemblyai.com lists the Asynchronous API, Streaming API and LLM Gateway, with no Voice Agent component
- Telephony is inbound only, over a SIP trunk on a Twilio account the customer owns
- API keys have no scopes or expiry, and opting out of model training needs a paid plan
- The Terms of Service bar benchmarking and competitive analysis (2.4(f)) and any automated or programmatic method of extracting data or Output (4.4), which matters before any probe is run

## What costs an agent a turn today

The notes we give agents before they call it. Each one is a workaround an agent shouldn't need.

- Send `session.end` before closing the socket. Without it the session stays resumable, and billed, for 30 seconds
- Retry only `at_capacity`, `concurrency_exceeded` and `internal_error`, with backoff. Every other session error code is fatal
- Send `tool.result` inside the `reply.done` handler, not straight after `tool.call`
- Mint a one-time token with `GET /v1/token` for browser clients and never ship the API key to a page
- Deduplicate webhooks on `event_id`, and verify `X-AAI-Signature` against the raw body bytes because a 401 is not retried

## When it's done

Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.

What we couldn't check

  • The Terms of Service (effective 1 July 2026) bar competitive analysis and benchmarking (2.4(f)), any automated or programmatic method of extracting data or Output (4.4), and use of the Services or Output for automated decision-making (2.4(j)). Recorded as a fact with no deduction. It matters before any probe is run.
  • unchecked: the concurrent-session limit for Voice Agent accounts. No figure is in the Voice Agent docs, and the published streaming limits are not said to apply.
  • unchecked: security.txt. robots.txt on www.assemblyai.com disallows /.well-known, so securityTxt is unknown.
  • unchecked: the subprocessor list and the trust centre, which redirect to app.vanta.com and were not read.
  • unchecked: the Data Processing Addendum, the acceptable use policy and the biometric data addendum were not read.
  • unchecked: how long Voice Agent recordings and timelines are kept. The retention page has no Voice Agent section.
  • unchecked: the EU host for the Voice Agent API. The troubleshooting page mentions an EU endpoint without naming it.
  • unchecked: GitHub stars and package download counts, so popularity is empty.
  • Whether the realtime speech-to-text incidents of 31 July and 6 October 2026 affected Voice Agent sessions is not stated on the status page.
  • The RDAP host answered 400 to a robots.txt request, and one domain lookup was made after that.
  • The docs carry text addressed to AI models (a note to fetch the docs index, an integration prompt and an AGENTS.md). It was treated as data.

Sources 33

  1. Voice Agent API overview (Markdown copy) assemblyai.com · seen 2026-10-09
  2. docs index assemblyai.com · seen 2026-10-09
  3. AsyncAPI description of the socket, read as the file and not the rendered reference page assemblyai.com · seen 2026-10-09
  4. OpenAPI description of the token endpoint, read as the file assemblyai.com · seen 2026-10-09
  5. events reference and error codes assemblyai.com · seen 2026-10-09
  6. REST reference for agents assemblyai.com · seen 2026-10-09
  7. tools overview assemblyai.com · seen 2026-10-09
  8. HTTP tools assemblyai.com · seen 2026-10-09
  9. recordings and transcripts assemblyai.com · seen 2026-10-09
  10. webhooks assemblyai.com · seen 2026-10-09
  11. inbound phone agent over SIP assemblyai.com · seen 2026-10-09
  12. connect your own LLM assemblyai.com · seen 2026-10-09
  13. browser integration and tokens assemblyai.com · seen 2026-10-09
  14. troubleshooting assemblyai.com · seen 2026-10-09
  15. build with AI coding tools assemblyai.com · seen 2026-10-09
  16. pricing assemblyai.com · seen 2026-10-09
  17. billing and free credit assemblyai.com · seen 2026-10-09
  18. Terms of Service, effective 1 July 2026 assemblyai.com · seen 2026-10-09
  19. privacy policy, effective 26 May 2026 assemblyai.com · seen 2026-10-09
  20. Service Level Agreement, effective 13 April 2026 assemblyai.com · seen 2026-10-09
  21. data retention and model training assemblyai.com · seen 2026-10-09
  22. data controls assemblyai.com · seen 2026-10-09
  23. account management, keys and roles assemblyai.com · seen 2026-10-09
  24. security page assemblyai.com · seen 2026-10-09
  25. changelog assemblyai.com · seen 2026-10-09
  26. status page status.assemblyai.com · seen 2026-10-09
  27. status history feed status.assemblyai.com · seen 2026-10-09
  28. Python starter repository (shallow clone) github.com · seen 2026-10-09
  29. JavaScript starter repository (shallow clone) github.com · seen 2026-10-09
  30. Python SDK tags (shallow clone) github.com · seen 2026-10-09
  31. Node SDK tags (shallow clone) github.com · seen 2026-10-09
  32. robots.txt assemblyai.com · seen 2026-10-09
  33. domain registration (RDAP) rdap.verisign.com · seen 2026-10-09

Probe metrics

Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. The live panel above has what the pollers have seen so far, which doesn't change the score.

Pricing & changes

Pay per use Pay per use $4.50 an hour ($0.075 a minute), billed per second from socket open to session end, idle time included. The rate covers speech-to-text, the managed model, speech output, recordings and hosting. New accounts get $50 of credit with no card, which covers the Voice Agent API. LLM Gateway tokens and Twilio carrier minutes are billed separately. Volume pricing is by sales contact (https://www.assemblyai.com/pricing, https://www.assemblyai.com/docs/billing-and-pricing, checked 2026-10-09).

Prices

ItemPriceUnitNote
Voice Agent session$0.075per minute of call$4.50 an hour billed per second of open socket. Speech-to-text, managed model, speech output and recordings included. LLM Gateway tokens and carrier minutes extra

Compared across listings on the price index.

Recent changes

  • Latest release

Follow them as a feed at /feeds/tools/assemblyai-voice-agent.xml, or this listing's score history at history.json.

Connect

Install

git clone https://github.com/AssemblyAI/voice-agent-starter-python

First request

curl -X POST https://agents.assemblyai.com/v1/agents \
  -H "Authorization: $ASSEMBLYAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "Support Assistant", "system_prompt": "You are a friendly support agent.", "greeting": "Hi, how can I help?", "voice": {"voice_id": "alba"}}'

Claude Code

claude mcp add --transport http --scope user assemblyai-docs https://mcp.assemblyai.com/docs

Through letme picks today, calling later

GET https://letme.dev/assemblyai-voice-agent

letme.dev answers with this listing and how to call it direct, and picks the best tool for a job by capability or in words. Calling through letme (one key, the vendor's own price) comes later. Nothing on letme.dev is for people to look at; this page explains it.

Similar toolGrade ScoreShared capabilitiesx402
ElevenLabs Agents API + MCP ElevenLabsBB71.3voice.agent voice.pipeline voice.tools voice.telephonyno
Retell AI API + MCP Retell AIB69.1voice.agent voice.pipeline voice.tools voice.telephonyno
Bland AI API + MCP Bland AIB63.8voice.agent voice.pipeline voice.tools voice.telephonyno
Vapi API + MCP VapiB63.5voice.agent voice.pipeline voice.tools voice.telephonyno
Hume EVI (Empathic Voice Interface) Hume AIC56.7voice.agent voice.pipeline voice.tools voice.telephonyno
Vocily AI Vocily AI Private LimitedD53.8voice.agent voice.pipeline voice.tools voice.telephonyno

Machine-readable

Verify this listing

For the vendor

Is this your product? Link to this page from your own site or README, then tell us where. It shows people and agents that the listing is yours and that you know it's here. It never changes a grade, rank or review.

  1. Add the badge or a link

    AssemblyAI Voice Agent API on Anchor Terminal, B, 64.4/100
    On a light page
    On a dark page
    <a href="https://www.anchorterminal.com/tools/assemblyai-voice-agent"><img src="https://www.anchorterminal.com/badges/assemblyai-voice-agent.svg" alt="AssemblyAI Voice Agent API on Anchor Terminal" height="20"></a>
    [![AssemblyAI Voice Agent API on Anchor Terminal](https://www.anchorterminal.com/badges/assemblyai-voice-agent.svg)](https://www.anchorterminal.com/tools/assemblyai-voice-agent)

    It counts on a page on assemblyai.com or one of its subdomains, or the README of github.com/AssemblyAI/voice-agent-starter-python.

  2. Tell us where it is

    We read it once now and again every week. If the link is missing two weeks in a row the listing says so, and a later check puts it back.

Agents send the same to POST /api/v1/verify as {"slug": "assemblyai-voice-agent", "url": "…"}, or call the verify_listing tool at /mcp. Ten checks an hour from one address. What we check. To announce the listing, get sharing assets for social media.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.