{
  "fixes": {
    "slug": "openai-realtime",
    "name": "OpenAI Realtime API",
    "listing": "https://www.anchorterminal.com/tools/openai-realtime",
    "markdown": "# Fix list: OpenAI Realtime API\n\nFrom Anchor Terminal's listing at https://www.anchorterminal.com/tools/openai-realtime, the October 2026 research run, assessed 9 October 2026. Grade B, 63.3 out of 100.\n\nThis is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.\n\nFor a coding agent working on OpenAI Realtime API: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.\n\n## 1. Payments \u0026 pricing, 20 out of 100, up to 10 more on the total\n\nWhy it scored 20: No x402, MPP or L402 in the docs or the OpenAPI file (0 of 40). Per-token prices for every Realtime model are published without a login (20). The rate limits page lists a Free tier with a $100 monthly usage limit, but whether it covers Realtime models without a card was not established, and the error guide ties access to prepaid credits (0 of 20). API keys are created in the platform settings after a browser sign-in (0).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):\n\nThe published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).\n\n- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).\n\nPayment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.\n\nOpen-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.\n\n## 2. Security \u0026 auth, 61 out of 100, up to 6.8 more on the total\n\nWhy it scored 61: Server connections send the API key as a Bearer header. Browsers use client secrets from `POST /v1/realtime/client_secrets` that last 10 seconds to 2 hours, 10 minutes by default, and can open several sessions until expiry. The browser WebSocket example passes the secret as a subprotocol, not in the query string. The docs index also lists workload identity federation, mutual TLS and an IP allowlist. Key scopes were not read (24 of 30). MCP tools take `allowed_tools` and `require_approval`, which defaults to `always` in the OpenAPI file, and a sideband connection keeps tool execution on the server. Session settings attached to a client secret can be overridden by the client connection (15 of 20). The model hears callers and reads MCP results. The SIP guide says to treat SIP headers as untrusted, and no prompt-injection guidance specific to Realtime was found (6 of 15). The docs index describes Admin APIs with audit log retrieval, sessions can write traces, and usage is reported per response (10 of 15). openai.com answered 403, so security.txt, any bug bounty and certifications are unchecked. We read a private vulnerability reporting policy on the OpenAPI repository and a 5 October changelog entry on HIPAA support (6 of 20).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-security):\n\n- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.\n\nModels are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.\n\n## 3. Reliability, 67 out of 100, up to 6.6 more on the total\n\nWhy it scored 67: Graded as a hosted API. status.openai.com lists Realtime as one of 13 API components with history (20). The page shows 100 per cent uptime for the Realtime component for July to October 2026, and the status feed names Realtime among the affected components of five resolved incidents on 24 July, 25 July (two), 17 September and 6 October. The feed gives only the final update, so durations were not read, and we score them as minor (20 of 30). The rate limits page gives usage tiers and monthly usage limits, but per-model limits for Realtime are shown only in account settings and the public table covers GPT-6 families only (5 of 15). The docs describe `Retry-After` on 429 and 503, exponential backoff with jitter and SDK retries, sessions emit `rate_limits.updated`, and SIP webhooks carry a `webhook-id` for deduplication. No idempotency key for call control was found (12 of 15). No SLA was found in the pages read, and openai.com answered 403 (0). The Realtime API has been generally available since 28 August 2025 (10).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):\n\nHosted APIs, MCP servers, models and platforms.\n\n- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.\n\nLocal packages, SDKs, frameworks and stdio MCP servers.\n\n- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.\n\nProtocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.\n\n## 4. Maintenance \u0026 community, 52 out of 100, up to 4.2 more on the total\n\nWhy it scored 52: The last dated changelog entry tagged `v1/realtime` that adds a feature is 28 July 2026, 73 days before the check, when `gpt-transcribe` became available for final transcripts of Realtime turns. `gpt-realtime-2.1` shipped on 6 July (20 of 30). Two entries tagged `v1/realtime` fall in the last 90 days, 28 July and 26 August, one short of the three the checklist asks for (0 of 20). Closed service with a changelog updated several times a week, plus a community forum and Discord linked from the docs, which we did not test (10 of 15). Current official SDKs, with `@openai/agents-realtime` 0.20.0 tagged on 8 October 2026 and 12 version tags since 29 July (15). The SDK is MIT but still 0.x, and its CI results were not read (7 of 10).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):\n\n- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.\n\nModels are read for deprecation notice periods and model churn rather than release counts.\n\n## 5. Agent ergonomics, 76 out of 100, up to 3.9 more on the total\n\nWhy it scored 76: Scored as a model API. Context can be capped with `truncation`, `retention_ratio` and `token_limits.post_instructions`, items can be deleted, prompt caching is automatic, and `response.done` reports token usage by modality (22 of 25). There are no lists to page. Output controls are `output_modalities`, `max_output_tokens`, optional input transcription and `allowed_tools` (14 of 20). The `error` event returns the `event_id` of the client event that caused it, and MCP tool listing and calls have their own failure events (15 of 20). Client events take an `event_id`, webhooks carry a `webhook-id`, and interruptions cancel the response. No session resumption after a dropped connection and no idempotency key for call control were found in the pages read (10 of 20). A session needs only a model, voice activity detection is automatic, and the Agents SDK starts a browser voice agent in about ten lines, with examples in five languages (15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):\n\n- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.\n\nModels are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.\n\n## 6. Schema \u0026 documentation, 84 out of 100, up to 2.6 more on the total\n\nWhy it scored 84: The OpenAPI 3.1 file in `openai/openai-openapi` (updated 8 October 2026) has nine `/realtime` paths and 167 Realtime schemas, with 13 client events and 48 server events typed. It is not an AsyncAPI document, so the WebSocket channel itself is not described (22 of 25). llms.txt and a Markdown twin of every docs page (10). The tools guide has a table of when to use `function`, `mcp` with `server_url` or a connector, and the transport guides say when to choose WebRTC, WebSocket or SIP (16 of 20). Session fields carry enums and ranges, such as `expires_after.seconds` from 10 to 7200, though `model` also accepts any string (13 of 15). Examples in JavaScript, Python, Java, Ruby and C#, an `error` event with type, code, message and param, and a list of common MCP failures. No catalogue of Realtime error codes was found (11 of 15). Dated changelog and dated model snapshots, less 3 because the OpenAPI file still lists `gpt-4o-realtime-preview` models and beta event schemas that the deprecations page says were removed in May 2026, and several guides mix GPT-Live and Realtime text on one page (12 of 15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):\n\nAPIs and MCP servers.\n\n- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.\n\nModels are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.\n\n## 7. Transparency \u0026 trust, 71 out of 100, up to 2.5 more on the total\n\nMade of editorial 58, provenance 84.\n\nWhy it scored 71: Closed service. The SDKs and the OpenAPI file are MIT. The service terms were not read because openai.com answered 403 (10 of 30). The data controls page says API data is not used for training unless the customer opts in, abuse monitoring logs are kept up to 30 days, `/v1/realtime` stores no application state and is eligible for Zero Data Retention with approval. The privacy policy and any DPA were not read, so agreement between them is unchecked (18 of 30). The deprecations page states at least six months' notice for generally available models and lists shutdown dates with replacements (20 of 20). Data residency exists for the United States and Europe, with a `sip-eu.api.openai.com` endpoint, and tracing on `/v1/realtime` is not EU residency compliant. The sub-processor list is linked but was not read (10 of 20).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):\n\n- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).\n\nThe other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.\n\nProvenance checks not met in full (half of this category, computed from checked facts):\n\n- Terms of service: published, but our reader couldn't read it (7 of 10)\n- Privacy policy: published, but our reader couldn't read it (7 of 10)\n- security.txt: could not be fetched (0 of 10)\n\n## What we couldn't check\n\nWhat we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.\n\n- unchecked: the service terms, privacy policy, DPA, sub-processor list, security and trust pages and security.txt on openai.com. The host answered 403 to one request on 9 October 2026 and we did not retry. `provenance.terms` and `provenance.privacy` are left out for that reason.\n- unchecked: whether an SLA is published for any tier. None was found in the developer docs. The rate limits page links Scale Tier and Reserved Tier pages on openai.com, which were not read.\n- unchecked: whether the Free usage tier gives access to Realtime models without a card or prepaid credits.\n- unchecked: duration and severity of the five incidents that name Realtime. The status feed shows only the final update of each.\n- unchecked: per-model rate limits and concurrent session limits for Realtime, which the docs place in account settings, and the maximum session length.\n- unchecked: the model pages for `gpt-realtime-2.1` and `gpt-realtime-2.1-mini` (context window, snapshot names), the WebRTC guide, the voice activity detection guide and the full event reference. We stopped at about twenty requests to developers.openai.com, five over the brief's guide of fifteen.\n- unchecked: GitHub stars and package download counts. The legal entity was not read. The domain openai.com was registered on 19 January 2007 per RDAP.\n- OpenAI shipped GPT-Live (`gpt-live-1`, `/v1/live/sessions`) as generally available on 10 September 2026 and publishes a guide for migrating from Realtime. No shutdown date for the Realtime API was found. GPT-Live is a separate product and is not graded here.\n- `gpt-realtime-mini-2025-10-06` was shut down on 23 July 2026 after an announcement on 22 April 2026, three months' notice for a dated snapshot. Whether that snapshot counts as a generally available model under the six-month policy is not clear.\n- The lead was right on the interface and the example model. It gave openai.com as the URL, and the documentation and reference are on developers.openai.com.\n\n## Weaknesses\n\n- openai.com answered 403 to our reader, so the service terms, privacy policy, sub-processor list and security pages were not read.\n- Realtime rate limits are shown only in account settings. The public table gives numbers for GPT-6 model families only.\n- Every response re-sends the whole conversation, so later turns cost more. Audio is $32 in and $64 out per 1M tokens on `gpt-realtime-2.1`.\n- The status feed names Realtime among affected components in five incidents between 24 July and 6 October 2026, with durations not shown in the feed.\n- Session settings attached to a client secret can be overridden by the client connection, per the client secret reference.\n\n## What costs an agent a turn today\n\nThe notes we give agents before they call it. Each one is a workaround an agent shouldn't need.\n\n- Create a client secret on a server with `POST /v1/realtime/client_secrets` and give browsers only the `ek_` value. Never ship the API key.\n- After a response with MCP calls finishes, send another `response.create`. The API does not create the follow-up response itself.\n- Set `truncation` with a `retention_ratio` below 1 and `token_limits.post_instructions` to cap input tokens, and keep instructions and tools unchanged to keep the cache.\n- On a WebSocket, stop playback on `input_audio_buffer.speech_started` and send `conversation.item.truncate` with `audio_end_ms`. WebRTC and SIP truncate on the server.\n- Use `gpt-realtime-2.1` or `gpt-realtime-2.1-mini`. `gpt-realtime` and `gpt-realtime-mini` shut down on 20 January 2027.\n\n## When it's done\n\nSend what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `\"kind\": \"dispute\"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.\n",
    "grade": "B",
    "score": 63.3,
    "assessed": "2026-10-09",
    "run": "October 2026 research run",
    "categories": [
      {
        "key": "payments",
        "name": "Payments \u0026 pricing",
        "score": 20,
        "maxGain": 10,
        "reason": "No x402, MPP or L402 in the docs or the OpenAPI file (0 of 40). Per-token prices for every Realtime model are published without a login (20). The rate limits page lists a Free tier with a $100 monthly usage limit, but whether it covers Realtime models without a card was not established, and the error guide ties access to prepaid credits (0 of 20). API keys are created in the platform settings after a browser sign-in (0).",
        "checklist": [
          "The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).",
          "- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).",
          "Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.",
          "Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-payments"
      },
      {
        "key": "security",
        "name": "Security \u0026 auth",
        "score": 61,
        "maxGain": 6.8,
        "reason": "Server connections send the API key as a Bearer header. Browsers use client secrets from `POST /v1/realtime/client_secrets` that last 10 seconds to 2 hours, 10 minutes by default, and can open several sessions until expiry. The browser WebSocket example passes the secret as a subprotocol, not in the query string. The docs index also lists workload identity federation, mutual TLS and an IP allowlist. Key scopes were not read (24 of 30). MCP tools take `allowed_tools` and `require_approval`, which defaults to `always` in the OpenAPI file, and a sideband connection keeps tool execution on the server. Session settings attached to a client secret can be overridden by the client connection (15 of 20). The model hears callers and reads MCP results. The SIP guide says to treat SIP headers as untrusted, and no prompt-injection guidance specific to Realtime was found (6 of 15). The docs index describes Admin APIs with audit log retrieval, sessions can write traces, and usage is reported per response (10 of 15). openai.com answered 403, so security.txt, any bug bounty and certifications are unchecked. We read a private vulnerability reporting policy on the OpenAPI repository and a 5 October changelog entry on HIPAA support (6 of 20).",
        "checklist": [
          "- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.",
          "Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-security"
      },
      {
        "key": "reliability",
        "name": "Reliability",
        "score": 67,
        "maxGain": 6.6,
        "reason": "Graded as a hosted API. status.openai.com lists Realtime as one of 13 API components with history (20). The page shows 100 per cent uptime for the Realtime component for July to October 2026, and the status feed names Realtime among the affected components of five resolved incidents on 24 July, 25 July (two), 17 September and 6 October. The feed gives only the final update, so durations were not read, and we score them as minor (20 of 30). The rate limits page gives usage tiers and monthly usage limits, but per-model limits for Realtime are shown only in account settings and the public table covers GPT-6 families only (5 of 15). The docs describe `Retry-After` on 429 and 503, exponential backoff with jitter and SDK retries, sessions emit `rate_limits.updated`, and SIP webhooks carry a `webhook-id` for deduplication. No idempotency key for call control was found (12 of 15). No SLA was found in the pages read, and openai.com answered 403 (0). The Realtime API has been generally available since 28 August 2025 (10).",
        "checklist": [
          "Hosted APIs, MCP servers, models and platforms.",
          "- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.",
          "Local packages, SDKs, frameworks and stdio MCP servers.",
          "- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.",
          "Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-reliability"
      },
      {
        "key": "maintenance",
        "name": "Maintenance \u0026 community",
        "score": 52,
        "maxGain": 4.2,
        "reason": "The last dated changelog entry tagged `v1/realtime` that adds a feature is 28 July 2026, 73 days before the check, when `gpt-transcribe` became available for final transcripts of Realtime turns. `gpt-realtime-2.1` shipped on 6 July (20 of 30). Two entries tagged `v1/realtime` fall in the last 90 days, 28 July and 26 August, one short of the three the checklist asks for (0 of 20). Closed service with a changelog updated several times a week, plus a community forum and Discord linked from the docs, which we did not test (10 of 15). Current official SDKs, with `@openai/agents-realtime` 0.20.0 tagged on 8 October 2026 and 12 version tags since 29 July (15). The SDK is MIT but still 0.x, and its CI results were not read (7 of 10).",
        "checklist": [
          "- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.",
          "Models are read for deprecation notice periods and model churn rather than release counts."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-maintenance"
      },
      {
        "key": "ergonomics",
        "name": "Agent ergonomics",
        "score": 76,
        "maxGain": 3.9,
        "reason": "Scored as a model API. Context can be capped with `truncation`, `retention_ratio` and `token_limits.post_instructions`, items can be deleted, prompt caching is automatic, and `response.done` reports token usage by modality (22 of 25). There are no lists to page. Output controls are `output_modalities`, `max_output_tokens`, optional input transcription and `allowed_tools` (14 of 20). The `error` event returns the `event_id` of the client event that caused it, and MCP tool listing and calls have their own failure events (15 of 20). Client events take an `event_id`, webhooks carry a `webhook-id`, and interruptions cancel the response. No session resumption after a dropped connection and no idempotency key for call control were found in the pages read (10 of 20). A session needs only a model, voice activity detection is automatic, and the Agents SDK starts a browser voice agent in about ten lines, with examples in five languages (15).",
        "checklist": [
          "- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.",
          "Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-ergonomics"
      },
      {
        "key": "schema",
        "name": "Schema \u0026 documentation",
        "score": 84,
        "maxGain": 2.6,
        "reason": "The OpenAPI 3.1 file in `openai/openai-openapi` (updated 8 October 2026) has nine `/realtime` paths and 167 Realtime schemas, with 13 client events and 48 server events typed. It is not an AsyncAPI document, so the WebSocket channel itself is not described (22 of 25). llms.txt and a Markdown twin of every docs page (10). The tools guide has a table of when to use `function`, `mcp` with `server_url` or a connector, and the transport guides say when to choose WebRTC, WebSocket or SIP (16 of 20). Session fields carry enums and ranges, such as `expires_after.seconds` from 10 to 7200, though `model` also accepts any string (13 of 15). Examples in JavaScript, Python, Java, Ruby and C#, an `error` event with type, code, message and param, and a list of common MCP failures. No catalogue of Realtime error codes was found (11 of 15). Dated changelog and dated model snapshots, less 3 because the OpenAPI file still lists `gpt-4o-realtime-preview` models and beta event schemas that the deprecations page says were removed in May 2026, and several guides mix GPT-Live and Realtime text on one page (12 of 15).",
        "checklist": [
          "APIs and MCP servers.",
          "- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.",
          "Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-schema"
      },
      {
        "key": "transparency",
        "name": "Transparency \u0026 trust",
        "score": 71,
        "maxGain": 2.5,
        "reason": "Closed service. The SDKs and the OpenAPI file are MIT. The service terms were not read because openai.com answered 403 (10 of 30). The data controls page says API data is not used for training unless the customer opts in, abuse monitoring logs are kept up to 30 days, `/v1/realtime` stores no application state and is eligible for Zero Data Retention with approval. The privacy policy and any DPA were not read, so agreement between them is unchecked (18 of 30). The deprecations page states at least six months' notice for generally available models and lists shutdown dates with replacements (20 of 20). Data residency exists for the United States and Europe, with a `sip-eu.api.openai.com` endpoint, and tracing on `/v1/realtime` is not EU residency compliant. The sub-processor list is linked but was not read (10 of 20).",
        "blend": "editorial 58, provenance 84",
        "checklist": [
          "- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).",
          "The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-transparency"
      }
    ],
    "provenance": [
      {
        "label": "Terms of service",
        "value": "published, but our reader couldn't read it",
        "points": 7,
        "max": 10
      },
      {
        "label": "Privacy policy",
        "value": "published, but our reader couldn't read it",
        "points": 7,
        "max": 10
      },
      {
        "label": "security.txt",
        "value": "could not be fetched",
        "points": 0,
        "max": 10
      }
    ],
    "unchecked": [
      "unchecked: the service terms, privacy policy, DPA, sub-processor list, security and trust pages and security.txt on openai.com. The host answered 403 to one request on 9 October 2026 and we did not retry. `provenance.terms` and `provenance.privacy` are left out for that reason.",
      "unchecked: whether an SLA is published for any tier. None was found in the developer docs. The rate limits page links Scale Tier and Reserved Tier pages on openai.com, which were not read.",
      "unchecked: whether the Free usage tier gives access to Realtime models without a card or prepaid credits.",
      "unchecked: duration and severity of the five incidents that name Realtime. The status feed shows only the final update of each.",
      "unchecked: per-model rate limits and concurrent session limits for Realtime, which the docs place in account settings, and the maximum session length.",
      "unchecked: the model pages for `gpt-realtime-2.1` and `gpt-realtime-2.1-mini` (context window, snapshot names), the WebRTC guide, the voice activity detection guide and the full event reference. We stopped at about twenty requests to developers.openai.com, five over the brief's guide of fifteen.",
      "unchecked: GitHub stars and package download counts. The legal entity was not read. The domain openai.com was registered on 19 January 2007 per RDAP.",
      "OpenAI shipped GPT-Live (`gpt-live-1`, `/v1/live/sessions`) as generally available on 10 September 2026 and publishes a guide for migrating from Realtime. No shutdown date for the Realtime API was found. GPT-Live is a separate product and is not graded here.",
      "`gpt-realtime-mini-2025-10-06` was shut down on 23 July 2026 after an announcement on 22 April 2026, three months' notice for a dated snapshot. Whether that snapshot counts as a generally available model under the six-month policy is not clear.",
      "The lead was right on the interface and the example model. It gave openai.com as the URL, and the documentation and reference are on developers.openai.com."
    ],
    "weaknesses": [
      "openai.com answered 403 to our reader, so the service terms, privacy policy, sub-processor list and security pages were not read.",
      "Realtime rate limits are shown only in account settings. The public table gives numbers for GPT-6 model families only.",
      "Every response re-sends the whole conversation, so later turns cost more. Audio is $32 in and $64 out per 1M tokens on `gpt-realtime-2.1`.",
      "The status feed names Realtime among affected components in five incidents between 24 July and 6 October 2026, with durations not shown in the feed.",
      "Session settings attached to a client secret can be overridden by the client connection, per the client secret reference."
    ],
    "agentNotes": [
      "Create a client secret on a server with `POST /v1/realtime/client_secrets` and give browsers only the `ek_` value. Never ship the API key.",
      "After a response with MCP calls finishes, send another `response.create`. The API does not create the follow-up response itself.",
      "Set `truncation` with a `retention_ratio` below 1 and `token_limits.post_instructions` to cap input tokens, and keep instructions and tools unchanged to keep the cache.",
      "On a WebSocket, stop playback on `input_audio_buffer.speech_started` and send `conversation.item.truncate` with `audio_end_ms`. WebRTC and SIP truncate on the server.",
      "Use `gpt-realtime-2.1` or `gpt-realtime-2.1-mini`. `gpt-realtime` and `gpt-realtime-mini` shut down on 20 January 2027."
    ],
    "recheck": "https://www.anchorterminal.com/builders/#disputes"
  },
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-10",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  }
}
