{
  "fixes": {
    "slug": "openai-decisions-api",
    "name": "OpenAI Decisions API",
    "listing": "https://www.anchorterminal.com/tools/openai-decisions-api",
    "markdown": "# Fix list: OpenAI Decisions API\n\nFrom Anchor Terminal's listing at https://www.anchorterminal.com/tools/openai-decisions-api, the October 2026 research run, assessed 8 October 2026. Grade BB, 71.5 out of 100.\n\nThis is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.\n\nFor a coding agent working on OpenAI Decisions API: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.\n\n## 1. Reliability, 53 out of 100, up to 9.4 more on the total\n\nWhy it scored 53: Read as a hosted model API. status.openai.com has a Decisions component under APIs, reading 100 per cent on 8 October (20). The endpoint was released on 6 October, so its record is two days long. The incident feed, read back to 10 July, names Decisions in no incident. The platform under it had API-wide error incidents on 25 July, 17 September and 29 September (5 hours 22 minutes), all before the release, so 8 of 30 for a clean record too short to count, the reading given to Clef on its first day. The GPT-6 Luna page lists 5,000 requests and 2,000,000 tokens a minute on the Build tier, 10,000 and 10,000,000 on Launch and 30,000 and 180,000,000 on Grow, but its endpoint table doesn't list Decisions and the guide gives no limits of its own (10 of 15). The rate-limits guide covers `Retry-After` on 429 and 503, backoff with jitter and SDK retries, and a decision has no side effects, so a retry is safe (15). No SLA covering a beta endpoint was found. The Scale Tier page answered 403 on 8 October (0). Public beta (0).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):\n\nHosted APIs, MCP servers, models and platforms.\n\n- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.\n\nLocal packages, SDKs, frameworks and stdio MCP servers.\n\n- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.\n\nProtocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.\n\n## 2. Payments \u0026 pricing, 30 out of 100, up to 8.8 more on the total\n\nWhy it scored 30: No x402, MPP or L402 in the guide, the reference or the pricing page (0). Per-token price published without a login, $0.10 per million input tokens with no output or cache charge (20). The rate-limits guide lists a Free tier capped at $100 a month in allowed countries, but the GPT-6 Luna page's rate-limit table starts at Build, which needs $5 of credit, and the Decisions guide doesn't mention a free allowance, so half, as for the OpenAI API listing (10). A person signs up in a browser and creates a key (0).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):\n\nThe published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).\n\n- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).\n\nPayment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.\n\nOpen-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.\n\n## 3. Security \u0026 auth, 83 out of 100, up to 3 more on the total\n\nWhy it scored 83: Read for a model API, on the lines used for Jev. Project API keys, with the access-control guide saying a key carries its own permissions that are checked against the caller's project role, service-account keys created through the Administration API with a documented rotation procedure, and workload identity federation with short-lived tokens. We didn't confirm that a restricted key can be limited to `/v1/decisions`, since the help centre article answered 403 (27 of 30). The data-controls guide lists `/v1/decisions` as not used for training, with abuse-monitoring logs kept up to 30 days and Zero Data Retention for eligible customers. An image flagged by the CSAM classifier is kept for review under any setting, and the Private Retention column reads pending confirmation (18 of 20). The endpoint judges text and images that often come from users. Neither the guide nor the reference mentions adversarial or injected content. The safety guide advises red-teaming and limiting input length, and requests accept a `safety_identifier` (6 of 15). Each response reports token usage, and the Admin APIs guide documents audit-log retrieval. We didn't confirm how Decisions calls appear in usage views (12 of 15). security.txt is PGP-signed and names a Bugcrowd bounty and a coordinated disclosure policy, though it has no Expires field, and trust.openai.com names SOC 2 Type 2 and ISO/IEC 27001, 27017, 27018, 27701 and 42001 (20).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-security):\n\n- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.\n\nModels are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.\n\n## 4. Agent ergonomics, 87 out of 100, up to 2.1 more on the total\n\nWhy it scored 87: Read as an API an agent calls for a decision, the reading used for Jev. Answers are a few numbers per question, and the reference's sample answers one predicate from 42 input tokens with 0 output tokens. Up to 128 inline images a request. The usage object reports cached tokens and the guide says cache reads and writes aren't charged. `/v1/decisions` isn't among the endpoints the Batch guide lists, and the guide states no context limit for the endpoint (20 of 25). The caller sets the whole output shape, with up to 200 independent questions about one input in a call (18 of 20). The error-codes guide gives types, codes and recovery steps for 400, 401, 403, 429, 500 and 503, and a question the model declines comes back as a typed `refusal` answer beside the others. Errors particular to the endpoint aren't documented (16 of 20). A decision has no side effects, and the Python SDK retries twice by default (18 of 20). Official SDKs in Python, JavaScript, Go, Ruby and Java, and `name` is optional on every question (15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):\n\n- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.\n\nModels are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.\n\n## 5. Maintenance \u0026 community, 77 out of 100, up to 2 more on the total\n\nWhy it scored 77: Read for a closed model, on the lines used for Jev. Released in beta on 6 October 2026 (30). OpenAI publishes minimum notice of 6 months for generally available models and 3 months for specialised variants, and says preview models can go with 2 weeks. Nothing states what a beta endpoint gets (6 of 12). Two days old with one model alias and no dated snapshot, too short to show churn (4 of 8). A dated changelog with several entries a month, including this release. The help centre answered 403 and issue replies weren't sampled (12 of 15). Official SDKs gained Decisions on release day, Python 3.26.0 and JavaScript 7.30.0 on 6 October, with 3.26.1 and 7.30.1 on 8 October, and the guide names Go 3.73.0, Ruby 0.101.0 and Java 4.78.0 (15). The Python SDK states Python 3.10 or later, the JavaScript SDK Node 22 or later, and the Python repository carries CI and breaking-change workflows (10).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):\n\n- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.\n\nModels are read for deprecation notice periods and model churn rather than release counts.\n\n## 6. Schema \u0026 documentation, 89 out of 100, up to 1.8 more on the total\n\nWhy it scored 89: OpenAI's OpenAPI 3.1 document (version 2.3.0, committed 8 October) has `POST /decisions` with `DecisionRequest` and `DecisionResponse` schemas (25). llms.txt indexes the guide, and the guide and the reference page each have a Markdown twin (10). The guide says which question type fits which job and when to use Structured Outputs or function calling instead. No page of known limits for the model on this endpoint was found (16 of 20). Questions are a union discriminated on `type`, with 1 to 200 questions, 2 to 255 choices and 2 to 10 levels set in the schema and no extra properties allowed (15). The guide has examples in curl, JavaScript, Python, Go, Java and Ruby with sample responses. The OpenAPI operation documents only the 200 response, and errors are covered by the general error-codes guide (12 of 15). A dated changelog carries the release on 6 October. The only model ID is the alias `gpt-6-luna` with no dated snapshot, and the beta has no version of its own (11 of 15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):\n\nAPIs and MCP servers.\n\n- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.\n\nModels are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.\n\n## 7. Transparency \u0026 trust, 83 out of 100, up to 1.5 more on the total\n\nMade of editorial 71, provenance 94.\n\nWhy it scored 83: Closed service with Apache-2.0 SDKs and an MIT OpenAPI repository. The services agreement answered 403 on 8 October, so its terms stand per the OpenAI API listing's earlier check (15). The data-controls guide, read on 8 October, gives `/v1/decisions` its own row and section covering training, 30-day abuse logs, Zero Data Retention, the 24-hour prompt cache and HIPAA eligibility. One column reads pending confirmation, and the DPA wasn't re-read (403), per the 4 October check on the OpenAI API listing (26 of 30). The deprecations page gives dated notices and minimum notice periods for models. It says nothing about beta endpoints, and the guide gives no date for general availability (14 of 20). The guide lists ten data-residency regions with Decisions in each, and regional processing for it in the United States and Europe. The subprocessor list answered 403 on 8 October, per the 4 October check (16 of 20).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):\n\n- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).\n\nThe other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.\n\nProvenance checks not met in full (half of this category, computed from checked facts):\n\n- Terms of service: published, but our reader couldn't read it (7 of 10)\n- Privacy policy: published, but our reader couldn't read it (7 of 10)\n\n## What we couldn't check\n\nWhat we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.\n\n- The scout gave the launch as unconfirmed, with press naming 29 September. The API changelog dates the beta release 6 October 2026, and openai-python added Decisions in 3.26.0 the same day. Whether a limited preview ran earlier wasn't established\n- Which rate limits apply to `/v1/decisions`. The GPT-6 Luna page gives the model's limits and its endpoint table doesn't list Decisions\n- The context limit on this endpoint. The guide gives none, and GPT-6 Luna's page gives 922,000 input tokens for the model\n- Whether the Free tier reaches `/v1/decisions`, and whether it needs a card\n- Whether a restricted project key can be limited to `/v1/decisions`, and how Decisions calls appear in usage views\n- Whether failed or refused questions are charged\n- The Mixpanel breach of November 2025 carries a deduction of 2 on three OpenAI listings and none on the OpenAI API listing. We took none here, since the endpoint is eleven months younger than the incident and the source page answered 403 on 8 October. An editor should settle this across the OpenAI listings\n- openai.com/.well-known/security.txt has no Expires field, which RFC 9116 requires. The provenance block says valid, as the other OpenAI listings do\n- unchecked: the services agreement, privacy policy, DPA and subprocessor list (openai.com answered 403 on 8 October). These stand per the OpenAI API listing's checks of 26 September and 4 October\n- unchecked: the Scale Tier SLA page (403), so no SLA was scored\n- unchecked: the help centre articles on API key permissions and support (help.openai.com answered 403)\n- unchecked: the Playground at platform.openai.com/decisions, which needs a login\n- unchecked: issue reply times on openai-python. The GitHub API refused us for its rate limit\n\n## Weaknesses\n\n- Public beta released on 6 October 2026. The guide expects general availability in the coming weeks and gives no date\n- One model alias, `gpt-6-luna`, with no dated snapshot to pin\n- Images must be inline base64 data URLs. Hosted image URLs and `file_id` inputs aren't accepted\n- Not among the endpoints the Batch guide lists, and no SLA covering it was found\n- No guidance on adversarial or injected content was found in the guide or the reference\n\n## What costs an agent a turn today\n\nThe notes we give agents before they call it. Each one is a workaround an agent shouldn't need.\n\n- Put independent questions about one input in a single `questions` array. Send a separate request when a question depends on an earlier answer\n- Check each answer's `type` before reading it. A question can come back as `refusal` while the others in the same request are answered\n- Encode images as base64 data URLs. Hosted URLs and `file_id` inputs are rejected\n- Give a `choice` question a fallback value such as `other`, and set thresholds from your own labelled examples\n- Use Python SDK 3.26.0 or JavaScript 7.30.0 or later, and follow `Retry-After` on 429 and 503\n\n## When it's done\n\nSend what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `\"kind\": \"dispute\"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.\n",
    "grade": "BB",
    "score": 71.5,
    "assessed": "2026-10-08",
    "run": "October 2026 research run",
    "categories": [
      {
        "key": "reliability",
        "name": "Reliability",
        "score": 53,
        "maxGain": 9.4,
        "reason": "Read as a hosted model API. status.openai.com has a Decisions component under APIs, reading 100 per cent on 8 October (20). The endpoint was released on 6 October, so its record is two days long. The incident feed, read back to 10 July, names Decisions in no incident. The platform under it had API-wide error incidents on 25 July, 17 September and 29 September (5 hours 22 minutes), all before the release, so 8 of 30 for a clean record too short to count, the reading given to Clef on its first day. The GPT-6 Luna page lists 5,000 requests and 2,000,000 tokens a minute on the Build tier, 10,000 and 10,000,000 on Launch and 30,000 and 180,000,000 on Grow, but its endpoint table doesn't list Decisions and the guide gives no limits of its own (10 of 15). The rate-limits guide covers `Retry-After` on 429 and 503, backoff with jitter and SDK retries, and a decision has no side effects, so a retry is safe (15). No SLA covering a beta endpoint was found. The Scale Tier page answered 403 on 8 October (0). Public beta (0).",
        "checklist": [
          "Hosted APIs, MCP servers, models and platforms.",
          "- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.",
          "Local packages, SDKs, frameworks and stdio MCP servers.",
          "- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.",
          "Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-reliability"
      },
      {
        "key": "payments",
        "name": "Payments \u0026 pricing",
        "score": 30,
        "maxGain": 8.8,
        "reason": "No x402, MPP or L402 in the guide, the reference or the pricing page (0). Per-token price published without a login, $0.10 per million input tokens with no output or cache charge (20). The rate-limits guide lists a Free tier capped at $100 a month in allowed countries, but the GPT-6 Luna page's rate-limit table starts at Build, which needs $5 of credit, and the Decisions guide doesn't mention a free allowance, so half, as for the OpenAI API listing (10). A person signs up in a browser and creates a key (0).",
        "checklist": [
          "The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).",
          "- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).",
          "Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.",
          "Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-payments"
      },
      {
        "key": "security",
        "name": "Security \u0026 auth",
        "score": 83,
        "maxGain": 3,
        "reason": "Read for a model API, on the lines used for Jev. Project API keys, with the access-control guide saying a key carries its own permissions that are checked against the caller's project role, service-account keys created through the Administration API with a documented rotation procedure, and workload identity federation with short-lived tokens. We didn't confirm that a restricted key can be limited to `/v1/decisions`, since the help centre article answered 403 (27 of 30). The data-controls guide lists `/v1/decisions` as not used for training, with abuse-monitoring logs kept up to 30 days and Zero Data Retention for eligible customers. An image flagged by the CSAM classifier is kept for review under any setting, and the Private Retention column reads pending confirmation (18 of 20). The endpoint judges text and images that often come from users. Neither the guide nor the reference mentions adversarial or injected content. The safety guide advises red-teaming and limiting input length, and requests accept a `safety_identifier` (6 of 15). Each response reports token usage, and the Admin APIs guide documents audit-log retrieval. We didn't confirm how Decisions calls appear in usage views (12 of 15). security.txt is PGP-signed and names a Bugcrowd bounty and a coordinated disclosure policy, though it has no Expires field, and trust.openai.com names SOC 2 Type 2 and ISO/IEC 27001, 27017, 27018, 27701 and 42001 (20).",
        "checklist": [
          "- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.",
          "Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-security"
      },
      {
        "key": "ergonomics",
        "name": "Agent ergonomics",
        "score": 87,
        "maxGain": 2.1,
        "reason": "Read as an API an agent calls for a decision, the reading used for Jev. Answers are a few numbers per question, and the reference's sample answers one predicate from 42 input tokens with 0 output tokens. Up to 128 inline images a request. The usage object reports cached tokens and the guide says cache reads and writes aren't charged. `/v1/decisions` isn't among the endpoints the Batch guide lists, and the guide states no context limit for the endpoint (20 of 25). The caller sets the whole output shape, with up to 200 independent questions about one input in a call (18 of 20). The error-codes guide gives types, codes and recovery steps for 400, 401, 403, 429, 500 and 503, and a question the model declines comes back as a typed `refusal` answer beside the others. Errors particular to the endpoint aren't documented (16 of 20). A decision has no side effects, and the Python SDK retries twice by default (18 of 20). Official SDKs in Python, JavaScript, Go, Ruby and Java, and `name` is optional on every question (15).",
        "checklist": [
          "- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.",
          "Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-ergonomics"
      },
      {
        "key": "maintenance",
        "name": "Maintenance \u0026 community",
        "score": 77,
        "maxGain": 2,
        "reason": "Read for a closed model, on the lines used for Jev. Released in beta on 6 October 2026 (30). OpenAI publishes minimum notice of 6 months for generally available models and 3 months for specialised variants, and says preview models can go with 2 weeks. Nothing states what a beta endpoint gets (6 of 12). Two days old with one model alias and no dated snapshot, too short to show churn (4 of 8). A dated changelog with several entries a month, including this release. The help centre answered 403 and issue replies weren't sampled (12 of 15). Official SDKs gained Decisions on release day, Python 3.26.0 and JavaScript 7.30.0 on 6 October, with 3.26.1 and 7.30.1 on 8 October, and the guide names Go 3.73.0, Ruby 0.101.0 and Java 4.78.0 (15). The Python SDK states Python 3.10 or later, the JavaScript SDK Node 22 or later, and the Python repository carries CI and breaking-change workflows (10).",
        "checklist": [
          "- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.",
          "Models are read for deprecation notice periods and model churn rather than release counts."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-maintenance"
      },
      {
        "key": "schema",
        "name": "Schema \u0026 documentation",
        "score": 89,
        "maxGain": 1.8,
        "reason": "OpenAI's OpenAPI 3.1 document (version 2.3.0, committed 8 October) has `POST /decisions` with `DecisionRequest` and `DecisionResponse` schemas (25). llms.txt indexes the guide, and the guide and the reference page each have a Markdown twin (10). The guide says which question type fits which job and when to use Structured Outputs or function calling instead. No page of known limits for the model on this endpoint was found (16 of 20). Questions are a union discriminated on `type`, with 1 to 200 questions, 2 to 255 choices and 2 to 10 levels set in the schema and no extra properties allowed (15). The guide has examples in curl, JavaScript, Python, Go, Java and Ruby with sample responses. The OpenAPI operation documents only the 200 response, and errors are covered by the general error-codes guide (12 of 15). A dated changelog carries the release on 6 October. The only model ID is the alias `gpt-6-luna` with no dated snapshot, and the beta has no version of its own (11 of 15).",
        "checklist": [
          "APIs and MCP servers.",
          "- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.",
          "Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-schema"
      },
      {
        "key": "transparency",
        "name": "Transparency \u0026 trust",
        "score": 83,
        "maxGain": 1.5,
        "reason": "Closed service with Apache-2.0 SDKs and an MIT OpenAPI repository. The services agreement answered 403 on 8 October, so its terms stand per the OpenAI API listing's earlier check (15). The data-controls guide, read on 8 October, gives `/v1/decisions` its own row and section covering training, 30-day abuse logs, Zero Data Retention, the 24-hour prompt cache and HIPAA eligibility. One column reads pending confirmation, and the DPA wasn't re-read (403), per the 4 October check on the OpenAI API listing (26 of 30). The deprecations page gives dated notices and minimum notice periods for models. It says nothing about beta endpoints, and the guide gives no date for general availability (14 of 20). The guide lists ten data-residency regions with Decisions in each, and regional processing for it in the United States and Europe. The subprocessor list answered 403 on 8 October, per the 4 October check (16 of 20).",
        "blend": "editorial 71, provenance 94",
        "checklist": [
          "- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).",
          "The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-transparency"
      }
    ],
    "provenance": [
      {
        "label": "Terms of service",
        "value": "published, but our reader couldn't read it",
        "points": 7,
        "max": 10
      },
      {
        "label": "Privacy policy",
        "value": "published, but our reader couldn't read it",
        "points": 7,
        "max": 10
      }
    ],
    "unchecked": [
      "The scout gave the launch as unconfirmed, with press naming 29 September. The API changelog dates the beta release 6 October 2026, and openai-python added Decisions in 3.26.0 the same day. Whether a limited preview ran earlier wasn't established",
      "Which rate limits apply to `/v1/decisions`. The GPT-6 Luna page gives the model's limits and its endpoint table doesn't list Decisions",
      "The context limit on this endpoint. The guide gives none, and GPT-6 Luna's page gives 922,000 input tokens for the model",
      "Whether the Free tier reaches `/v1/decisions`, and whether it needs a card",
      "Whether a restricted project key can be limited to `/v1/decisions`, and how Decisions calls appear in usage views",
      "Whether failed or refused questions are charged",
      "The Mixpanel breach of November 2025 carries a deduction of 2 on three OpenAI listings and none on the OpenAI API listing. We took none here, since the endpoint is eleven months younger than the incident and the source page answered 403 on 8 October. An editor should settle this across the OpenAI listings",
      "openai.com/.well-known/security.txt has no Expires field, which RFC 9116 requires. The provenance block says valid, as the other OpenAI listings do",
      "unchecked: the services agreement, privacy policy, DPA and subprocessor list (openai.com answered 403 on 8 October). These stand per the OpenAI API listing's checks of 26 September and 4 October",
      "unchecked: the Scale Tier SLA page (403), so no SLA was scored",
      "unchecked: the help centre articles on API key permissions and support (help.openai.com answered 403)",
      "unchecked: the Playground at platform.openai.com/decisions, which needs a login",
      "unchecked: issue reply times on openai-python. The GitHub API refused us for its rate limit"
    ],
    "weaknesses": [
      "Public beta released on 6 October 2026. The guide expects general availability in the coming weeks and gives no date",
      "One model alias, `gpt-6-luna`, with no dated snapshot to pin",
      "Images must be inline base64 data URLs. Hosted image URLs and `file_id` inputs aren't accepted",
      "Not among the endpoints the Batch guide lists, and no SLA covering it was found",
      "No guidance on adversarial or injected content was found in the guide or the reference"
    ],
    "agentNotes": [
      "Put independent questions about one input in a single `questions` array. Send a separate request when a question depends on an earlier answer",
      "Check each answer's `type` before reading it. A question can come back as `refusal` while the others in the same request are answered",
      "Encode images as base64 data URLs. Hosted URLs and `file_id` inputs are rejected",
      "Give a `choice` question a fallback value such as `other`, and set thresholds from your own labelled examples",
      "Use Python SDK 3.26.0 or JavaScript 7.30.0 or later, and follow `Retry-After` on 429 and 503"
    ],
    "recheck": "https://www.anchorterminal.com/builders/#disputes"
  },
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-08",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  }
}
