{
  "fixes": {
    "slug": "amazon-textract",
    "name": "Amazon Textract",
    "listing": "https://www.anchorterminal.com/tools/amazon-textract",
    "markdown": "# Fix list: Amazon Textract\n\nFrom Anchor Terminal's listing at https://www.anchorterminal.com/tools/amazon-textract, the October 2026 research run, assessed 8 October 2026. Grade BB, 73.5 out of 100.\n\nThis is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.\n\nFor a coding agent working on Amazon Textract: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.\n\n## 1. Payments \u0026 pricing, 35 out of 100, up to 8.1 more on the total\n\nWhy it scored 35: No x402, MPP or L402 (0). Per-page prices by Region published without login, with a price feed the pricing page loads (20). The pricing page gives new AWS customers three months of free pages (1,000 a month for text detection, 100 for forms, tables, layout, queries, expense and identity analysis, 2,000 for lending). The Free Tier FAQ says most new customers don't need a payment method, though AWS may still ask for one, so 15 of 20. A person signs up for the AWS account. IAM can mint keys by API only after that (0).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):\n\nThe published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).\n\n- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).\n\nPayment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.\n\nOpen-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.\n\n## 2. Agent ergonomics, 70 out of 100, up to 4.9 more on the total\n\nWhy it scored 70: Graded on the Textract API through the AWS CLI and SDKs. `FeatureTypes` selects what a call analyses, and queries return a named answer per question. The response is a flat list of blocks for every page, line and word, each with geometry and a confidence value, with no field selection, no way to omit word blocks and no Markdown or text output (12 of 25). `Get` operations page with `MaxResults` (1,000 blocks at most) and `NextToken`, and the two adapter list operations page and filter. No operation lists jobs and no block filter exists (13 of 20). 18 typed errors with HTTP codes and a recovery hint each, though throttling uses 400 and 500, not 429 (16 of 20). `ClientRequestToken` on the `Start` operations and on adapter creation, with `IdempotentParameterMismatchException` when a token is reused with different input. Synchronous calls change nothing but bill each attempt (18 of 20). `DetectDocumentText` needs only `Document`, and SDKs exist in every major language. Any PDF or TIFF of more than one page needs S3, an asynchronous job and polling or an SNS topic, and every call needs SigV4 (11 of 15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):\n\n- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.\n\nModels are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.\n\n## 3. Security \u0026 auth, 74 out of 100, up to 4.6 more on the total\n\nWhy it scored 74: IAM with identity policies per action, temporary credentials from AWS STS and rotation. Resource-level permissions cover adapters only, and Textract has no service-specific condition keys. No secret travels in a URL (30). A policy can allow `textract:DetectDocumentText` alone, and the service reads documents without changing them, writing to S3 only when `OutputConfig` names a bucket. The guide names `AmazonTextractFullAccess` and we found no read-only managed policy in it. The only destructive calls delete adapters (16 of 20). Returns text from untrusted documents, with no prompt-injection guidance found (0 of 15). CloudTrail logs every operation with the caller's identity, leaving out image bytes and response fields (15). In scope for SOC 1, 2 and 3 (page updated 11 August 2026), and the guide names HIPAA, ISO and PCI programmes. Vulnerability disclosure programme on HackerOne. aws.amazon.com's security.txt expired on 24 September 2026 and we found no paid bug bounty (18 of 20). We take 5 off the total of 79, a departure from the checklist, because section 50.3 of the AWS Service Terms lets AWS store and use processed documents to improve the service by default, and the opt-out is a policy on the whole AWS organisation.\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-security):\n\n- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.\n\nModels are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.\n\n## 4. Maintenance \u0026 community, 58 out of 100, up to 3.7 more on the total\n\nWhy it scored 58: The Textract API model last changed on 11 August 2026, a documentation update, 58 days before the check, and the JavaScript SDK client shipped 3.1148.0 on 8 October (20 of 30). That is the only model change since 10 July. The client had 64 versions in that period, published with the whole SDK. The guide's document history has no entry after 21 April 2022 and the newest Textract item in AWS's What's New feed is 30 June 2025 (5 of 20). Closed service with AWS re:Post and paid support plans, and a document history that is four years stale (8 of 15). Official SDKs released on most working days (15). The JavaScript client needs Node 20 or later and comes from the SDK's release pipeline. The `amazon-textract-textractor` helper library released 1.10.0 on 11 August 2026 (10).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):\n\n- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.\n\nModels are read for deprecation notice periods and model churn rather than release counts.\n\n## 5. Schema \u0026 documentation, 82 out of 100, up to 2.9 more on the total\n\nWhy it scored 82: No OpenAPI, but the Smithy model is published in aws/api-models-aws, 25 operations under API version 2018-06-27 on the awsJson1_1 protocol (25). llms.txt for the developer guide and the API reference, a Markdown twin of every page and the whole guide as one Markdown file (10). Operation descriptions state purpose and the guide has a section on picking an operation for a use case and on synchronous against asynchronous calls. Few say when not to use them (15 of 20). 11 enumerated types (such as `FeatureType` and `BlockType`), 39 length, pattern or range constraints and 45 required members, with one map and no JSON carried inside strings (13 of 15). The reference pages list each operation's errors with HTTP status codes but carry no examples, and the model has none. The guide has request and response samples and SDK code (11 of 15). The API is versioned, but the guide's document history stops at 21 April 2022 and omits signatures, the lending workflow, layout and Custom Queries, all launched later. Model changes show only as commits in aws/api-models-aws (8 of 15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):\n\nAPIs and MCP servers.\n\n- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.\n\nModels are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.\n\n## 6. Transparency \u0026 trust, 82 out of 100, up to 1.6 more on the total\n\nMade of editorial 69, provenance 95.\n\nWhy it scored 82: Closed service under the AWS Customer Agreement and section 50 of the Service Terms, with Apache-2.0 SDKs (20 of 30). The guide says asynchronous results are kept encrypted for 7 days in a bucket Textract owns unless `OutputConfig` names yours, and that adapter training content is deleted when training ends. Section 50.3 of the Service Terms lets AWS store and use processed content to improve the service, in other Regions too, with an opt-out through an AI services opt-out policy that also deletes stored content. The Textract guide's data protection page does not mention this, and we found no retention statement for synchronous calls (17 of 30). The Customer Agreement commits to 12 months' notice before material functionality of a generally available service is discontinued. No dated Textract deprecation notices found (15 of 20). The sub-processor page, last updated 28 July 2026, lists Textract under service improvement with processing in Australia, Germany, Ireland, Singapore, the United Kingdom and the USA, and promises 30 days' notice of new sub-processors. Region is chosen per request across 16 Regions (17 of 20).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):\n\n- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).\n\nThe other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.\n\nProvenance checks not met in full (half of this category, computed from checked facts):\n\n- security.txt: published but past its Expires date (5 of 10)\n\n## 7. Reliability, 96 out of 100, up to 0.8 more on the total\n\nWhy it scored 96: AWS Health Dashboard with per-service, per-Region history (20). AWS's dashboard history file, read on 8 October 2026, lists eleven events since 10 July and none names Amazon Textract. The service is not sold in the Middle East Regions whose disruptions have been open since March. The one event that names it is the US East (N. Virginia) multi-service outage of 20 October 2025, outside the 90-day window. The public dashboard shows broad events only (30). Default quotas are published per operation and Region, such as 25 `DetectDocumentText` calls a second in US East (N. Virginia), 1 in most other Regions, and 600 concurrent asynchronous jobs (15). The guide recommends five SDK retries with exponential backoff and jitter, and `ClientRequestToken` makes a repeated `Start` call return the same `JobId`. Throttling arrives as `ProvisionedThroughputExceededException` with HTTP 400 or `ThrottlingException` with HTTP 500, neither with Retry-After, and the API model marks no error as retryable (11 of 15). The Amazon Textract SLA, last updated 5 May 2022, commits 99.9 per cent monthly uptime per Region with credits of 10, 25 or 100 per cent (10). Generally available since 29 May 2019 (10).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):\n\nHosted APIs, MCP servers, models and platforms.\n\n- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.\n\nLocal packages, SDKs, frameworks and stdio MCP servers.\n\n- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.\n\nProtocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.\n\n## What we couldn't check\n\nWhat we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.\n\n- unchecked: accuracy, processing time and table quality. We ran no documents through the API.\n- unchecked: the AWS Data Processing Addendum and the list of services the AI services opt-out policy covers. We read the Service Terms and the opt-out overview page only.\n- unchecked: whether an AWS managed read-only policy for Textract exists. The guide names `AmazonTextractFullAccess` only, and we did not read the managed policy reference.\n- unchecked: whether failed or rejected pages are billed. The pricing page read does not say.\n- unchecked: the remaining API reference pages. We read two operation pages and the common errors page, and took the rest from the Smithy model.\n- No retention statement for documents sent to synchronous operations was found in the pages read.\n- No MCP server for Textract was found. `awslabs.amazon-textract-mcp-server` returns 404 on PyPI and in the awslabs/mcp repository.\n- The PyPI download figure is for `amazon-textract-textractor`, a helper library in the aws-samples organisation on GitHub. `boto3` covers every AWS service, so no Textract figure exists for it.\n- The lead called the interface a REST API. The model declares the awsJson1_1 protocol, JSON over POST with an `X-Amz-Target` header, signed with SigV4.\n- The Machine Learning Language SLA does not name Textract. The service has its own SLA at aws.amazon.com/textract/sla/.\n\n## Weaknesses\n\n- AWS may store and use processed documents to improve the service, in other Regions too, unless an AI services opt-out policy is set on the AWS organisation\n- Synchronous calls take one page of at most 10 MB. Multipage PDF and TIFF files must sit in S3 and run as asynchronous jobs\n- Text detection covers English, French, German, Italian, Portuguese and Spanish only. Handwriting and queries are English only, and vertical text is not read\n- Output is a flat list of JSON blocks linked by ID. No Markdown or plain-text output and no way to leave out word blocks or geometry\n- The guide's document history stops at 21 April 2022, and the newest Textract entry in AWS's What's New feed is dated 30 June 2025\n\n## What costs an agent a turn today\n\nThe notes we give agents before they call it. Each one is a workaround an agent shouldn't need.\n\n- Sign with SigV4 for service `textract` at `textract.\u003cregion\u003e.amazonaws.com`, or call `aws textract detect-document-text`. The AWS CLI can't send image bytes, so reference an S3 object\n- Use `StartDocumentTextDetection` or `StartDocumentAnalysis` for any PDF or TIFF of more than one page, pass a `ClientRequestToken`, then page `Get` calls with `NextToken` (1,000 blocks at most each)\n- Fetch results within 7 days of starting a job, or set `OutputConfig` to write them to your own S3 bucket\n- Request only the `FeatureTypes` you need. Forms cost $50 per 1,000 pages against $15 for tables or queries in US East\n- Back off on `ProvisionedThroughputExceededException` (HTTP 400) and `ThrottlingException` (HTTP 500). Neither carries Retry-After, and default quotas are 1 call a second in most Regions\n- Set the AI services opt-out policy on the AWS organisation before sending customer documents\n\n## When it's done\n\nSend what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `\"kind\": \"dispute\"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.\n",
    "grade": "BB",
    "score": 73.5,
    "assessed": "2026-10-08",
    "run": "October 2026 research run",
    "categories": [
      {
        "key": "payments",
        "name": "Payments \u0026 pricing",
        "score": 35,
        "maxGain": 8.1,
        "reason": "No x402, MPP or L402 (0). Per-page prices by Region published without login, with a price feed the pricing page loads (20). The pricing page gives new AWS customers three months of free pages (1,000 a month for text detection, 100 for forms, tables, layout, queries, expense and identity analysis, 2,000 for lending). The Free Tier FAQ says most new customers don't need a payment method, though AWS may still ask for one, so 15 of 20. A person signs up for the AWS account. IAM can mint keys by API only after that (0).",
        "checklist": [
          "The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).",
          "- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).",
          "Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.",
          "Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-payments"
      },
      {
        "key": "ergonomics",
        "name": "Agent ergonomics",
        "score": 70,
        "maxGain": 4.9,
        "reason": "Graded on the Textract API through the AWS CLI and SDKs. `FeatureTypes` selects what a call analyses, and queries return a named answer per question. The response is a flat list of blocks for every page, line and word, each with geometry and a confidence value, with no field selection, no way to omit word blocks and no Markdown or text output (12 of 25). `Get` operations page with `MaxResults` (1,000 blocks at most) and `NextToken`, and the two adapter list operations page and filter. No operation lists jobs and no block filter exists (13 of 20). 18 typed errors with HTTP codes and a recovery hint each, though throttling uses 400 and 500, not 429 (16 of 20). `ClientRequestToken` on the `Start` operations and on adapter creation, with `IdempotentParameterMismatchException` when a token is reused with different input. Synchronous calls change nothing but bill each attempt (18 of 20). `DetectDocumentText` needs only `Document`, and SDKs exist in every major language. Any PDF or TIFF of more than one page needs S3, an asynchronous job and polling or an SNS topic, and every call needs SigV4 (11 of 15).",
        "checklist": [
          "- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.",
          "Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-ergonomics"
      },
      {
        "key": "security",
        "name": "Security \u0026 auth",
        "score": 74,
        "maxGain": 4.6,
        "reason": "IAM with identity policies per action, temporary credentials from AWS STS and rotation. Resource-level permissions cover adapters only, and Textract has no service-specific condition keys. No secret travels in a URL (30). A policy can allow `textract:DetectDocumentText` alone, and the service reads documents without changing them, writing to S3 only when `OutputConfig` names a bucket. The guide names `AmazonTextractFullAccess` and we found no read-only managed policy in it. The only destructive calls delete adapters (16 of 20). Returns text from untrusted documents, with no prompt-injection guidance found (0 of 15). CloudTrail logs every operation with the caller's identity, leaving out image bytes and response fields (15). In scope for SOC 1, 2 and 3 (page updated 11 August 2026), and the guide names HIPAA, ISO and PCI programmes. Vulnerability disclosure programme on HackerOne. aws.amazon.com's security.txt expired on 24 September 2026 and we found no paid bug bounty (18 of 20). We take 5 off the total of 79, a departure from the checklist, because section 50.3 of the AWS Service Terms lets AWS store and use processed documents to improve the service by default, and the opt-out is a policy on the whole AWS organisation.",
        "checklist": [
          "- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.",
          "Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-security"
      },
      {
        "key": "maintenance",
        "name": "Maintenance \u0026 community",
        "score": 58,
        "maxGain": 3.7,
        "reason": "The Textract API model last changed on 11 August 2026, a documentation update, 58 days before the check, and the JavaScript SDK client shipped 3.1148.0 on 8 October (20 of 30). That is the only model change since 10 July. The client had 64 versions in that period, published with the whole SDK. The guide's document history has no entry after 21 April 2022 and the newest Textract item in AWS's What's New feed is 30 June 2025 (5 of 20). Closed service with AWS re:Post and paid support plans, and a document history that is four years stale (8 of 15). Official SDKs released on most working days (15). The JavaScript client needs Node 20 or later and comes from the SDK's release pipeline. The `amazon-textract-textractor` helper library released 1.10.0 on 11 August 2026 (10).",
        "checklist": [
          "- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.",
          "Models are read for deprecation notice periods and model churn rather than release counts."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-maintenance"
      },
      {
        "key": "schema",
        "name": "Schema \u0026 documentation",
        "score": 82,
        "maxGain": 2.9,
        "reason": "No OpenAPI, but the Smithy model is published in aws/api-models-aws, 25 operations under API version 2018-06-27 on the awsJson1_1 protocol (25). llms.txt for the developer guide and the API reference, a Markdown twin of every page and the whole guide as one Markdown file (10). Operation descriptions state purpose and the guide has a section on picking an operation for a use case and on synchronous against asynchronous calls. Few say when not to use them (15 of 20). 11 enumerated types (such as `FeatureType` and `BlockType`), 39 length, pattern or range constraints and 45 required members, with one map and no JSON carried inside strings (13 of 15). The reference pages list each operation's errors with HTTP status codes but carry no examples, and the model has none. The guide has request and response samples and SDK code (11 of 15). The API is versioned, but the guide's document history stops at 21 April 2022 and omits signatures, the lending workflow, layout and Custom Queries, all launched later. Model changes show only as commits in aws/api-models-aws (8 of 15).",
        "checklist": [
          "APIs and MCP servers.",
          "- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.",
          "Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-schema"
      },
      {
        "key": "transparency",
        "name": "Transparency \u0026 trust",
        "score": 82,
        "maxGain": 1.6,
        "reason": "Closed service under the AWS Customer Agreement and section 50 of the Service Terms, with Apache-2.0 SDKs (20 of 30). The guide says asynchronous results are kept encrypted for 7 days in a bucket Textract owns unless `OutputConfig` names yours, and that adapter training content is deleted when training ends. Section 50.3 of the Service Terms lets AWS store and use processed content to improve the service, in other Regions too, with an opt-out through an AI services opt-out policy that also deletes stored content. The Textract guide's data protection page does not mention this, and we found no retention statement for synchronous calls (17 of 30). The Customer Agreement commits to 12 months' notice before material functionality of a generally available service is discontinued. No dated Textract deprecation notices found (15 of 20). The sub-processor page, last updated 28 July 2026, lists Textract under service improvement with processing in Australia, Germany, Ireland, Singapore, the United Kingdom and the USA, and promises 30 days' notice of new sub-processors. Region is chosen per request across 16 Regions (17 of 20).",
        "blend": "editorial 69, provenance 95",
        "checklist": [
          "- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).",
          "The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-transparency"
      },
      {
        "key": "reliability",
        "name": "Reliability",
        "score": 96,
        "maxGain": 0.8,
        "reason": "AWS Health Dashboard with per-service, per-Region history (20). AWS's dashboard history file, read on 8 October 2026, lists eleven events since 10 July and none names Amazon Textract. The service is not sold in the Middle East Regions whose disruptions have been open since March. The one event that names it is the US East (N. Virginia) multi-service outage of 20 October 2025, outside the 90-day window. The public dashboard shows broad events only (30). Default quotas are published per operation and Region, such as 25 `DetectDocumentText` calls a second in US East (N. Virginia), 1 in most other Regions, and 600 concurrent asynchronous jobs (15). The guide recommends five SDK retries with exponential backoff and jitter, and `ClientRequestToken` makes a repeated `Start` call return the same `JobId`. Throttling arrives as `ProvisionedThroughputExceededException` with HTTP 400 or `ThrottlingException` with HTTP 500, neither with Retry-After, and the API model marks no error as retryable (11 of 15). The Amazon Textract SLA, last updated 5 May 2022, commits 99.9 per cent monthly uptime per Region with credits of 10, 25 or 100 per cent (10). Generally available since 29 May 2019 (10).",
        "checklist": [
          "Hosted APIs, MCP servers, models and platforms.",
          "- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.",
          "Local packages, SDKs, frameworks and stdio MCP servers.",
          "- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.",
          "Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-reliability"
      }
    ],
    "provenance": [
      {
        "label": "security.txt",
        "value": "published but past its Expires date",
        "points": 5,
        "max": 10
      }
    ],
    "unchecked": [
      "unchecked: accuracy, processing time and table quality. We ran no documents through the API.",
      "unchecked: the AWS Data Processing Addendum and the list of services the AI services opt-out policy covers. We read the Service Terms and the opt-out overview page only.",
      "unchecked: whether an AWS managed read-only policy for Textract exists. The guide names `AmazonTextractFullAccess` only, and we did not read the managed policy reference.",
      "unchecked: whether failed or rejected pages are billed. The pricing page read does not say.",
      "unchecked: the remaining API reference pages. We read two operation pages and the common errors page, and took the rest from the Smithy model.",
      "No retention statement for documents sent to synchronous operations was found in the pages read.",
      "No MCP server for Textract was found. `awslabs.amazon-textract-mcp-server` returns 404 on PyPI and in the awslabs/mcp repository.",
      "The PyPI download figure is for `amazon-textract-textractor`, a helper library in the aws-samples organisation on GitHub. `boto3` covers every AWS service, so no Textract figure exists for it.",
      "The lead called the interface a REST API. The model declares the awsJson1_1 protocol, JSON over POST with an `X-Amz-Target` header, signed with SigV4.",
      "The Machine Learning Language SLA does not name Textract. The service has its own SLA at aws.amazon.com/textract/sla/."
    ],
    "weaknesses": [
      "AWS may store and use processed documents to improve the service, in other Regions too, unless an AI services opt-out policy is set on the AWS organisation",
      "Synchronous calls take one page of at most 10 MB. Multipage PDF and TIFF files must sit in S3 and run as asynchronous jobs",
      "Text detection covers English, French, German, Italian, Portuguese and Spanish only. Handwriting and queries are English only, and vertical text is not read",
      "Output is a flat list of JSON blocks linked by ID. No Markdown or plain-text output and no way to leave out word blocks or geometry",
      "The guide's document history stops at 21 April 2022, and the newest Textract entry in AWS's What's New feed is dated 30 June 2025"
    ],
    "agentNotes": [
      "Sign with SigV4 for service `textract` at `textract.\u003cregion\u003e.amazonaws.com`, or call `aws textract detect-document-text`. The AWS CLI can't send image bytes, so reference an S3 object",
      "Use `StartDocumentTextDetection` or `StartDocumentAnalysis` for any PDF or TIFF of more than one page, pass a `ClientRequestToken`, then page `Get` calls with `NextToken` (1,000 blocks at most each)",
      "Fetch results within 7 days of starting a job, or set `OutputConfig` to write them to your own S3 bucket",
      "Request only the `FeatureTypes` you need. Forms cost $50 per 1,000 pages against $15 for tables or queries in US East",
      "Back off on `ProvisionedThroughputExceededException` (HTTP 400) and `ThrottlingException` (HTTP 500). Neither carries Retry-After, and default quotas are 1 call a second in most Regions",
      "Set the AI services opt-out policy on the AWS organisation before sending customer documents"
    ],
    "recheck": "https://www.anchorterminal.com/builders/#disputes"
  },
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-09",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  }
}
