# Amazon Textract (slim) > Amazon Textract is AWS's document OCR and analysis API. It reads printed and handwritten text from scans and PDFs and returns tables, form fields, layout elements, answers to queries, and invoice, receipt and identity document fields as JSON. - Full: https://www.anchorterminal.com/tools/amazon-textract.md (~9,450 tokens) · this version ~2,130 tokens · JSON https://www.anchorterminal.com/tools/amazon-textract.json · canonical https://www.anchorterminal.com/tools/amazon-textract - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **BB · 73.5/100 · rank #85 of 842 · #1 in Document parsing & extraction · agent-ready · confidence medium** Assessment: IAM policies limit a credential to single operations, CloudTrail logs every call, and page prices start at $1.50 per 1,000 in US East. Multipage files need S3 and an asynchronous job. Text detection covers six languages. Under the AWS Service Terms AWS may store and use documents to improve the service unless an AI services opt-out policy is set. ## Facts - Kind: HTTP API · vendor: Amazon Web Services · category: Document parsing & extraction · legal entity: Amazon Web Services, Inc. · provenance 95/100 - Endpoint: `https://textract.us-east-1.amazonaws.com` (HTTP) - Auth: API key · pricing: Pay per use · x402: no · licence: Proprietary service under the AWS Customer Agreement. The AWS SDKs and the Textractor helper library are Apache-2.0 - Probe metrics: not measured yet (probes haven't run) - API: Amazon Textract API version 2018-06-27, 25 operations, awsJson1_1 protocol (JSON over POST), signed with SigV4. Reached through the AWS CLI (`aws textract`) and the AWS SDKs - Endpoint: `textract..amazonaws.com`, with dual-stack `textract..api.aws` and FIPS endpoints in US and Canada Regions. 16 Regions on the endpoints page, two GovCloud Regions included - Operations: `DetectDocumentText` for text, `AnalyzeDocument` for tables, forms, queries, signatures and layout, `AnalyzeExpense` for invoices and receipts, `AnalyzeID` for US identity documents, and `StartLendingAnalysis` for mortgage packages, which also classifies pages - Price per page: US East (N. Virginia), per 1,000 pages. Text $1.50, tables $15, queries $15, forms $50, layout $4, signatures $3.50, invoices and receipts $10, identity documents $25, lending $70 - Output: JSON blocks (page, line, word, table, cell, key-value set, selection element, query result, layout element) linked by ID. No Markdown or plain-text output - Page references: Each block has a page number, a bounding box, a polygon and a confidence value - Structured extraction: Key-value pairs from forms, up to 15 natural-language queries a page synchronously and 30 asynchronously, and Custom Queries adapters trained on 5 to 2,500 of your documents - Limits: JPEG, PNG, PDF and TIFF. Synchronous calls take 1 page up to 10 MB. Asynchronous jobs take PDF and TIFF up to 500 MB and 3,000 pages from S3 - Languages: English, French, German, Italian, Portuguese and Spanish. Handwriting and queries in English only. No vertical text - Free tier: Three months for new AWS customers. 1,000 pages a month for text detection, 100 a month for forms, tables, layout, queries, expense and identity analysis, 2,000 for lending. None for Custom Queries - Rate limits: Per account and Region. 25 `DetectDocumentText` and 10 `AnalyzeDocument` calls a second in US East (N. Virginia) and US West (Oregon), 1 a second in most other Regions. 600 concurrent asynchronous jobs in those two Regions, 100 elsewhere. Adjustable through Service Quotas - Idempotency: `ClientRequestToken` on the `Start` operations returns the same `JobId` for a repeated call - Data retention: Asynchronous results kept 7 days unless written to your S3 bucket. AWS may store documents to improve the service unless an AI services opt-out policy is set - Audit: CloudTrail logs every operation with the caller's identity. Image bytes and response fields are not logged - Agent tooling: No MCP server for Textract found. The AWS CLI and SDKs are the documented routes, with the Apache-2.0 `amazon-textract-textractor` Python library for reading responses - SLA: Amazon Textract SLA, 99.9 per cent monthly uptime per Region, credits of 10, 25 or 100 per cent - Compliance: In scope for SOC 1, 2 and 3 (page updated 11 August 2026). The guide names HIPAA, ISO and PCI programmes, with an account opt-out needed for PCI DSS data - Prices: Text detection $1.50 per 1,000 pages; Document analysis, tables $15 per 1,000 pages; Document analysis, forms $50 per 1,000 pages; Layout $4 per 1,000 pages; Invoices and receipts $10 per 1,000 pages; Identity documents $25 per 1,000 pages; Lending workflow $70 per 1,000 pages - Scores: Reliability 96, Performance pending, Schema & documentation 82, Agent ergonomics 70, Security & auth 74, Payments & pricing 35, Task success pending, Maintenance & community 58, Transparency & trust 82 · total over the 7 assessed categories - Why: Reliability, AWS Health Dashboard with per-service, per-Region history (20). · Schema & documentation, No OpenAPI, but the Smithy model is published in aws/api-models-aws, 25 operations under API version 2018-06-27 on the awsJson1_1 protocol… · Agent ergonomics, Graded on the Textract API through the AWS CLI and SDKs. · Security & auth, IAM with identity policies per action, temporary credentials from AWS STS and rotation. · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, The Textract API model last changed on 11 August 2026, a documentation update, 58 days before the check, and the JavaScript SDK client shipp… · Transparency & trust, Closed service under the AWS Customer Agreement and section 50 of the Service Terms, with Apache-2.0 SDKs (20 of 30). - Sources: 39, open questions: 10, both in the full twin - Capabilities: docs.parse, docs.ocr, docs.extract, docs.tables - JSON: https://www.anchorterminal.com/api/v1/tools/amazon-textract.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/amazon-textract.svg` or a link to https://www.anchorterminal.com/tools/amazon-textract from a page on aws.amazon.com or one of its subdomains, or the README of github.com/aws/api-models-aws, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Sign with SigV4 for service `textract` at `textract..amazonaws.com`, or call `aws textract detect-document-text`. The AWS CLI can't send image bytes, so reference an S3 object 2. Use `StartDocumentTextDetection` or `StartDocumentAnalysis` for any PDF or TIFF of more than one page, pass a `ClientRequestToken`, then page `Get` calls with `NextToken` (1,000 blocks at most each) 3. Fetch results within 7 days of starting a job, or set `OutputConfig` to write them to your own S3 bucket 4. Request only the `FeatureTypes` you need. Forms cost $50 per 1,000 pages against $15 for tables or queries in US East 5. Back off on `ProvisionedThroughputExceededException` (HTTP 400) and `ThrottlingException` (HTTP 500). Neither carries Retry-After, and default quotas are 1 call a second in most Regions 6. Set the AI services opt-out policy on the AWS organisation before sending customer documents ## Connect ```bash pip install boto3 # or: npm i @aws-sdk/client-textract ``` ```bash aws textract detect-document-text \ --document '{"S3Object":{"Bucket":"my-bucket","Name":"scan.png"}}' \ --region us-east-1 ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/amazon-textract ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Azure Document Intelligence | B | 66 | docs.parse, docs.ocr, docs.extract, docs.tables | https://www.anchorterminal.com/tools/azure-document-intelligence.min.md | | Extend API + MCP | B | 62.9 | docs.parse, docs.ocr, docs.extract, docs.tables | https://www.anchorterminal.com/tools/extend.min.md | | Reducto API + MCP | B | 62.9 | docs.parse, docs.ocr, docs.extract, docs.tables | https://www.anchorterminal.com/tools/reducto.min.md | | LlamaParse API + MCP | C | 59.2 | docs.parse, docs.ocr, docs.extract, docs.tables | https://www.anchorterminal.com/tools/llamaparse.min.md | | Mistral OCR API | C | 58.8 | docs.parse, docs.ocr, docs.extract, docs.tables | https://www.anchorterminal.com/tools/mistral-ocr.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)