# Document parsing, OCR and extraction APIs for AI agents > 10 document parsing & extraction listings ranked by the Anchor benchmark. Leader Reducto API + MCP (B). APIs that turn scans, PDFs, invoices and forms into text, tables and structured fields an agent can use. Compared on text accuracy, table structure, field extraction, page references, processing time and price per page. - Canonical: https://www.anchorterminal.com/categories/document-extraction - Markdown: https://www.anchorterminal.com/categories/document-extraction.md (~2,850 tokens) - Slim: https://www.anchorterminal.com/categories/document-extraction.min.md (~530 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/categories/document-extraction.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-04 APIs that turn scans, PDFs, invoices and forms into text, tables and structured fields an agent can use. Compared on text accuracy, table structure, field extraction, page references, processing time and price per page. - Tools ranked: 10 · agent-ready (BB or better): 0 · accept x402: 0 · hosted endpoints: 10 · desk reviews by the panel: 20 - JSON: https://www.anchorterminal.com/api/v1/tools.json (list) · https://www.anchorterminal.com/api/v1/rankings.json (ranked) · https://www.anchorterminal.com/api/v1/x402.json (payable) · https://www.anchorterminal.com/api/v1/capabilities.json (by capability) - Grades run AA, A, BB, B, C, D, E, F · methodology: https://www.anchorterminal.com/benchmark/ - Capabilities in this category: docs.parse, docs.ocr, docs.extract, docs.tables, docs.chunk - https://letme.dev/docs.parse picks the top-graded tool in this list and says how to call it direct; calling through letme comes later (https://www.anchorterminal.com/letme/index.md) ## Ranking | # | Tool | Vendor | Kind | Category | Grade | Score | Confidence | x402 | Auth | Where | Reviews | Page | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 207 | Reducto API + MCP | Reducto | HTTP API | Documents | B | 63.2 | medium | no | API key | hosted + local | 4/5 (2) | https://www.anchorterminal.com/tools/reducto.md | | 211 | Extend API + MCP | Extend | HTTP API | Documents | B | 62.9 | medium | no | OAuth or key | hosted | 4/5 (2) | https://www.anchorterminal.com/tools/extend.md | | 227 | Mindee API | Mindee | HTTP API | Documents | C | 61.5 | medium | no | API key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/mindee.md | | 262 | LlamaParse API + MCP | LlamaIndex | HTTP API | Documents | C | 59.6 | medium | no | OAuth or key | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/llamaparse.md | | 271 | Mistral OCR API | Mistral AI | Model API | Documents | C | 59 | medium | no | API key | hosted | 4.5/5 (2) | https://www.anchorterminal.com/tools/mistral-ocr.md | | 307 | Adobe PDF Services / PDF Extract API | Adobe | HTTP API | Documents | C | 56 | medium | no | OAuth | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/adobe-pdf-extract.md | | 315 | Veryfi API + MCP | Veryfi | HTTP API | Documents | C | 55.4 | medium | no | API key | hosted + local | 3/5 (2) | https://www.anchorterminal.com/tools/veryfi.md | | 390 | Unstructured API + MCP | Unstructured | HTTP API | Documents | D | 47 | medium | no | OAuth or key | hosted | 2.5/5 (2) | https://www.anchorterminal.com/tools/unstructured.md | | 413 | Nanonets API + MCP | Nanonets | HTTP API | Documents | E | 42.6 | medium | no | OAuth or key | hosted | 2/5 (2) | https://www.anchorterminal.com/tools/nanonets.md | | 430 | PDF.co API + MCP | PDF.co (Artifex Software) | HTTP API | PDF | E | 38.2 | medium | no | API key | hosted + local | 3/5 (2) | https://www.anchorterminal.com/tools/pdf-co.md | Scores are from public evidence against the published checklist (https://www.anchorterminal.com/benchmark/), with Performance and Task success pending. p95 latency and context cost come from our probes, which haven't run yet. ## Summaries ### 207. Reducto API + MCP, B (63.2) Hosted parse, extract, split, classify and edit endpoints for PDFs, scans, spreadsheets and Office files, built on its own r-1 parsing model. Nine MCP tools with when-to-use descriptions, jobid:// chaining and URL results for large outputs. Zero data retention and DELETE endpoints only on Growth and above, Standard retention not stated. - Page: https://www.anchorterminal.com/tools/reducto · Markdown: https://www.anchorterminal.com/tools/reducto.md · JSON: https://www.anchorterminal.com/api/v1/tools/reducto.json - Capabilities: docs.parse, docs.ocr, docs.extract, docs.tables, docs.chunk, docs.classify · endpoint: `https://platform.reducto.ai` ### 211. Extend API + MCP, B (62.9) Hosted parse, extract, classify, split and PDF form-fill APIs with versioned processors, evaluation sets and workflows. Hosted MCP with OAuth scoped to workspaces and to test or production, and a tools filter with nine groups. Parse is billed on top of Extract, Split and Classify, so Performance Extract costs 5 credits a page. - Page: https://www.anchorterminal.com/tools/extend · Markdown: https://www.anchorterminal.com/tools/extend.md · JSON: https://www.anchorterminal.com/api/v1/tools/extend.json - Capabilities: docs.parse, docs.ocr, docs.extract, docs.tables, docs.chunk, docs.classify · endpoint: `https://api.extend.ai` ### 227. Mindee API, C (61.5) Hosted extraction, classification, split, crop and raw-text OCR models that you define in the Mindee platform, then call by model ID through an async REST API. Extracted data kept 12 hours by default (1 to 24 configurable), deletable on fetch, source files never stored. Every call needs a model_id created in the web platform first. - Page: https://www.anchorterminal.com/tools/mindee · Markdown: https://www.anchorterminal.com/tools/mindee.md · JSON: https://www.anchorterminal.com/api/v1/tools/mindee.json - Capabilities: docs.ocr, docs.extract, docs.classify · endpoint: `https://api-v2.mindee.net/v2` ### 262. LlamaParse API + MCP, C (59.6) LlamaIndex's hosted Parse, Extract, Classify, Split and Index APIs on one LlamaCloud key, billed in credits by tier. Free plan with 10,000 credits a month, no card. Breaking SDK changes shipped as minor releases (files.get renamed, classify v1 removed). - Page: https://www.anchorterminal.com/tools/llamaparse · Markdown: https://www.anchorterminal.com/tools/llamaparse.md · JSON: https://www.anchorterminal.com/api/v1/tools/llamaparse.json - Capabilities: docs.parse, docs.ocr, docs.extract, docs.tables, docs.chunk, docs.classify · endpoint: `https://api.cloud.llamaindex.ai/api/v2` ### 271. Mistral OCR API, C (59) Mistral's OCR API for extracting content from documents. Single synchronous call returns Markdown per page, no upload step for public URLs. OCR API at 99.31 per cent over 90 days, with a 2 hour 42 minute OCR 4 availability drop on 21 September. - Page: https://www.anchorterminal.com/tools/mistral-ocr · Markdown: https://www.anchorterminal.com/tools/mistral-ocr.md · JSON: https://www.anchorterminal.com/api/v1/tools/mistral-ocr.json - Capabilities: docs.parse, docs.ocr, docs.extract, docs.tables · endpoint: `https://api.mistral.ai/v1` ### 307. Adobe PDF Services / PDF Extract API, C (56) Adobe's REST API for reading and changing PDFs. Public OpenAPI 3.0.1 spec covering Extract, PDF to Markdown and 20 other operations. No published paid price, paid use goes through sales. - Page: https://www.anchorterminal.com/tools/adobe-pdf-extract · Markdown: https://www.anchorterminal.com/tools/adobe-pdf-extract.md · JSON: https://www.anchorterminal.com/api/v1/tools/adobe-pdf-extract.json - Capabilities: docs.parse, docs.ocr, docs.extract, docs.tables, pdf.convert, pdf.merge, pdf.forms, pdf.generate, pdf.extract · endpoint: `https://pdf-services.adobe.io` ### 315. Veryfi API + MCP, C (55.4) Receipt, invoice, bank statement, cheque and W-2 extraction API priced per document, with fraud checks as add-ons. Typed fields and line items for receipts, invoices, cheques, bank statements and tax forms. API outage on 29 September and degraded, incomplete extraction on 1 October. - Page: https://www.anchorterminal.com/tools/veryfi · Markdown: https://www.anchorterminal.com/tools/veryfi.md · JSON: https://www.anchorterminal.com/api/v1/tools/veryfi.json - Capabilities: docs.ocr, docs.extract, docs.tables, docs.classify · endpoint: `https://api.veryfi.com/api/v8/partner` ### 390. Unstructured API + MCP, D (47) Unstructured Transform is a hosted parse and schema-extract API with an official hosted MCP server that partitions, chunks and embeds files. Apache-2.0 library partitions 45+ file types locally at no cost, tagged 11 times since 3 July. No status page, no published rate limits and no SLA. - Page: https://www.anchorterminal.com/tools/unstructured · Markdown: https://www.anchorterminal.com/tools/unstructured.md · JSON: https://www.anchorterminal.com/api/v1/tools/unstructured.json - Capabilities: docs.parse, docs.ocr, docs.extract, docs.tables, docs.chunk · endpoint: `https://transform.unstructured.io/api/v2` ### 413. Nanonets API + MCP, E (42.6) OCR and field extraction from PDFs, scans and images. Extraction API returns Markdown, CSV or JSON, with named fields or a JSON schema. No public changelog and no dated release in the last 90 days. - Page: https://www.anchorterminal.com/tools/nanonets · Markdown: https://www.anchorterminal.com/tools/nanonets.md · JSON: https://www.anchorterminal.com/api/v1/tools/nanonets.json - Capabilities: docs.parse, docs.ocr, docs.extract, docs.tables, docs.classify · endpoint: `https://extraction-api.nanonets.com/api/v2` ### 430. PDF.co API + MCP, E (38.2) Credit-billed REST API for merging, splitting, converting, filling and securing PDFs, plus PDF to text, CSV, JSON and Excel, template parsing and an AI invoice parser. OpenAPI 3.0.1 document with structured error bodies for 400 to 454, plus llms.txt and Markdown docs. No status page. status.pdf.co redirects to betterstack.com/uptime. - Page: https://www.anchorterminal.com/tools/pdf-co · Markdown: https://www.anchorterminal.com/tools/pdf-co.md · JSON: https://www.anchorterminal.com/api/v1/tools/pdf-co.json - Capabilities: pdf.convert, pdf.merge, pdf.forms, pdf.generate, pdf.extract, docs.ocr, docs.tables, docs.extract · endpoint: `https://api.pdf.co/v1` ## How we test this category The same scans, invoices, tables and long PDFs through every API. We score text accuracy, table structure, field extraction, page references, processing time and price per page. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence. ## Indexed, not reviewed (3) Sorted into this category from public catalogues, with facts and our own checks but no score, grade or rank (https://www.anchorterminal.com/indexed/index.md). | Listing | Kind | What it does | Why it's here | | --- | --- | --- | --- | | [diemdesk.com MCP server](https://www.anchorterminal.com/tools/diemdesk-mcp.md) | MCP server | Convert Office files and PDFs, OCR a scan, or capture a web page — from your assistant. | vendor's own | | [nestr.io MCP server](https://www.anchorterminal.com/tools/nestr-mcp.md) | MCP server | Connect AI to Nestr for Holacracy, Sociocracy, and self-organizing teams. | vendor's own | | [Receipt Extraction](https://www.anchorterminal.com/tools/acjlabs-receipt-extraction.md) | MCP server | AU GST/ABN receipt extraction — assigns entertainment/ITC tax codes per line, not just OCR. | vendor's own |