# Diffbot (slim) > Diffbot is a hosted set of web data APIs from Diffbot Technologies Corp. of Menlo Park, California. Its Extract API classifies a page and returns typed JSON, alongside Crawl, Web Search, Natural Language and Knowledge Graph APIs. - Full: https://www.anchorterminal.com/tools/diffbot.md (~8,400 tokens) · this version ~1,930 tokens · JSON https://www.anchorterminal.com/tools/diffbot.json · canonical https://www.anchorterminal.com/tools/diffbot - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-10 **C · 54.3/100 · rank #687 of 950 · #12 in Web scraping & crawling · not agent-ready · confidence medium** Assessment: Extract returns typed JSON for articles, products, discussions and other page types with no selectors to write, under a public OpenAPI 3.1 contract and a free plan that needs no card. The API token travels in the URL query string, tokens carry no scopes, and no security page, certification or security.txt was found. ## Facts - Kind: HTTP API · vendor: Diffbot Technologies Corp. · category: Web scraping & crawling · legal entity: Diffbot Technologies Corp. · provenance 86/100 - Endpoint: `https://api.diffbot.com/v3` (HTTP) - Auth: API key · pricing: Freemium · x402: no · licence: Proprietary service under Diffbot's Terms of Use. The Python client, the TypeScript client and the agent skills are MIT - Probe metrics: not measured yet (probes haven't run) - Extract endpoints: `GET https://api.diffbot.com/v3/analyze` classifies a page and extracts it by type. `/article`, `/product`, `/discussion`, `/job`, `/image`, `/video`, `/list` and `/event` force one schema, and `/{api}` runs a custom rule set. Raw HTML or text can be sent by POST - Other APIs: Crawl and Bulk Extract (Plus and above), Knowledge Graph search in DQL and Enhance at `kg.diffbot.com`, Natural Language at `nl.diffbot.com`, Web Search at `llm.diffbot.com` - Plans: Free $0 with 10,000 credits a month. Startup $299 with 250,000 credits. Plus $899 with 1,000,000 credits, Crawl with 25 active jobs and 3 seats. Enterprise by sales. Monthly, cancel at any time - Credits: Extract a page 1, through the data centre proxy 2. Crawl spidering 0. Natural Language document up to 10,000 characters 1. Knowledge Graph entity export or Enhance 25, facet record or Enhance with refresh 100. Web search 1 - Rate limits: Free 5 calls a minute, Startup 5 a second, Plus 25 a second. Free also caps Extract at 10,000 calls a month, Natural Language at 500 and DQL and Enhance at 400 entities each - Errors: JSON with `errorCode` and `error`. Docs pages cover 401, 404, 429, 457 and two 500 cases. The 429 page says to pause one second and back off. The OpenAPI file documents only 200 and 500 - Request options: `mode`, `fallback`, `fields` (links, extlinks, meta, querystring, breadcrumb), `discussion`, `timeout` (default 30,000 ms), `proxy`, `proxyAuth`, `useProxy`, `callback` - Clients: Python `diffbot` 3.0.1 in the repository (tag v3.0.0 of 28 September 2026, Python 3.9+, sync and async, a `db` CLI) and TypeScript `@diffbot/typescript` 0.2.0 (Node 18+), both MIT - Agent skills: Ten skills in `diffbot/diffbot-skills` (v1.1.1 of 13 August 2026, MIT) for Claude Code, GitHub Copilot, Snowflake Cortex Code, Factory Droid and others. They call the `db` CLI and read the token from `~/.diffbot/credentials` - Status: status.diffbot.com, a Pingdom public report with 13 checks and seven days of daily uptime. No incident notes - Certifications: None found on the pages read. No security page is linked from the home page or footer - SLA: None found on the pricing page or in the terms. The terms disclaim uninterrupted service - Prices: Extracted page, Startup overage $1 per 1,000 pages; Extracted page through the data centre proxy, Startup overage $2 per 1,000 pages; Extracted page, Plus overage $0.90 per 1,000 pages; Startup plan $299 per month (plan); Plus plan $899 per month (plan) - Scores: Reliability 55, Performance pending, Schema & documentation 83, Agent ergonomics 63, Security & auth 21, Payments & pricing 40, Task success pending, Maintenance & community 57, Transparency & trust 68 · total over the 7 assessed categories - Why: Reliability, Graded as a hosted API. · Schema & documentation, The Extract APIs have a public OpenAPI 3.1.0 file, version 1.1.0, with 13 operations, and the reference pages are generated from it. · Agent ergonomics, Graded on the REST API, with the skills noted. · Security & auth, One token per account, sent in the URL query string as the documented and only scheme in the OpenAPI file. · Payments & pricing, No x402, MPP or L402 was found (0 of 40). · Maintenance & community, The official Python client was tagged v3.0.0 on 28 September 2026, eleven days before the check. · Transparency & trust, A closed service under undated Terms of Use that cover the Site and the Service, cap liability at $10 and grant Diffbot a perpetual licence… - Sources: 27, open questions: 9, both in the full twin - Capabilities: web.extract, web.scrape, web.crawl, scraping.proxies, web.search, data.knowledge, data.company - JSON: https://www.anchorterminal.com/api/v1/tools/diffbot.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/diffbot.svg` or a link to https://www.anchorterminal.com/tools/diffbot from a page on diffbot.com or one of its subdomains, or the README of github.com/diffbot/diffbot-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Call `GET https://api.diffbot.com/v3/analyze?url=&token=` when the page type is unknown. Use `/v3/article` or `/v3/product` to force one schema 2. Keep the token out of logs and shared URLs. It sits in the query string, so any proxy or log that records the request URL records the token 3. Stay under 5 calls a minute on Free, 5 a second on Startup and 25 a second on Plus. On 429 wait at least one second and back off 4. Check a 200 response for `errorCode`. The official client raises `ExtractionError` when a 200 reports a failed extraction 5. Set `useProxy=default` only when a site blocks the fetch. A proxied page costs 2 credits against 1 6. Install `diffbot` from PyPI, not `diffbot-python`, which stopped at 0.3.0 ## Connect ```bash python3 -m pip install diffbot ``` ```bash curl --request GET \ --url 'https://api.diffbot.com/v3/analyze?token=&url=https%3A%2F%2Fwww.diffbot.com%2Fdocs%2Fintroduction' \ --header 'accept: application/json' ``` ```bash /plugin marketplace add diffbot/diffbot-skills /plugin install diffbot@diffbot-skills ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/diffbot ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Spider | BB | 74.1 | web.scrape, web.extract, web.crawl, web.search, scraping.proxies | https://www.anchorterminal.com/tools/spider-cloud.min.md | | ScrapeGraphAI | B | 68.3 | web.scrape, web.extract, web.crawl, web.search, scraping.proxies | https://www.anchorterminal.com/tools/scrapegraphai.min.md | | Hyperbrowser | B | 66.9 | web.scrape, web.crawl, web.extract, web.search, scraping.proxies | https://www.anchorterminal.com/tools/hyperbrowser.min.md | | Olostep | B | 63 | web.scrape, web.extract, web.crawl, web.search, scraping.proxies | https://www.anchorterminal.com/tools/olostep.min.md | | Browserless | BB | 70.4 | scraping.proxies, web.scrape, web.crawl, web.search | https://www.anchorterminal.com/tools/browserless.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)