Head to head · Extraction · October 2026 research run

Diffbot vs ScraperAPI

Diffbot scores 54.3 (C) on agent readiness against ScraperAPI's 52.1 (D), and leads in 3 of 7 scored categories. ScraperAPI leads on agent ergonomics and security & auth. Both do extraction.

Best web scraping and crawling APIs for AI agents · All 167 scrapers comparisons

Which one, for what

Diffbot C

Good for Agents that want typed article, product or discussion data from arbitrary pages without writing selectors, and teams that also want the Knowledge Graph.

Ahead on

  • Reliability, 55 against 45
  • Schema & documentation, 83 against 67
  • Transparency & trust, 68 against 54

Watch for

The token is sent as the token query parameter on Extract and Knowledge Graph calls. It is the only scheme in the OpenAPI file

ScraperAPI D

Good for Agents that want a single GET to fetch pages through proxies, with parsed JSON for Amazon, Google, Walmart, eBay and Redfin and a batch API of up to 50,000 URLs.

Ahead on

  • Agent ergonomics, 74 against 63
  • Security & auth, 29 against 21

Also in its favour

  • Runs on your own machine

Watch for

No status page is linked from the home, pricing, support or docs pages, although the pricing page lists a 99.9% uptime guarantee

Score by category

CategoryWeight this runDiffbotScraperAPIEdge
Reliability16%205545Diffbot +10
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28367Diffbot +16
Agent ergonomics13%16.26374ScraperAPI +11
Security & auth14%17.52129ScraperAPI +8
Payments & pricing10%12.54040even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.85761ScraperAPI +4
Transparency & trust7%8.86854Diffbot +14
Negative events≤1500
Total54.3 · C52.1 · D

Facts side by side

FactDiffbotScraperAPI
KindHTTP APIHTTP API
VendorDiffbot Technologies Corp.ScraperAPI (saas.group)
Hosted endpointhttps://api.diffbot.com/v3https://api.scraperapi.com
TransportsHTTPHTTP, Streamable HTTP, stdio
AuthAPI keyAPI key
PricingFreemiumFreemium
Price for extraction$0.90 per 1,000 pagesnot published
x402nono
LicenceProprietary service under Diffbot's Terms of Use. The Python client, the TypeScript client and the agent skills are MITnone
Tools exposednone28
Read-only variant documentednoyes
llms.txtyesyes
Last release2026-06-152026-08-07
Terms last updatedno date givencouldn't be read
Privacy policy last updated2025-08-29no date given
Customer content may train modelsnot found in the textcouldn't be read
Terms restrict automated accessyescouldn't be read
Terms restrict benchmarkingnot found in the textcouldn't be read
Terms or service can change without noticenot found in the textcouldn't be read
Arbitration or class-action waiveryescouldn't be read
Popularity9 npm/wk, 45 PyPI/wk5 stars

Verdicts

Diffbot

Extract returns typed JSON for articles, products, discussions and other page types with no selectors to write, under a public OpenAPI 3.1 contract and a free plan that needs no card. The API token travels in the URL query string, tokens carry no scopes, and no security page, certification or security.txt was found.

ScraperAPI

One GET request with a key and a URL returns a page, and the MCP server's 28 tools carry read-only and destructive hints. No status page, OpenAPI file, security.txt or certification was found, and the documented default puts the API key in the URL query string.

Before you call either

Diffbot

  1. Call GET https://api.diffbot.com/v3/analyze?url=<encoded URL>&token=<token> when the page type is unknown. Use /v3/article or /v3/product to force one schema
  2. Keep the token out of logs and shared URLs. It sits in the query string, so any proxy or log that records the request URL records the token
  3. Stay under 5 calls a minute on Free, 5 a second on Startup and 25 a second on Plus. On 429 wait at least one second and back off
  4. Check a 200 response for errorCode. The official client raises ExtractionError when a 200 reports a failed extraction
  5. Set useProxy=default only when a site blocks the fetch. A proxied page costs 2 credits against 1
  6. Install diffbot from PyPI, not diffbot-python, which stopped at 0.3.0

ScraperAPI

  1. Set the client timeout to 70 seconds. The API retries for that long before it returns 500, and a request cancelled earlier is still charged
  2. Send the key as the x-sapi-api_key header to keep it out of URLs and logs
  3. Pass max_cost on each request, or call /account/urlcost first, because domain and anti-bot surcharges vary from 1 to 75 credits
  4. Use output_format=markdown or text to cut response size, and autoparse=true only on supported domains
  5. On 429 reduce concurrency to the plan limit (5 on Free, 20 on Hobby). The request is rejected, not queued

Questions

Which is better for AI agents, Diffbot or ScraperAPI?

Diffbot scores 54.3 (C) on agent readiness against ScraperAPI's 52.1 (D), and leads in 3 of 7 scored categories. ScraperAPI leads on agent ergonomics and security & auth.

Do Diffbot and ScraperAPI need an API key?

Both need an API key.

Can an agent call Diffbot and ScraperAPI without installing anything?

Yes. Diffbot has a hosted endpoint at https://api.diffbot.com/v3 and ScraperAPI at https://api.scraperapi.com.

Other comparisons with Diffbot or ScraperAPI

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.