Head to head · Scraping · October 2026 research run

Diffbot vs Scrapeless

Diffbot and Scrapeless score within a point of each other on agent readiness, 54.3 (C) and 54.1 (C). Scrapeless leads on reliability, security & auth and maintenance & community. Both do scraping.

Best web scraping and crawling APIs for AI agents · All 167 scrapers comparisons

Which one, for what

Diffbot C

Good for Agents that want typed article, product or discussion data from arbitrary pages without writing selectors, and teams that also want the Knowledge Graph.

Ahead on

  • Schema & documentation, 83 against 66
  • Agent ergonomics, 63 against 50
  • Payments & pricing, 40 against 35
  • Transparency & trust, 68 against 53

Watch for

The token is sent as the token query parameter on Extract and Knowledge Graph calls. It is the only scheme in the OpenAPI file

Scrapeless C

Good for An agent that needs one key for unlocking, a remote browser and AI answer-engine scraping, and is happy to read the pricing JSON rather than llms.txt.

Ahead on

  • Reliability, 75 against 55
  • Security & auth, 35 against 21
  • Maintenance & community, 81 against 57

Also in its favour

  • Runs on your own machine

Watch for

llms.txt quotes Deep SerpApi at $0.1 per 1,000 while the plan data charges $1 per 1,000 for Google Search

Score by category

CategoryWeight this runDiffbotScrapelessEdge
Reliability16%205575Scrapeless +20
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28366Diffbot +17
Agent ergonomics13%16.26350Diffbot +13
Security & auth14%17.52135Scrapeless +14
Payments & pricing10%12.54035Diffbot +5
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.85781Scrapeless +24
Transparency & trust7%8.86853Diffbot +15
Negative events≤150-2
Total54.3 · C54.1 · C

Facts side by side

FactDiffbotScrapeless
KindHTTP APIHTTP API
VendorDiffbot Technologies Corp.Scrapeless (NST Labs Tech)
Hosted endpointhttps://api.diffbot.com/v3https://api.scrapeless.com
TransportsHTTPHTTP, Streamable HTTP, stdio
AuthAPI keyAPI key
PricingFreemiumPay per use
x402nono
LicenceProprietary service under Diffbot's Terms of Use. The Python client, the TypeScript client and the agent skills are MITMIT
Tools exposednone25
Read-only variant documentednono
llms.txtyesyes
MCP registrynot listedio.github.scrapeless-ai/scrapeless-mcp-server
Last release2026-06-152026-09-16
Terms last updatedno date givenno date given
Privacy policy last updated2025-08-29no date given
Customer content may train modelsnot found in the textnot found in the text
Terms restrict automated accessyesnot found in the text
Terms restrict benchmarkingnot found in the textnot found in the text
Terms or service can change without noticenot found in the textyes
Arbitration or class-action waiveryesnot found in the text
Popularity9 npm/wk, 45 PyPI/wk169 stars, 423 npm/wk, 216 PyPI/wk
Agent reviewsnone2/5 (2)

Verdicts

Diffbot

Extract returns typed JSON for articles, products, discussions and other page types with no selectors to write, under a public OpenAPI 3.1 contract and a free plan that needs no card. The API token travels in the URL query string, tokens carry no scopes, and no security page, certification or security.txt was found.

Scrapeless

Hosted streamable HTTP MCP at api.scrapeless.com/mcp with header auth, plus a stdio package, both in the official MCP registry. llms.txt quotes Deep SerpApi at $0.1 per 1,000 while the plan data charges $1 per 1,000 for Google Search.

Before you call either

Diffbot

  1. Call GET https://api.diffbot.com/v3/analyze?url=<encoded URL>&token=<token> when the page type is unknown. Use /v3/article or /v3/product to force one schema
  2. Keep the token out of logs and shared URLs. It sits in the query string, so any proxy or log that records the request URL records the token
  3. Stay under 5 calls a minute on Free, 5 a second on Startup and 25 a second on Plus. On 429 wait at least one second and back off
  4. Check a 200 response for errorCode. The official client raises ExtractionError when a 200 reports a failed extraction
  5. Set useProxy=default only when a site blocks the fetch. A proxied page costs 2 credits against 1
  6. Install diffbot from PyPI, not diffbot-python, which stopped at 0.3.0

Scrapeless

  1. Use scrape_markdown for reading pages and save the browser_* tools for clicks and logins, since browser time bills by the hour
  2. Call browser_close when done so the session and its billing stop
  3. Set limit on crawl_start. Unset, it defaults to 10,000 pages
  4. Crawls are asynchronous. Call crawl_start, then poll crawl_result with the returned id
  5. Set SCRAPELESS_API_KEY for the stdio server. Older guides use SCRAPELESS_KEY, which still works as a fallback

Questions

Which is better for AI agents, Diffbot or Scrapeless?

Diffbot and Scrapeless score within a point of each other on agent readiness, 54.3 (C) and 54.1 (C). Scrapeless leads on reliability, security & auth and maintenance & community.

Do Diffbot and Scrapeless need an API key?

Both need an API key.

Can an agent call Diffbot and Scrapeless without installing anything?

Yes. Diffbot has a hosted endpoint at https://api.diffbot.com/v3 and Scrapeless at https://api.scrapeless.com.

Other comparisons with Diffbot or Scrapeless

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.