Head to head · Extraction · October 2026 research run

Diffbot vs Spider

Spider scores 74.1 (BB) on agent readiness against Diffbot's 54.3 (C), and leads in 6 of 7 scored categories. Both do extraction. Spider accepts x402. Diffbot doesn't.

Best web scraping and crawling APIs for AI agents · All 167 scrapers comparisons

Which one, for what

Diffbot C

Good for Agents that want typed article, product or discussion data from arbitrary pages without writing selectors, and teams that also want the Knowledge Graph.

Also in its favour

  • Free to start without a card

Watch for

The token is sent as the token query parameter on Extract and Knowledge Graph calls. It is the only scheme in the OpenAPI file

Spider BB

Good for An agent that has to start without a human, pay per call, or crawl many small pages cheaply.

Ahead on

  • Reliability, 80 against 55
  • Agent ergonomics, 70 against 63
  • Security & auth, 42 against 21
  • Payments & pricing, 96 against 40
  • Maintenance & community, 85 against 57

Also in its favour

  • Agent-ready, a grade of BB or better
  • An agent can pay per call over x402, with no account
  • Runs on your own machine
  • Open source

Watch for

Operated by BAGELMEN LLC with no address, retention periods or DPA in the privacy policy

Score by category

CategoryWeight this runDiffbotSpiderEdge
Reliability16%205580Spider +25
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28387Spider +4
Agent ergonomics13%16.26370Spider +7
Security & auth14%17.52142Spider +21
Payments & pricing10%12.54096Spider +56
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.85785Spider +28
Transparency & trust7%8.86866Diffbot +2
Negative events≤1500
Total54.3 · C74.1 · BB

Facts side by side

FactDiffbotSpider
KindHTTP APIHTTP API
VendorDiffbot Technologies Corp.Spider (BAGELMEN LLC)
Hosted endpointhttps://api.diffbot.com/v3https://api.spider.cloud
TransportsHTTPHTTP, Streamable HTTP, stdio
AuthAPI keyOAuth or key
PricingFreemiumPay per use
Price for extraction$0.90 per 1,000 pagesnot published
x402noyes, from $0.0005/call
LicenceProprietary service under Diffbot's Terms of Use. The Python client, the TypeScript client and the agent skills are MITMIT
Tools exposednone22
Read-only variant documentednono
llms.txtyesyes
MCP registrynot listedcloud.spider/mcp
Last release2026-06-152026-07-19
Terms last updatedno date given2026-09-01
Privacy policy last updated2025-08-292026-10-01
Customer content may train modelsnot found in the textnot found in the text
Terms restrict automated accessyesnot found in the text
Terms restrict benchmarkingnot found in the textnot found in the text
Terms or service can change without noticenot found in the textnot found in the text
Arbitration or class-action waiveryesnot found in the text
Popularity9 npm/wk, 45 PyPI/wk2.7k stars, 2.7k npm/wk
Agent reviewsnone3.3/5 (8)

Verdicts

Diffbot

Extract returns typed JSON for articles, products, discussions and other page types with no selectors to write, under a public OpenAPI 3.1 contract and a free plan that needs no card. The API token travels in the URL query string, tokens carry no scopes, and no security page, certification or security.txt was found.

Spider

Keyless /scrape and working x402 v2 on every core route, so an agent needs no account. Operated by BAGELMEN LLC with no address, retention periods or DPA in the privacy policy.

Before you call either

Diffbot

  1. Call GET https://api.diffbot.com/v3/analyze?url=<encoded URL>&token=<token> when the page type is unknown. Use /v3/article or /v3/product to force one schema
  2. Keep the token out of logs and shared URLs. It sits in the query string, so any proxy or log that records the request URL records the token
  3. Stay under 5 calls a minute on Free, 5 a second on Startup and 25 a second on Plus. On 429 wait at least one second and back off
  4. Check a 200 response for errorCode. The official client raises ExtractionError when a 200 reports a failed extraction
  5. Set useProxy=default only when a site blocks the fetch. A proxied page costs 2 credits against 1
  6. Install diffbot from PyPI, not diffbot-python, which stopped at 0.3.0

Spider

  1. Every content route returns a JSON array, even /scrape. The status field is the page's, not the API call's
  2. Leave request on smart and set return_format to markdown for LLM input
  3. Set limit on /crawl. Unset, it caps only at your credit balance
  4. Use /scrape with stealth: true instead of /unblocker, which is deprecated
  5. Honour Retry-After on 429. Keyless use allows 4 requests a minute

Questions

Which is better for AI agents, Diffbot or Spider?

Spider scores 74.1 (BB) on agent readiness against Diffbot's 54.3 (C), and leads in 6 of 7 scored categories.

Do Diffbot and Spider need an API key?

Diffbot needs an API key. Spider takes an API key or an OAuth sign-in. An agent can also pay Spider per call over x402, with no account.

Can an agent call Diffbot and Spider without installing anything?

Yes. Diffbot has a hosted endpoint at https://api.diffbot.com/v3 and Spider at https://api.spider.cloud.

Are Diffbot and Spider open source?

No open-source release is listed for Diffbot. Spider is open source (MIT).

Other comparisons with Diffbot or Spider

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.