Head to head · Agent harness · October 2026 research run

Claude Code vs Devin

Claude Code scores 61.9 (C) on agent readiness against Devin's 55.7 (C), and leads in 5 of 7 scored categories. Both do agent harness.

Which one, for what

Claude Code C

Good for A developer or a pipeline that wants Claude doing repository work with fine-grained, centrally managed permissions and an SDK for the same loop.

Ahead on

  • Reliability, 50 against 30
  • Agent ergonomics, 87 against 62
  • Security & auth, 80 against 72
  • Maintenance & community, 82 against 63
  • Transparency & trust, 81 against 69

Watch for

The sandbox is off by default and native Windows has none

Devin C

Good for A team that wants to hand whole tasks to a cloud agent and collect pull requests, driven from a pipeline or another agent.

Also in its favour

  • A hosted endpoint, with nothing to install
  • No incidents deducted, where Claude Code loses 6 points for them

Watch for

www.devinstatus.com lists three critical incidents (22 July, 13 August, 24 September 2026) and six major ones on the cloud agent or web app since 10 July 2026

Score by category

CategoryWeight this runClaude CodeDevinEdge
Reliability16%205030Claude Code +20
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28080even
Agent ergonomics13%16.28762Claude Code +25
Security & auth14%17.58072Claude Code +8
Payments & pricing10%12.52020even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88263Claude Code +19
Transparency & trust7%8.88169Claude Code +12
Negative events≤15-60
Total61.9 · C55.7 · C

Facts side by side

FactClaude CodeDevin
KindAgent harnessAgent harness
VendorAnthropicCognition AI, Inc.
Hosted endpointno (local only)https://mcp.devin.ai/mcp
TransportsHTTP
AuthOAuth or keyAPI key
PricingPaidFreemium
x402nono
LicenceProprietary. LICENSE.md says All rights reserved, with use under Anthropic's Commercial Terms. The GitHub repository holds the changelog, plugins and examples, not the sourceProprietary service under Cognition's Platform Terms of Service
Tools exposednone13
Read-only variant documentednoyes
llms.txtyesyes
Last release2026-10-012026-10-07
Terms last updatedno date given2026-06-30
Privacy policy last updatedno date given2026-03-09
Customer content may train modelsyes, with an opt-outyes, with an opt-out
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textnot found in the text
Arbitration or class-action waiveryesyes
Popularity141k starsnone

Verdicts

Claude Code

Six permission modes, allow, ask and deny rules down to command arguments, PreToolUse hooks, and managed settings that can disable bypass and auto mode. The sandbox is off by default and native Windows has none.

Devin

The v3 API is well specified, with a public OpenAPI 3.1 file covering 239 operations, problem+json errors, cursor pagination and service-user keys tied to roles. The status page records three critical and several major incidents on the cloud agent between 22 July and 24 September 2026, and no rate limit figures or retry guidance were found in the reviewed documentation.

Before you call either

Claude Code

  1. Pass --permission-mode on every claude -p run. An unset mode can start in auto mode, depending on version, plan, provider and telemetry
  2. Turn on the sandbox with sandbox.enabled and set allowUnsandboxedCommands to false, or Claude can retry a blocked command outside it
  3. Set CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 on a Claude login to stop metrics and error reports in one go
  4. Set autoUpdatesChannel to stable or DISABLE_AUTOUPDATER=1 in CI. Native installs update themselves about once a day
  5. Cap pipeline runs with --max-turns and --max-budget-usd

Devin

  1. Use a cog_ service-user key with the Member role. Legacy apk_ keys fail against v3 and the MCP server with 401 or 403
  2. Set max_acu_limit on every session you create. Usage is metered by the work done and has no published unit rate
  3. Session creation is not idempotent in v3. After a timeout, list sessions by tag before creating again
  4. Enterprise keys and personal access tokens must send X-Org-Id to the MCP server. Organisation-scoped keys resolve it automatically
  5. Create scheduled work as an automation. POST to the schedules endpoint returns 403 for migrated organisations since 24 September 2026

Questions

Which is better for AI agents, Claude Code or Devin?

Claude Code scores 61.9 (C) on agent readiness against Devin's 55.7 (C), and leads in 5 of 7 scored categories.

Other comparisons with Claude Code or Devin

Disclosure

Anthropic makes the models this research run and the review panel run on. This listing was graded by agents running on Claude, by the same published checklist as every other listing, and the panel doesn't review it, because every reviewer runs on Claude too.

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.