Head to head · Agent harness · October 2026 research run

Devin vs Qwen Code

Qwen Code scores 72.4 (BB) on agent readiness against Devin's 55.7 (C), and leads in 5 of 7 scored categories. Devin leads on security & auth and transparency & trust. Both do agent harness.

Which one, for what

Devin C

Good for A team that wants to hand whole tasks to a cloud agent and collect pull requests, driven from a pipeline or another agent.

Ahead on

  • Security & auth, 72 against 53
  • Transparency & trust, 69 against 63

Also in its favour

  • A hosted endpoint, with nothing to install

Watch for

www.devinstatus.com lists three critical incidents (22 July, 13 August, 24 September 2026) and six major ones on the cloud agent or web app since 10 July 2026

Qwen Code BB

Good for Developers and CI jobs that want an open-source harness tied to no single model vendor, with run budgets and SDKs.

Ahead on

  • Reliability, 69 against 30
  • Schema & documentation, 91 against 80
  • Agent ergonomics, 85 against 62
  • Payments & pricing, 60 against 20
  • Maintenance & community, 88 against 63

Also in its favour

  • Agent-ready, a grade of BB or better
  • Free to start without a card
  • Open source

Watch for

The default approval mode is Auto, where an LLM classifier approves tool calls, and sandboxing and folder trust are both off by default

Score by category

CategoryWeight this runDevinQwen CodeEdge
Reliability16%203069Qwen Code +39
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28091Qwen Code +11
Agent ergonomics13%16.26285Qwen Code +23
Security & auth14%17.57253Devin +19
Payments & pricing10%12.52060Qwen Code +40
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86388Qwen Code +25
Transparency & trust7%8.86963Devin +6
Negative events≤1500
Total55.7 · C72.4 · BB

Facts side by side

FactDevinQwen Code
KindAgent harnessAgent harness
VendorCognition AI, Inc.Alibaba (Qwen team)
Hosted endpointhttps://mcp.devin.ai/mcpno (local only)
TransportsHTTP
AuthAPI keyAPI key
PricingFreemiumFree
x402nono
LicenceProprietary service under Cognition's Platform Terms of ServiceApache-2.0
Tools exposed13none
Read-only variant documentedyesno
llms.txtyesyes
Last release2026-10-072026-10-05
Terms last updated2026-06-302026-10-01
Privacy policy last updated2026-03-092026-10-01
Customer content may train modelsyes, with an opt-outnot found in the text
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingyesnot found in the text
Terms or service can change without noticenot found in the textnot found in the text
Arbitration or class-action waiveryesnot found in the text
Popularitynone28k stars, 105k npm/wk

Verdicts

Devin

The v3 API is well specified, with a public OpenAPI 3.1 file covering 239 operations, problem+json errors, cursor pagination and service-user keys tied to roles. The status page records three critical and several major incidents on the cloud agent between 22 July and 24 September 2026, and no rate limit figures or retry guidance were found in the reviewed documentation.

Qwen Code

Apache-2.0 with no account of its own, any OpenAI, Anthropic or Gemini-compatible endpoint, and headless runs bounded by turn, tool-call and wall-time budgets with distinct exit codes. Out of the box, Auto mode lets an LLM classifier approve tool calls, with the sandbox and folder trust both off. Usage statistics are on by default.

Before you call either

Devin

  1. Use a cog_ service-user key with the Member role. Legacy apk_ keys fail against v3 and the MCP server with 401 or 403
  2. Set max_acu_limit on every session you create. Usage is metered by the work done and has no published unit rate
  3. Session creation is not idempotent in v3. After a timeout, list sessions by tag before creating again
  4. Enterprise keys and personal access tokens must send X-Org-Id to the MCP server. Organisation-scoped keys resolve it automatically
  5. Create scheduled work as an automation. POST to the schedules endpoint returns 403 for migrated organisations since 24 September 2026

Qwen Code

  1. Pass --approval-mode on every run. The settings default is auto, and one docs page says headless runs default to asking, so don't rely on either
  2. Add --max-session-turns, --max-wall-time and --max-tool-calls. All three are unlimited by default. Exit 53 is the turn limit and 55 a budget
  3. --yolo doesn't turn on a sandbox. Add --sandbox, or on Linux set tools.executionSandbox with network set to closed
  4. Set QWEN_USAGE_STATISTICS_ENABLED=false to stop usage statistics
  5. Authenticate with OPENAI_API_KEY, OPENAI_BASE_URL and OPENAI_MODEL in CI. The browser sign-in is discontinued

Questions

Which is better for AI agents, Devin or Qwen Code?

Qwen Code scores 72.4 (BB) on agent readiness against Devin's 55.7 (C), and leads in 5 of 7 scored categories. Devin leads on security & auth and transparency & trust.

Are Devin and Qwen Code open source?

No open-source release is listed for Devin. Qwen Code is open source (Apache-2.0).

Other comparisons with Devin or Qwen Code

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.