Head to head · Agent multi agent · October 2026 research run

Paperclip vs Qwen Code

Qwen Code scores 72.4 (BB) on agent readiness against Paperclip's 59 (C), and leads in 4 of 7 scored categories. Paperclip leads on security & auth. Both do agent multi agent.

Which one, for what

Paperclip C

Good for Someone running several coding or operations agents who wants one place for tasks, budgets, approvals and history across harnesses.

Ahead on

  • Security & auth, 64 against 53

Watch for

claude_local defaults dangerouslySkipPermissions to true and codex_local defaults to bypassing approvals and the sandbox, with agents running on the host

Qwen Code BB

Good for Developers and CI jobs that want an open-source harness tied to no single model vendor, with run budgets and SDKs.

Ahead on

  • Schema & documentation, 91 against 81
  • Agent ergonomics, 85 against 70
  • Maintenance & community, 88 against 78

Also in its favour

  • Agent-ready, a grade of BB or better
  • No incidents deducted, where Paperclip loses 10 points for them

Watch for

The default approval mode is Auto, where an LLM classifier approves tool calls, and sandboxing and folder trust are both off by default

Score by category

CategoryWeight this runPaperclipQwen CodeEdge
Reliability16%206669Qwen Code +3
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28191Qwen Code +10
Agent ergonomics13%16.27085Qwen Code +15
Security & auth14%17.56453Paperclip +11
Payments & pricing10%12.56060even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.87888Qwen Code +10
Transparency & trust7%8.86563Paperclip +2
Negative events≤15-100
Total59 · C72.4 · BB

Facts side by side

FactPaperclipQwen Code
KindAgent harnessAgent harness
VendorPaperclip Labs, Inc.Alibaba (Qwen team)
Hosted endpointno (local only)no (local only)
Transports
AuthOAuth or keyAPI key
PricingFreeFree
x402nono
LicenceMITApache-2.0
Read-only variant documentednono
llms.txtyesyes
Last release2026-10-052026-10-05
Terms last updated2026-07-232026-10-01
Privacy policy last updated2026-07-232026-10-01
Customer content may train modelsyes, with an opt-outnot found in the text
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingnot found in the textnot found in the text
Terms or service can change without noticenot found in the textnot found in the text
Arbitration or class-action waiveryesnot found in the text
Popularity99k stars, 60k npm/wk28k stars, 105k npm/wk

Verdicts

Paperclip

Board approvals, budgets with a hard stop and an activity log sit above whichever harnesses do the work, and eleven stable versions shipped in 90 days. The Claude Code and Codex adapters skip permission prompts and the sandbox by default, and twelve security advisories, five of them critical, have been published since April 2026.

Qwen Code

Apache-2.0 with no account of its own, any OpenAI, Anthropic or Gemini-compatible endpoint, and headless runs bounded by turn, tool-call and wall-time budgets with distinct exit codes. Out of the box, Auto mode lets an LLM classifier approve tool calls, with the sandbox and folder trust both off. Usage statistics are on by default.

Before you call either

Paperclip

  1. Set PAPERCLIP_TELEMETRY_DISABLED=1 or DO_NOT_TRACK=1 before the first start. Telemetry is on by default
  2. Set dangerouslySkipPermissions and dangerouslyBypassApprovalsAndSandbox to false on agents that read untrusted input, or run them in a sandbox provider
  3. Install with Node.js 24.11 or newer, and install and sign in to each harness CLI on the host first. Paperclip assumes they are there
  4. Use --bind lan or --bind tailnet at onboarding for anything beyond one machine. The default local_trusted mode treats every request as the board admin
  5. Treat 409 on task checkout as owned by another agent and pick different work. The API docs say not to retry

Qwen Code

  1. Pass --approval-mode on every run. The settings default is auto, and one docs page says headless runs default to asking, so don't rely on either
  2. Add --max-session-turns, --max-wall-time and --max-tool-calls. All three are unlimited by default. Exit 53 is the turn limit and 55 a budget
  3. --yolo doesn't turn on a sandbox. Add --sandbox, or on Linux set tools.executionSandbox with network set to closed
  4. Set QWEN_USAGE_STATISTICS_ENABLED=false to stop usage statistics
  5. Authenticate with OPENAI_API_KEY, OPENAI_BASE_URL and OPENAI_MODEL in CI. The browser sign-in is discontinued

Questions

Which is better for AI agents, Paperclip or Qwen Code?

Qwen Code scores 72.4 (BB) on agent readiness against Paperclip's 59 (C), and leads in 4 of 7 scored categories. Paperclip leads on security & auth.

Are Paperclip and Qwen Code open source?

Yes. Paperclip is open source (MIT). Qwen Code is open source (Apache-2.0).

Other comparisons with Paperclip or Qwen Code

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.