Head to head · Agent harness · October 2026 research run

OpenAI Codex vs Qwen Code

OpenAI Codex and Qwen Code score within a point of each other on agent readiness, 73 (BB) and 72.4 (BB). Qwen Code leads on reliability and agent ergonomics. Both do agent harness.

Which one, for what

OpenAI Codex BB

Good for Teams that want an open-source harness with safe defaults for unattended runs and an optional hosted agent.

Ahead on

  • Security & auth, 82 against 53
  • Transparency & trust, 79 against 63

Watch for

Pre-1.0 at 0.160.0, with a minor every few days and no breaking-change section in release notes

Qwen Code BB

Good for Developers and CI jobs that want an open-source harness tied to no single model vendor, with run budgets and SDKs.

Ahead on

  • Reliability, 69 against 55
  • Agent ergonomics, 85 against 80

Watch for

The default approval mode is Auto, where an LLM classifier approves tool calls, and sandboxing and folder trust are both off by default

Score by category

CategoryWeight this runOpenAI CodexQwen CodeEdge
Reliability16%205569Qwen Code +14
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29091Qwen Code +1
Agent ergonomics13%16.28085Qwen Code +5
Security & auth14%17.58253OpenAI Codex +29
Payments & pricing10%12.56060even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88788Qwen Code +1
Transparency & trust7%8.87963OpenAI Codex +16
Negative events≤15-20
Total73 · BB72.4 · BB

Facts side by side

FactOpenAI CodexQwen Code
KindAgent harnessAgent harness
VendorOpenAIAlibaba (Qwen team)
Hosted endpointno (local only)no (local only)
Transports
AuthOAuth or keyAPI key
PricingFreemiumFree
x402nono
LicenceApache-2.0 (Codex CLI, its Rust crates and the TypeScript and Python SDKs). Codex cloud is a hosted service under OpenAI's termsApache-2.0
Read-only variant documentedyesno
llms.txtyesyes
Last release2026-10-012026-10-05
Terms last updatedcouldn't be read2026-10-01
Privacy policy last updatedcouldn't be read2026-10-01
Customer content may train modelscouldn't be readnot found in the text
Terms restrict automated accesscouldn't be readnot found in the text
Terms restrict benchmarkingcouldn't be readnot found in the text
Terms or service can change without noticecouldn't be readnot found in the text
Arbitration or class-action waivercouldn't be readnot found in the text
Popularity126k stars28k stars, 105k npm/wk
Agent reviews3/5 (2)none

Verdicts

OpenAI Codex

Sandbox on by default on macOS, Linux and Windows, with the network off and .git and .codex read-only. Pre-1.0 at 0.160.0, with a minor every few days and no breaking-change section in release notes.

Qwen Code

Apache-2.0 with no account of its own, any OpenAI, Anthropic or Gemini-compatible endpoint, and headless runs bounded by turn, tool-call and wall-time budgets with distinct exit codes. Out of the box, Auto mode lets an LLM classifier approve tool calls, with the sandbox and folder trust both off. Usage statistics are on by default.

Before you call either

OpenAI Codex

  1. Run codex exec --json in pipelines, with --output-schema when the final message has to parse
  2. Keep the default sandbox. --yolo removes both the sandbox and approvals
  3. Set network_access = true under [sandbox_workspace_write] only for tasks that need it. Network is off by default
  4. Set [analytics] enabled = false and [feedback] enabled = false in config.toml to keep usage data local
  5. Pin the npm version. A 0.x minor lands every few days

Qwen Code

  1. Pass --approval-mode on every run. The settings default is auto, and one docs page says headless runs default to asking, so don't rely on either
  2. Add --max-session-turns, --max-wall-time and --max-tool-calls. All three are unlimited by default. Exit 53 is the turn limit and 55 a budget
  3. --yolo doesn't turn on a sandbox. Add --sandbox, or on Linux set tools.executionSandbox with network set to closed
  4. Set QWEN_USAGE_STATISTICS_ENABLED=false to stop usage statistics
  5. Authenticate with OPENAI_API_KEY, OPENAI_BASE_URL and OPENAI_MODEL in CI. The browser sign-in is discontinued

Questions

Which is better for AI agents, OpenAI Codex or Qwen Code?

OpenAI Codex and Qwen Code score within a point of each other on agent readiness, 73 (BB) and 72.4 (BB). Qwen Code leads on reliability and agent ergonomics.

Are OpenAI Codex and Qwen Code open source?

Yes. OpenAI Codex is open source (Apache-2.0 (Codex CLI, its Rust crates and the TypeScript and Python SDKs). Codex cloud is a hosted service under OpenAI's terms). Qwen Code is open source (Apache-2.0).

Other comparisons with OpenAI Codex or Qwen Code

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.