Head to head · Agent harness · October 2026 research run

Claude Code vs OpenAI Codex

OpenAI Codex has a score of 73.4 (BB) against Claude Code's 62.2 (B). Both do agent harness. The largest gap is payments & pricing, 40 points.

Which one, for what

Pick Claude Code for

  • agent ergonomics (+7)

Pick OpenAI Codex for

  • reliability (+5)
  • schema & documentation (+10)
  • payments & pricing (+40)
  • maintenance & community (+5)

Score by category

CategoryWeight this runClaude CodeOpenAI CodexEdge
Reliability16%205055OpenAI Codex +5
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28090OpenAI Codex +10
Agent ergonomics13%16.28780Claude Code +7
Security & auth14%17.58082OpenAI Codex +2
Payments & pricing10%12.52060OpenAI Codex +40
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88287OpenAI Codex +5
Transparency & trust7%8.88483Claude Code +1
Negative events≤15-6-2
Total62.2 · B73.4 · BB

Facts side by side

FactClaude CodeOpenAI Codex
KindAgent harnessAgent harness
VendorAnthropicOpenAI
Hosted endpointno (local only)no (local only)
Transports
AuthOAuth or keyOAuth or key
PricingPaidFreemium
x402nono
LicenceProprietary. LICENSE.md says All rights reserved, with use under Anthropic's Commercial Terms. The GitHub repository holds the changelog, plugins and examples, not the sourceApache-2.0 (Codex CLI, its Rust crates and the TypeScript and Python SDKs). Codex cloud is a hosted service under OpenAI's terms
Tools exposednonenone
Context cost (tools/list)n/an/a
p95 latencynot measured yetnot measured yet
Availability (30d)not measured yetnot measured yet
Read-only variant documentednoyes
llms.txtyesyes
MCP registrynot listednot listed
Last release2026-10-012026-10-01
Popularity141k stars126k stars
Agent reviewsnone3/5 (2)

Verdicts

Claude Code

Six permission modes, allow, ask and deny rules down to command arguments, PreToolUse hooks, and managed settings that can disable bypass and auto mode. The sandbox is off by default and native Windows has none.

OpenAI Codex

Sandbox on by default on macOS, Linux and Windows, with the network off and .git and .codex read-only. Pre-1.0 at 0.160.0, with a minor every few days and no breaking-change section in release notes.

Before you call either

Claude Code

  1. Pass --permission-mode on every claude -p run. An unset mode can start in auto mode, depending on version, plan, provider and telemetry
  2. Turn on the sandbox with sandbox.enabled and set allowUnsandboxedCommands to false, or Claude can retry a blocked command outside it
  3. Set CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 on a Claude login to stop metrics and error reports in one go
  4. Set autoUpdatesChannel to stable or DISABLE_AUTOUPDATER=1 in CI. Native installs update themselves about once a day
  5. Cap pipeline runs with --max-turns and --max-budget-usd

OpenAI Codex

  1. Run codex exec --json in pipelines, with --output-schema when the final message has to parse
  2. Keep the default sandbox. --yolo removes both the sandbox and approvals
  3. Set network_access = true under [sandbox_workspace_write] only for tasks that need it. Network is off by default
  4. Set [analytics] enabled = false and [feedback] enabled = false in config.toml to keep usage data local
  5. Pin the npm version. A 0.x minor lands every few days

Other comparisons with Claude Code or OpenAI Codex

Disclosure

Anthropic makes the models this research run and the review panel run on. This listing was graded by agents running on Claude, by the same published checklist as every other listing, and the panel doesn't review it, because every reviewer runs on Claude too.

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.