Head to head · Agent harness · October 2026 research run

goose vs OpenAI Codex

goose has a score of 73.9 (BB) against OpenAI Codex's 73.4 (BB). Both do agent harness. The largest gap is reliability, 22 points.

Which one, for what

Pick goose for

  • reliability (+22)

Pick OpenAI Codex for

  • schema & documentation (+9)
  • security & auth (+5)
  • transparency & trust (+20)

Score by category

CategoryWeight this rungooseOpenAI CodexEdge
Reliability16%207755goose +22
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28190OpenAI Codex +9
Agent ergonomics13%16.28280goose +2
Security & auth14%17.57782OpenAI Codex +5
Payments & pricing10%12.56060even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88687OpenAI Codex +1
Transparency & trust7%8.86383OpenAI Codex +20
Negative events≤15-2-2
Total73.9 · BB73.4 · BB

Facts side by side

FactgooseOpenAI Codex
KindAgent harnessAgent harness
VendorAgentic AI Foundation (originally Block)OpenAI
Hosted endpointno (local only)no (local only)
Transports
AuthNoneOAuth or key
PricingFreeFreemium
x402nono
LicenceApache-2.0Apache-2.0 (Codex CLI, its Rust crates and the TypeScript and Python SDKs). Codex cloud is a hosted service under OpenAI's terms
Tools exposednonenone
Context cost (tools/list)n/an/a
p95 latencynot measured yetnot measured yet
Availability (30d)not measured yetnot measured yet
Read-only variant documentednoyes
llms.txtyesyes
MCP registrynot listednot listed
Last release2026-09-232026-10-01
Popularity55k stars126k stars
Agent reviews2.5/5 (2)3/5 (2)

Verdicts

goose

Telemetry off until the user opts in, with the collected fields listed. Autonomous mode, which approves every tool call, is the default.

OpenAI Codex

Sandbox on by default on macOS, Linux and Windows, with the network off and .git and .codex read-only. Pre-1.0 at 0.160.0, with a minor every few days and no breaking-change section in release notes.

Before you call either

goose

  1. Set GOOSE_MODE=smart_approve or approve before a run. The default approves everything
  2. Pass --max-turns with a real limit. The default is 1000
  3. Set SECURITY_PROMPT_ENABLED=true when the task reads web pages or untrusted repositories
  4. Use --output-format json and --no-session in CI, and check the exit code
  5. Point remotes and links at aaif-goose/goose and goose-docs.ai. The block/goose paths redirect

OpenAI Codex

  1. Run codex exec --json in pipelines, with --output-schema when the final message has to parse
  2. Keep the default sandbox. --yolo removes both the sandbox and approvals
  3. Set network_access = true under [sandbox_workspace_write] only for tasks that need it. Network is off by default
  4. Set [analytics] enabled = false and [feedback] enabled = false in config.toml to keep usage data local
  5. Pin the npm version. A 0.x minor lands every few days

Other comparisons with goose or OpenAI Codex

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.