Codex vs Claude Code: the decision is reversible

By Marco Kohns, Co-founder of ENLIX, lecturer in AI and growth

· 12 min read

Codex and Claude Code get compared as though you were signing up for the next two years. That is the most expensive mistake in this question. The demanding part of your setup is not the tool, it is the written-down project knowledge, and with both of them that sits in the repository as ordinary Markdown.

The short answer: take the tool that comes with the subscription you already pay for, and treat a switch as an option rather than a defeat. What genuinely separates them afterwards is three things: how much you approve per run, what the entry costs, and whether you have a counter-check that can judge a delegated result without you.

What is the difference between Codex and Claude Code?

Both are agents that read a codebase, change files and run commands. They differ in the vendor, in the subscription that unlocks them, and in the default for how much they may do without asking. In the kind of task they accept, barely at all.

Anthropic describes Claude Code as "an agentic coding tool that reads your codebase, edits files, runs commands, and integrates with your development tools", available "in your terminal, IDE, desktop app, and browser" (Anthropic, Claude Code documentation).

OpenAI describes the Codex command line as a way to "Inspect, edit, and run code from your terminal", and lists the ChatGPT desktop app, ChatGPT on the web, the Codex IDE extension and Codex cloud alongside it (OpenAI, Codex CLI).

SurfaceClaude CodeCodex
Terminalyesyes
IDE extensionVS Code, JetBrainsyes
Desktop appyesthrough the ChatGPT app
Browseryesyes, as Codex cloud
Where each one runs, according to its own documentation. As of 28 September 2026.Source: Anthropic, Claude Code documentation, OpenAI, Codex documentation

That table is why the usual framing of "cloud versus terminal" does not hold. Both now run in both places. Deciding on the strength of the interface means deciding on something both of them offer.

Why is the decision easier to reverse than it looks?

Because the file holding your project knowledge belongs to no vendor. Codex reads an AGENTS.md, Claude Code reads a CLAUDE.md and, under a clearly stated condition, an AGENTS.md as well. Both are Markdown files in the repository that travel with the project.

OpenAI puts it this way: "Codex reads AGENTS.md files before doing any work" (OpenAI, Custom instructions with AGENTS.md). Anthropic's documentation names the condition on its side explicitly: Claude Code reads the AGENTS.md as the project instructions when there is no CLAUDE.md in your working directory or above it. If both exist, only the CLAUDE.md is read by default, and anyone who wants both imports one into the other. Version 2.1.277 or later is required (Anthropic, How Claude remembers your project).

It gets interesting at the question of which file wins when several are lying around. Both vendors solve that in layers, and both draw the same line between your machine and your repository.

Which instructions files get read, and in what order

sits in the repository and travels with a clone

Codex

  1. ~/.codex/AGENTS.override.md
  2. ~/.codex/AGENTS.md
  3. AGENTS.md from the Git root downwards
  4. AGENTS.md in the current directory

2 of 4 layers sit in the repository

Claude Code

  1. your organisation's managed policy
  2. ~/.claude/CLAUDE.md
  3. ./CLAUDE.md, otherwise ./AGENTS.md
  4. CLAUDE.md in a subdirectory, on demand

2 of 4 layers sit in the repository

Order and filenames per each vendor's documentation, read on 28 September 2026. With Codex, the closer a file sits to your working directory, the later it appears in the combined prompt and the more it counts.

The practical consequence sits in the lower half of both columns. Half the layers live in the repository and survive a change of tool. What sits on your machine, under ~/.codex or ~/.claude, is your personal habit and can be rewritten in a quarter of an hour.

That is exactly what separates this from the question of Claude Code or Cursor: there you are choosing between two ways of working, here between two vendors of the same way of working.

How much of your project knowledge is tied to the tool at all?

Less than it feels like. For this article we counted the repository this website lives in, and the result is the clearest number there is against commitment panic.

Project documentation in this repository, words per file

CLAUDE.md3652URL-STRUCTURE.md2883HANDOVER.md2587GO-LIVE.md2495PLATFORM-SHIFT.md860README.md155

Six top-level documents, 12,632 words in total, counted on 28 September 2026 in this website's repository.

Project documentation in this repository, words per file
CLAUDE.md3652
URL-STRUCTURE.md2883
HANDOVER.md2587
GO-LIVE.md2495
PLATFORM-SHIFT.md860
README.md155

3,652 of the 12,632 words sit in the instructions file, so a good quarter. The other three quarters are ordinary project documentation that any agent can read, because it is Markdown in the repository: the URL decisions, the handover state, the ordered runbook to the first sale. Change tools and three quarters of your setup comes along untouched, while you rewrite the last quarter.

Two more numbers from the same run, both uncomfortable. There is no editor configuration anywhere in the repository, no .vscode, no .cursor, no .idea. And our CLAUDE.md runs to 414 lines, while Anthropic's own documentation names "under 200 lines per CLAUDE.md file" as the target and says plainly that longer files consume more context and reduce adherence (Anthropic, How Claude remembers your project).

So we sit at twice the recommendation of the vendor whose tool we run. That is not a footnote. It is the most common reason a rule goes unfollowed, and it is a tidying job that a move to a second tool would force anyway. What belongs in a file like that, and what is better off as its own skill, we have written up separately.

What actually separates them day to day?

The default for how much the agent may do before it asks you. That is the difference you feel every working day, and comparison tables usually carry it as a footnote.

With Codex you set the boundaries per run. OpenAI describes the /permissions command as the place where you choose "what Codex is allowed to do" and "set the boundaries for each run", including when Codex may "edit files or run commands without asking" (OpenAI, Codex CLI). The approach is: mark out the fence once, then let it run.

Claude Code asks before actions by default, and documents a separate mechanism for the case where an instruction really has to be enforced. The documentation is notably candid here: Claude treats instructions files "as context, not enforced configuration", and anyone who wants to block an action regardless of what the model decides uses a hook for it (Anthropic, How Claude remembers your project).

For your decision that means: an instructions file is a request in both cases, not a lock. If you need a hard boundary, something like "this file is never touched", you build it outside the instructions file either way. That is the insight that saved us the most time, and it appears in none of the comparison tables we read for this article.

What does it cost to get started?

Not the same, contrary to what almost everyone writes. The two pricing pages spell out different entry points, and the gap sits exactly where somebody without a budget begins.

Claude CodeCodex
Free plannot includedincluded
Cheapest paid plan with accessPro, $17 per month billed annually, $20 monthlyGo, $8 per month
Next tierMax, from $100 per monthPlus, $20 per month
Teamfrom $20 per seat billed annuallyfrom $20 per user billed annually
Entry prices, as each pricing page spells them out. As of 28 September 2026.Source: claude.com/pricing, OpenAI, Pricing

OpenAI's pricing page states that Codex is "included in your ChatGPT Free, Go, Plus, Pro, Business, Edu, or Enterprise plan" (OpenAI, Pricing). On Anthropic's pricing page, Claude Code is absent from the Free plan's list and present from Pro upwards (claude.com/pricing).

So the widespread claim that "both start at $20" appears on neither page. It holds for Plus against Pro and misses that Codex has two tiers below that. If you already pay for ChatGPT, you have already paid for Codex, and if you already pay for Claude Pro, you have already paid for Claude Code. For most people that is the most honest basis for deciding this question.

What stays open: both pages name only "from $100" for the top individual tier and never write out the higher figure. And both count in usage limits rather than tokens, which vary by model and by task. If you bill through the programming interface instead and want to know what your way of working costs there, the Claude Code cost calculator works it out.

How do you know a delegated result holds?

By a chain that reaches a verdict without you. This is where the two tools behave identically and where most setups come apart: an agent that cannot check whether it is finished hands back a proposal that you end up reading anyway.

What that looks like here can be counted. This article went through the same chain as every change to this website: 183 tests across 8 test files, measured in the run of 28 September 2026, then a full build in which every marketing route has to be served statically. If either step fails, nothing is published. The run that wrote this text could have stopped there and ended empty-handed, and that is by design.

From which follows a rule that has nothing to do with the choice of tool. The quality of your delegation is the quality of your counter-check. If you have tests, a build or any other machine-run verification, you can hand over a great deal with either tool. If you have none of that, either one hands you back a proposal and you check it yourself.

So if you are starting today, the first investment is not the subscription. It is a single command that checks your project and ends in a clear yes or no. After that the tool question really does come down to interface.

When do you take Codex, when Claude Code?

By your starting position, not by the feature list. Three cases cover most situations.

Take Codex if you pay for ChatGPT anyway, if you would rather mark out the boundaries once per run than confirm things one at a time, and if you want the entry to cost as little as possible. The Go plan at $8 is the cheapest paid route to either of these agents that we found on the two pricing pages.

Take Claude Code if you pay for Claude Pro or Max anyway, if you want the question before an action as a safety net rather than a brake, and if the range of surfaces matters to you, from terminal to editor to browser.

Take both if your project benefits from two independent readings and the second bill does not bother you. Keep the shared rules in one file and point to it from the second rather than duplicating them. Both vendors' documentation allows for this case explicitly.

What appears in none of the three: the question of which model leads on which benchmark. That order changes several times a year. Your way of working does not.

How do you test both in one week?

With a task you have to do anyway, and a week. The routine is deliberately small, because a comparison that costs a day of preparation never happens.

  1. Monday. Write your project rules into a single file in the repository. Build steps, conventions, the two or three decisions nobody can see from the outside. Keep it under 200 lines.
  2. Tuesday and Wednesday. Do the task with whichever tool your existing subscription includes. Note every point where you had to step in.
  3. Thursday. Take the same file and the same repository and do the next task with the other one. You set nothing up again, and that is the actual test.
  4. Friday. Compare not the results but the number of interventions and the places where you had to explain yourself.

By Friday evening you have an answer nobody can work out for you, and something more durable alongside it: a project file both tools understand. You keep that whichever way the decision goes.

If you would rather walk that week guided than alone, it is close to how our course The Claude Code System is built: four weeks, weekly live sessions, and in week one everyone writes their first instructions file. The free live webinar gives you a sense of it first. Every course is listed under Courses, and the people behind it are on About.

Frequently asked

Which is better, Codex or Claude Code?

For most tasks there is little between them, and the decision is cheaper to undo than most comparisons suggest. Decide on three things instead: how much you want to approve per run, which subscription you are already paying for, and whether you have a counter-check that can judge a delegated result without you.

Does Claude Code read an AGENTS.md?

Yes, under one condition. Anthropic's documentation states that Claude Code reads a repository's AGENTS.md as the project instructions when there is no CLAUDE.md in your working directory or above it. If both exist, it reads only the CLAUDE.md by default. Claude Code version 2.1.277 or later is required.

What does each one cost to start with?

According to OpenAI's pricing page, Codex is included in every ChatGPT plan, including the Free plan and the Go plan at $8 per month. Claude Code is not part of the Free plan; the cheapest paid route is Pro at $17 per month on annual billing, $20 billed monthly. The common claim that both start at $20 appears on neither pricing page.

Can I run both in the same repository?

Yes. Both read their instructions from ordinary Markdown files in the project, and neither documentation asks for exclusivity. If you run both permanently, keep the shared rules in one file and point to it from the second rather than duplicating them.

Do I have to commit to one of them for good?

No, and that is the point of this article. The expensive part of your setup is the project documentation, and it sits as Markdown in the repository rather than in one program's settings. What a switch actually costs you is getting used to a different interface, not rebuilding your project knowledge.

Written by

Marco Kohns

Co-founder of ENLIX, lecturer in AI and growth

Marco worked as a growth product manager at a Silicon Valley scale-up and has been teaching that way of working ever since. Today he runs ENLIX with Tobias and builds two products of his own on the same systems, which is what the courses open up.

  • Growth product manager at a Silicon Valley scale-up, Series A to B, backed by a16z, General Catalyst and Sapphire, with users in over 100 countries and more than 20,000 cities
  • Peer-reviewed research in the Journal of Business Research on generative AI in growth, with Prof. René Bohnsack. The research began in summer 2022, months before ChatGPT was public
  • Executive education lecturer at Católica-Lisbon, over 10 seminars, more than 1,500 people taught

Keep reading