Codex is included with every paid ChatGPT plan: Free, Go at $8 a month, Plus at $20, Pro at $200, and Business at $20 per user per month on annual billing, metered by a rolling five-hour window with weekly caps and overflow bought as credits. Devin is a standalone product from Cognition: Free, Pro at $20, Max at $200, Teams from an $80 monthly minimum, and Enterprise billed in Agent Compute Units at a contracted rate, metered by agent work rather than by time window. Codex is usually already paid for. Devin earns a separate line item on enterprise deployment, per-task cost attribution, and parallel fleet execution.
- Codex is included with ChatGPT plans. For most teams the marginal cost of trying it is zero.
- Devin is a separate purchase, and has to justify itself as one.
- The meters differ: Codex measures a rolling time window, Devin measures agent work.
- Only one of those meters lets you say what a specific task cost.
- Devin Enterprise offers a deployment inside your own cloud. Codex does not.
- Codex is OpenAI models only. Devin runs several vendors plus Cognition’s own SWE models.
Why this is the closest matchup in the category
Most comparisons in this space are cross-category and therefore easy: an editor against a terminal agent, a review bot against an autocomplete. This one is not. Devin and Codex both accept a scoped brief, work in an environment you did not have to sit in front of, and hand back a reviewable change. Both run locally and in the cloud. Both are reachable from a CLI, a web surface, and an API.
| Codex | Devin | |
|---|---|---|
| Vendor | OpenAI | Cognition |
| Sold as | Included with ChatGPT | A standalone product |
| Free tier | Yes, small volumes | Yes, light quota |
| Individual entry | $20/mo with ChatGPT Plus | $20/mo Pro |
| Individual top | $200/mo with ChatGPT Pro | $200/mo Max |
| Team entry | $20 per user per month, annual (Business) | $80/mo minimum, $40 full seats |
| Occasional users | Full seat price | Free flex seats |
| Enterprise | ChatGPT Enterprise, custom | Custom, billed in ACUs |
| Metered by | Rolling five-hour window plus weekly caps | Agent work, as credits or ACUs |
| Overflow | Credits from a published rate card | On-demand credits, which roll over |
| Models | OpenAI only | Several vendors plus Cognition SWE models |
| Runs in your own cloud | No | VPC deployment at Enterprise |
| Surfaces | CLI, IDE extension, web, app, code review | Cloud, CLI, Desktop editor, Review, API |
The meters measure different things
This is the most consequential difference and the least discussed. Codex meters a rolling five-hour window with weekly caps on top. Devin meters agent work, expressed as on-demand credits on self-serve plans and Agent Compute Units on enterprise order forms.
A time-window meter answers "can I keep going right now". A work meter answers "what did that cost". They are not interchangeable, and which question your organization needs to answer should decide the product.
| Question | Codex | Devin |
|---|---|---|
| Can I keep working right now? | Clear: check the window | Clear: check the credit balance |
| What did this task cost? | Not directly answerable | Per-session consumption |
| What did this project cost? | Not directly answerable | Aggregable from sessions |
| Can I cap a person? | Plan tier is the cap | Per-user usage policies at Enterprise |
| Can I cap the org? | Sum of seats | Org ACU limit; work halts at it |
| Does unused allowance carry? | No, windows reset | Yes, credits roll over |
| Is the cap hard or soft? | Hard: wait for the window | Hard: activity stops at the limit |
If per-task and per-repository attribution is the thing you actually need, note that neither vendor solves it across the tools your team also runs. The cost allocation guide covers the general shape of that problem.
Where Codex wins
- It is already paid for. Codex ships with ChatGPT from Free upward. There is no separate Codex subscription and no procurement conversation for a team that already has ChatGPT.
- The entry cost of trying is zero. That is not a small advantage. Most agent rollouts fail at the trial stage because getting a trial approved takes longer than the trial.
- It is closer to the terminal. The CLI is the primary surface, and for engineers who already live there the adoption curve is short.
- The window is legible. A rolling five-hour meter is easy to reason about and easy to plan a day around, once you know it exists.
- One vendor, one identity. If ChatGPT Business or Enterprise is already deployed with SSO, Codex inherits it rather than adding an identity surface.
- It runs headless. CI and scripted use are first-class, which matters more than it sounds once agent work becomes routine.
The full ladder, the two meters, and the credit rate card are in the Codex pricing guide.
Where Devin earns a separate line item
- Deployment inside your own cloud. Devin Enterprise offers a VPC option so code stays within your boundary. Codex has no equivalent, and for some security reviews this ends the comparison immediately.
- Per-session cost accounting. Consumption is visible per session and aggregable, which makes chargeback and per-project attribution possible rather than aspirational.
- Hard organizational limits. Enterprise admins set org-level ACU caps and per-user usage policies. At the limit, activity halts and users are told to contact an admin.
- Free flex seats. Unlimited occasional users at no fixed cost, drawing from a shared credit pool. For a long tail of light users this is structurally cheaper than any per-seat model.
- Model breadth. Several frontier vendors plus Cognition’s own SWE series, against OpenAI models only.
- Parallel fleet execution. Devin is designed for many agents working at once rather than for one session you are attending.
- An editor in the bundle. Devin Desktop is the editor formerly called Windsurf, included in the same product line and the same seat.
How to run the comparison properly
Both have free tiers, so this can be decided with evidence in two afternoons rather than with a spreadsheet over two weeks.
Pick three tasks you have already finished
Use work whose correct answer you know: one small bug fix, one dependency or migration chore, one feature of a size you would normally scope in a ticket. Known answers make the comparison honest.
Write one brief and use it for both
Same wording, same acceptance criteria, same base commit. Rewriting the prompt for each vendor is how a trial quietly becomes a demonstration of the conclusion you wanted.
Time the review, not the run
Record how long it took a human to read, correct, and accept each result. Agent wall-clock time is the number vendors report and the number that matters least.
Record consumption in each vendor’s own unit
Codex reports against the rolling window; Devin reports credits or ACUs per session. Do not convert them into a single number. The point is to see which unit answers the questions you will be asked at renewal.
Test the failure case deliberately
Give both a task that is under-specified on purpose. What an agent does with an ambiguous brief tells you more about living with it than three successes do.
Check the constraint that is not negotiable
If code cannot leave your cloud, or if per-project attribution is a finance requirement, that decides it regardless of the trial results. Establish those first so the trial answers the remaining question.
Related reading: Claude Code against Devin for the session-versus-assignment axis, Devin enterprise pricing for the ACU order form, and the autonomous agent field guide for the wider category.
Questions people ask
Is Codex free if I have ChatGPT?
Codex is included with every paid ChatGPT plan and available at small volumes on Free and Go. There is no separate Codex subscription, so for most teams the marginal cost of trying it is zero.
Is Devin better than Codex?
On capability they are close enough that the answer changes with every model release. On commercial shape they differ sharply: Codex is bundled into a subscription most teams already hold, while Devin is a standalone purchase with per-session cost accounting, hard organizational limits, and a deployment option inside your own cloud.
What is an ACU?
An Agent Compute Unit is Cognition’s unit of agent work. Enterprise customers are billed in ACUs at the rate set in their order form. On self-serve plans the equivalent unit is an on-demand credit, and Cognition documents that a credit carries the same dollar value as the ACU it replaced.
Can I see what a single task cost?
On Devin, yes: consumption is reported per session and aggregable at the organization level. On Codex, not directly, because the meter is a rolling five-hour window with weekly caps rather than a per-task unit.
Which one can run inside our own cloud?
Devin, at Enterprise, offers a VPC deployment so code stays within your boundary. Codex does not offer an equivalent. If that is a hard requirement, it decides the comparison on its own.
Which models does each run?
Codex runs OpenAI models only. Devin offers frontier models from several vendors alongside Cognition’s own SWE series. Model breadth matters less than people expect on routine work and more than expected on unusual work.
What happens when each one runs out?
Codex pauses until the rolling window refills, or you buy credits from the published rate card. Devin continues on on-demand credits, which roll over rather than expiring, and an organization at its ACU limit stops until an admin raises it. Both are hard stops rather than silent overage.
Should a small team buy Devin as well as Codex?
Only against a specific requirement: cloud deployment, per-project cost attribution, a hard org-wide cap, or a long tail of occasional users who would be free flex seats. Absent one of those, use the Codex you already pay for.
Sources
Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.
- Devin pricing
- ChatGPT plans and pricing
- OpenAI Codex documentation
- Devin documentation: self-serve plans Plan quota, on-demand credits, rollover, and the ACU equivalence note.
- Devin documentation: enterprise billing ACU billing, org limits, and per-user usage policies.