Devin vs Codex: two cloud agents, two meters

Devin and Codex are the two products that most closely occupy the same idea: describe a task, let an agent work in its own environment, review a pull request. They differ far less in what they do than in how they are sold, and for most teams that is the deciding difference rather than a footnote to it.

By the Continuum team. We build a workbench that runs Claude Code, Codex, and their peers, so the model rates quoted here are the ones our own cost analytics ship with.

The short version

Codex is included with every paid ChatGPT plan: Free, Go at $8 a month, Plus at $20, Pro at $200, and Business at $20 per user per month on annual billing, metered by a rolling five-hour window with weekly caps and overflow bought as credits. Devin is a standalone product from Cognition: Free, Pro at $20, Max at $200, Teams from an $80 monthly minimum, and Enterprise billed in Agent Compute Units at a contracted rate, metered by agent work rather than by time window. Codex is usually already paid for. Devin earns a separate line item on enterprise deployment, per-task cost attribution, and parallel fleet execution.

What you need to know
  • Codex is included with ChatGPT plans. For most teams the marginal cost of trying it is zero.
  • Devin is a separate purchase, and has to justify itself as one.
  • The meters differ: Codex measures a rolling time window, Devin measures agent work.
  • Only one of those meters lets you say what a specific task cost.
  • Devin Enterprise offers a deployment inside your own cloud. Codex does not.
  • Codex is OpenAI models only. Devin runs several vendors plus Cognition’s own SWE models.

Why this is the closest matchup in the category

Most comparisons in this space are cross-category and therefore easy: an editor against a terminal agent, a review bot against an autocomplete. This one is not. Devin and Codex both accept a scoped brief, work in an environment you did not have to sit in front of, and hand back a reviewable change. Both run locally and in the cloud. Both are reachable from a CLI, a web surface, and an API.

checked aug 2026
CodexDevin
VendorOpenAICognition
Sold asIncluded with ChatGPTA standalone product
Free tierYes, small volumesYes, light quota
Individual entry$20/mo with ChatGPT Plus$20/mo Pro
Individual top$200/mo with ChatGPT Pro$200/mo Max
Team entry$20 per user per month, annual (Business)$80/mo minimum, $40 full seats
Occasional usersFull seat priceFree flex seats
EnterpriseChatGPT Enterprise, customCustom, billed in ACUs
Metered byRolling five-hour window plus weekly capsAgent work, as credits or ACUs
OverflowCredits from a published rate cardOn-demand credits, which roll over
ModelsOpenAI onlySeveral vendors plus Cognition SWE models
Runs in your own cloudNoVPC deployment at Enterprise
SurfacesCLI, IDE extension, web, app, code reviewCloud, CLI, Desktop editor, Review, API

The meters measure different things

This is the most consequential difference and the least discussed. Codex meters a rolling five-hour window with weekly caps on top. Devin meters agent work, expressed as on-demand credits on self-serve plans and Agent Compute Units on enterprise order forms.

A time-window meter answers "can I keep going right now". A work meter answers "what did that cost". They are not interchangeable, and which question your organization needs to answer should decide the product.

QuestionCodexDevin
Can I keep working right now?Clear: check the windowClear: check the credit balance
What did this task cost?Not directly answerablePer-session consumption
What did this project cost?Not directly answerableAggregable from sessions
Can I cap a person?Plan tier is the capPer-user usage policies at Enterprise
Can I cap the org?Sum of seatsOrg ACU limit; work halts at it
Does unused allowance carry?No, windows resetYes, credits roll over
Is the cap hard or soft?Hard: wait for the windowHard: activity stops at the limit

If per-task and per-repository attribution is the thing you actually need, note that neither vendor solves it across the tools your team also runs. The cost allocation guide covers the general shape of that problem.

Where Codex wins

  • It is already paid for. Codex ships with ChatGPT from Free upward. There is no separate Codex subscription and no procurement conversation for a team that already has ChatGPT.
  • The entry cost of trying is zero. That is not a small advantage. Most agent rollouts fail at the trial stage because getting a trial approved takes longer than the trial.
  • It is closer to the terminal. The CLI is the primary surface, and for engineers who already live there the adoption curve is short.
  • The window is legible. A rolling five-hour meter is easy to reason about and easy to plan a day around, once you know it exists.
  • One vendor, one identity. If ChatGPT Business or Enterprise is already deployed with SSO, Codex inherits it rather than adding an identity surface.
  • It runs headless. CI and scripted use are first-class, which matters more than it sounds once agent work becomes routine.

The full ladder, the two meters, and the credit rate card are in the Codex pricing guide.

Where Devin earns a separate line item

  • Deployment inside your own cloud. Devin Enterprise offers a VPC option so code stays within your boundary. Codex has no equivalent, and for some security reviews this ends the comparison immediately.
  • Per-session cost accounting. Consumption is visible per session and aggregable, which makes chargeback and per-project attribution possible rather than aspirational.
  • Hard organizational limits. Enterprise admins set org-level ACU caps and per-user usage policies. At the limit, activity halts and users are told to contact an admin.
  • Free flex seats. Unlimited occasional users at no fixed cost, drawing from a shared credit pool. For a long tail of light users this is structurally cheaper than any per-seat model.
  • Model breadth. Several frontier vendors plus Cognition’s own SWE series, against OpenAI models only.
  • Parallel fleet execution. Devin is designed for many agents working at once rather than for one session you are attending.
  • An editor in the bundle. Devin Desktop is the editor formerly called Windsurf, included in the same product line and the same seat.

How to run the comparison properly

Both have free tiers, so this can be decided with evidence in two afternoons rather than with a spreadsheet over two weeks.

01

Pick three tasks you have already finished

Use work whose correct answer you know: one small bug fix, one dependency or migration chore, one feature of a size you would normally scope in a ticket. Known answers make the comparison honest.

02

Write one brief and use it for both

Same wording, same acceptance criteria, same base commit. Rewriting the prompt for each vendor is how a trial quietly becomes a demonstration of the conclusion you wanted.

03

Time the review, not the run

Record how long it took a human to read, correct, and accept each result. Agent wall-clock time is the number vendors report and the number that matters least.

04

Record consumption in each vendor’s own unit

Codex reports against the rolling window; Devin reports credits or ACUs per session. Do not convert them into a single number. The point is to see which unit answers the questions you will be asked at renewal.

05

Test the failure case deliberately

Give both a task that is under-specified on purpose. What an agent does with an ambiguous brief tells you more about living with it than three successes do.

06

Check the constraint that is not negotiable

If code cannot leave your cloud, or if per-project attribution is a finance requirement, that decides it regardless of the trial results. Establish those first so the trial answers the remaining question.

Related reading: Claude Code against Devin for the session-versus-assignment axis, Devin enterprise pricing for the ACU order form, and the autonomous agent field guide for the wider category.

Questions people ask

Is Codex free if I have ChatGPT?

Codex is included with every paid ChatGPT plan and available at small volumes on Free and Go. There is no separate Codex subscription, so for most teams the marginal cost of trying it is zero.

Is Devin better than Codex?

On capability they are close enough that the answer changes with every model release. On commercial shape they differ sharply: Codex is bundled into a subscription most teams already hold, while Devin is a standalone purchase with per-session cost accounting, hard organizational limits, and a deployment option inside your own cloud.

What is an ACU?

An Agent Compute Unit is Cognition’s unit of agent work. Enterprise customers are billed in ACUs at the rate set in their order form. On self-serve plans the equivalent unit is an on-demand credit, and Cognition documents that a credit carries the same dollar value as the ACU it replaced.

Can I see what a single task cost?

On Devin, yes: consumption is reported per session and aggregable at the organization level. On Codex, not directly, because the meter is a rolling five-hour window with weekly caps rather than a per-task unit.

Which one can run inside our own cloud?

Devin, at Enterprise, offers a VPC deployment so code stays within your boundary. Codex does not offer an equivalent. If that is a hard requirement, it decides the comparison on its own.

Which models does each run?

Codex runs OpenAI models only. Devin offers frontier models from several vendors alongside Cognition’s own SWE series. Model breadth matters less than people expect on routine work and more than expected on unusual work.

What happens when each one runs out?

Codex pauses until the rolling window refills, or you buy credits from the published rate card. Devin continues on on-demand credits, which roll over rather than expiring, and an organization at its ACU limit stops until an admin raises it. Both are hard stops rather than silent overage.

Should a small team buy Devin as well as Codex?

Only against a specific requirement: cloud deployment, per-project cost attribution, a hard org-wide cap, or a long tail of occasional users who would be free flex seats. Absent one of those, use the Codex you already pay for.

Sources

Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.

  1. Devin pricing
  2. ChatGPT plans and pricing
  3. OpenAI Codex documentation
  4. Devin documentation: self-serve plans Plan quota, on-demand credits, rollover, and the ACU equivalence note.
  5. Devin documentation: enterprise billing ACU billing, org limits, and per-user usage policies.
Try it

What did that
task cost?

Continuum prices every agent session into one ledger by repository, provider, and model, with a live quota gauge per subscription.

free app · your subscriptions · local-first