Run codex doctor first. It checks runtime, auth, network reachability, config, terminal, and MCP in one command, and it has a JSON mode for support tickets. After that, assume the sandbox: network errors during package installs, silent write failures, and commands that fail only inside Codex are almost always workspace-write doing its job, because it disables network access by default. Check sandbox_mode and approval_policy before assuming anything is broken. A Codex stuck thinking is four separate failures wearing one spinner: a genuinely long think, a dropped response stream, an MCP handshake that never finished, or a wedged sandbox child. The other frequent surprise is billing: codex login status tells you whether the session is on your ChatGPT plan or on an API key, which bills separately at standard API rates.
- Start with
codex doctor. One command, and it has a--jsonmode for bug reports. - Network errors on install:
workspace-writedisables the network by default. - Cannot write anything: the sandbox is
read-only. - Billed per token: the session is signed in with an API key, not your plan.
- Test with full access once, to isolate. Then put the boundary back.
- On Linux and WSL2 the sandbox needs
bubblewrapinstalled. - Stuck on Thinking: press Esc, not Ctrl+C. Esc interrupts the turn and keeps the session.
Start with codex doctor
Codex ships a single diagnostic that covers most of what you would otherwise check by hand. It landed in v0.131.0 and has grown since; as of August 2026 it reports runtime, authentication, network reachability, terminal environment, configuration, Git state, and MCP server resolution in one grouped, colour-coded report.
codex --version
codex doctor
# structured and redacted, for a bug report or a support ticket
codex doctor --json
The dominant cause
| Symptom | Actually | Fix |
|---|---|---|
ECONNREFUSED during npm install | Network off in workspace-write | Enable network_access |
EAI_AGAIN, DNS failures, hanging fetches | Same | Same |
| Cannot write any file | Sandbox is read-only | --sandbox workspace-write |
| Cannot write outside the project | Working as designed | Add a writable_roots entry |
Cannot write to /tmp | exclude_slash_tmp or exclude_tmpdir_env_var set | Unset it, or add the path |
| A command works in your shell, not in Codex | Sandbox restriction | Check the mode first |
| Constant approval prompts | approval_policy is untrusted | Use on-request |
| No prompts at all, and edits everywhere | Someone left --yolo on | Put the boundary back |
# does it work with restrictions removed?
codex --sandbox danger-full-access --ask-for-approval never
# if yes, it was policy, not a bug. Fix the config; do not stay here.
| Setting | Values | Controls |
|---|---|---|
sandbox_mode | read-only, workspace-write, danger-full-access | What Codex can do |
approval_policy | untrusted, on-request, never | When it has to ask |
They are independent, which is why the failure is confusing: you can have a permissive approval policy and still be blocked by the sandbox, and the message you see comes from the command that failed rather than from Codex explaining its own policy. As of August 2026 the default pairing in a version-controlled folder is workspace-write with on-request.
The network default
This is the single highest-traffic Codex complaint, and it is one line of configuration. In workspace-write, outbound network access is off unless you turn it on.
sandbox_mode = "workspace-write"
approval_policy = "on-request"
[sandbox_workspace_write]
network_access = true
writable_roots = ["/Users/you/.cache/pnpm"]
codex -c sandbox_workspace_write.network_access=true
Billing and authentication
| Symptom | Cause | Fix |
|---|---|---|
| Charged per token despite a ChatGPT plan | Signed in with an API key, not the plan | codex logout, then codex login |
| Login never completes | No browser, or a proxy | codex login --device-auth |
| Asked to log in every launch | ~/.codex/auth.json not writable | Check ownership of ~/.codex |
401 after a while | Session expired | codex logout, then codex login |
| Works locally, fails on a server | No browser on the host | Device auth, or an API key |
| Signed in as the wrong workspace | Multiple ChatGPT workspaces | codex login status to confirm |
codex login status # the active method: ChatGPT sign-in, or an API key
codex logout && codex login
# headless box with no browser
codex login --device-auth
# or explicitly with a key, without leaving it in your shell profile
printenv OPENAI_API_KEY | codex login --with-api-key
Where the sandbox actually comes from
The sandbox is not one implementation, and knowing which one you are on explains most of the platform-specific weirdness.
| Platform | Mechanism | You need to |
|---|---|---|
| macOS | The built-in Seatbelt framework | Nothing. It works out of the box. |
| Linux and WSL2 | bubblewrap, plus kernel filtering | Install bubblewrap with your package manager |
| Windows, PowerShell | The native Windows sandbox | Nothing |
| Windows, WSL2 | The Linux implementation | Install bubblewrap inside the distribution |
sudo apt-get install bubblewrap # Ubuntu and Debian
sudo dnf install bubblewrap # Fedora
which bwrap # Codex uses the first bwrap on PATH
Codex stuck thinking: four different failures, one spinner
The spinner is the least informative thing in the TUI. A Codex stuck thinking is not one bug, it is four, and they have nothing in common except that the interface looks identical for all of them. Work out which before you restart anything, because restarting fixes one of the four and loses your session in the other three.
| What you see | It is | Fix |
|---|---|---|
| Spinner for 5 to 10 minutes, then a normal answer | The model genuinely thinking | Lower model_reasoning_effort |
Reconnecting... 2/5, then Stream disconnected before completion | The response stream died and Codex is retrying | Network, proxy, or provider. See below. |
| Hangs at launch, before any output, with MCP servers configured | An MCP handshake that never completed | Raise startup_timeout_sec |
| Stop button does nothing, and the state survives a restart | A sandbox child wedged, often at 100% CPU | Kill the process, start a new thread |
Take the four in order of how often they are the answer.
Rule out a long think first
/status
It prints the session configuration, including the model and reasoning effort. On the highest efforts a hard prompt genuinely takes minutes, and the TUI shows nothing while it does. Users have reported 5 to 10 minute waits that were real work, not a hang. Drop to a lower effort with /model and see whether the same prompt answers in seconds.
Read the reconnect counter, if there is one
Reconnecting... 2/5 (7m 20s - esc to interrupt)
Stream disconnected before completion: Operation timed out
That is a dropped SSE stream, not a stuck model. Codex reconnects a bounded number of times and then gives up, which is why a flaky link surfaces as a very long silence followed by a single error. Check your provider status page and anything doing TLS interception between you and it.
Suspect MCP if it hangs before the first token
MCP client for `context7` timed out after 10 seconds.
Add or adjust `startup_timeout_sec` in your config.toml
Every MCP server gets 10 seconds to finish the initialize handshake. A cold npx has to reach the registry, download the package, and on Windows walk every new file past the real-time scanner, which routinely blows that budget. The same servers start fine when you run the command by hand, which is exactly why this one is hard to spot.
If the stop button is dead and a restart does not clear it, look for a wedged child
ps aux | grep -i codex
# a child pinned near 100% CPU is the one holding the turn open
A sandboxed command that grabbed a file lock and never let go leaves the turn unfinishable. Because the thread state persists, the conversation comes back stuck after you relaunch the app. Kill the child, then start a new thread rather than trying to revive that one.
[mcp_servers.context7]
command = "npx"
args = ["-y", "@upstash/context7-mcp"]
startup_timeout_sec = 30 # default 10, the initialize handshake
tool_timeout_sec = 120 # default 60, one tool call
# network tuning is PER PROVIDER, not global
[model_providers.my-provider]
name = "My provider"
base_url = "https://example.com/v1"
request_max_retries = 4 # default 4, HTTP requests
stream_max_retries = 5 # default 5, SSE stream interruptions
stream_idle_timeout_ms = 300000 # default 300000, five minutes
# log_dir defaults to $CODEX_HOME/log; setting it turns on the plaintext TUI log
codex -c log_dir=./.codex-log
tail -F ./.codex-log/codex-tui.log
# and confirm the basics are healthy while you are here
codex doctor
Everything else
| Symptom | Check |
|---|---|
codex: command not found | New terminal, then PATH, then a moved Node prefix |
| Vanished after upgrading Node | npm global prefix moved. Reinstall, or use Homebrew. |
| Behaviour changed overnight | codex --version, then the release notes |
| Slow in WSL | The project is on /mnt/c. Move it. |
| MCP server missing | codex doctor reports stdio command resolution and permissions |
| Answers feel shallow on a hard task | model_reasoning_effort. Values run minimal to xhigh. |
| Need to see what it actually did | Turn on a log directory and tail it |
# log_dir defaults to $CODEX_HOME/log. Setting it explicitly also turns on
# the opt-in plaintext TUI log, codex-tui.log, in that directory:
codex -c log_dir=./.codex-log
tail -F ./.codex-log/codex-tui.log
# non-interactive mode prints its messages inline, so there is no file to watch
codex exec "run the test suite and summarise failures"
Questions people ask
What should I run first when Codex CLI is not working?
codex doctor. As of August 2026 it checks runtime, authentication, network reachability for the provider you are actually using, terminal environment, configuration, Git state, and MCP servers in one command, and codex doctor --json produces a redacted report you can attach to a bug report.
Why does npm install fail inside Codex?
Because workspace-write disables outbound network access by default. Set network_access = true under [sandbox_workspace_write] in config.toml when you need it, or pass it for one run with codex -c sandbox_workspace_write.network_access=true.
Why can Codex not write files?
The sandbox is probably read-only. Run with --sandbox workspace-write, or set sandbox_mode in config.toml. If it can write inside the project but not outside, that is workspace-write working correctly; add the path to writable_roots rather than removing the sandbox.
How do I tell a sandbox restriction from a real bug?
Run once with --sandbox danger-full-access --ask-for-approval never. If the failure disappears, it was policy rather than a defect. Then fix the configuration narrowly instead of staying without a boundary.
Why is Codex charging me when I have a ChatGPT plan?
The session is authenticated with an API key rather than your ChatGPT sign-in, and OpenAI bills API key usage through your Platform account at standard API rates. Run codex login status to see the active method, then codex logout and codex login to sign in with ChatGPT. To pin a machine to one rail, set forced_login_method to chatgpt in config.toml.
How do I log in to Codex on a server with no browser?
Use codex login --device-auth, which completes the flow on another device. Alternatively pipe a key in with printenv OPENAI_API_KEY | codex login --with-api-key, or copy ~/.codex/auth.json from an already authenticated machine and treat it like a password.
Does the Codex sandbox work on Windows?
Yes. In PowerShell it uses the native Windows sandbox, and in WSL2 it uses the Linux implementation, which needs bubblewrap installed inside the distribution. Codex prints a startup warning when it cannot enforce the sandbox, and that warning is worth reading rather than dismissing.
Where are the Codex CLI logs?
log_dir defaults to $CODEX_HOME/log, and setting it explicitly also turns on the opt-in plaintext TUI log. Start with codex -c log_dir=./.codex-log and tail ./.codex-log/codex-tui.log. Non-interactive codex exec prints its messages inline instead, and honours RUST_LOG.
Why is Codex stuck thinking and not responding?
Four different failures show the same spinner. The model may genuinely be thinking, which on high reasoning effort takes minutes with no output. The response stream may have dropped, which you can tell from a Reconnecting counter followed by "Stream disconnected before completion". An MCP server may have missed its 10 second startup handshake, which hangs the session before the first token. Or a sandboxed child may be wedged, in which case the stop button does nothing and the state survives a restart.
How do I interrupt Codex when it is stuck on Thinking?
Press Esc first. It interrupts the current turn and leaves the session alive, so you keep the conversation and can steer it elsewhere. Ctrl+C exits Codex outright, which is what you want only once the turn has stopped. If you do exit, codex resume reopens a recent chat from the current repository.
Can an MCP server make Codex hang at startup?
Yes. Each server gets 10 seconds to complete the initialize handshake, and a cold npx that has to reach the registry and download a package routinely misses it, especially on Windows where every new file passes the real-time scanner. Raise startup_timeout_sec under [mcp_servers.<id>] in config.toml. The per-tool timeout is separate and defaults to 60 seconds via tool_timeout_sec.
How do I make Codex retry longer on a dropped stream?
request_max_retries, stream_max_retries, and stream_idle_timeout_ms tune it, with defaults of 4, 5, and 300000 milliseconds. They are only valid inside a [model_providers.<id>] block, and the built-in openai provider is reserved, so you cannot raise them for the default sign-in. Setting them at the root of config.toml appears to work and changes nothing.
Sources
Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.