Claude Code rate limit reached: what to do right now

You are blocked mid-task and the message names a reset time. Four different things produce that message, and they need four different responses.

By the Continuum team. We build a workbench that runs Claude Code, Codex, and their peers, so the model rates quoted here are the ones our own cost analytics ship with.

The short version

Read the message before doing anything else. "You have hit your session limit" is the rolling 5-hour window and clears on its own. "You have hit your weekly limit" needs a change, not patience. "You have hit your Opus limit" is model-specific and switching model with /model keeps you working immediately. A 429 is not a plan limit at all: it is API throughput on your key or cloud project. Two more messages get filed here and belong to neither family. "Claude reached its tool-use limit for this turn." caps the tool calls in one reply and clears when you click Continue. "API Error: 400 due to tool use concurrency issues." is a malformed request, not a throughput problem, and it is fixed with /model, /rewind, or /clear.

What you need to know
  • The message names the limit: session, weekly, Opus, or a 429. Four problems.
  • A weekly or session limit is shared across models, so /model does not restore access.
  • An Opus limit is the exception: /model gets you working again in one keystroke.
  • /usage shows the bars and reset times. /usage-credits buys headroom past them.
  • Commit first. Being cut off mid-edit is the expensive part.
  • Usage is shared with Claude chat and Cowork on the same account.
  • Two messages here are not limits at all: the per-turn tool-use cap, and a 400 that names concurrency but means a malformed request.

Which limit did you hit

Claude Code tells you exactly which ceiling you reached and when it resets. The wording matters more than it looks, because three of these four are unrelated mechanisms.

Message wording as of August 2026, and what each one means.
You seeIt isWait?Do
You've hit your session limit · resets 3:45pmThe rolling 5-hour windowHoursTake the break. Review the diff.
You've hit your weekly limit · resets Mon 12:00amThe weekly plan allowanceDays/usage-credits, or change how you work
You've hit your Opus limit · resets 3:45pmA model-specific allowanceNo/model to Sonnet and keep going
API Error: Request rejected (429)API throughput on your key or cloud projectSecondsLower concurrency; check the provider console
API Error: Repeated 529 Overloaded errorsService capacity, not youMinutesCheck the status page; try another model

The 5-hour window is rolling, which people misread constantly. Capacity returns continuously as older usage ages past the five-hour mark rather than all at once at a round hour. The message gives you the point at which you are back in business, not the start of a fresh allowance.

The next five minutes

01

Commit whatever is in the tree

git add -A && git commit -m "wip: agent progress before usage limit"

An interrupted agent leaves a partial change. Committing it means you can reason about it later instead of reconstructing what happened from memory.

02

Look at the actual numbers

/usage

You get plan usage bars with reset times, plus a breakdown attributing recent usage to skills, subagents, plugins, and individual MCP servers, and flags for behaviours such as long context or cache misses that account for 10 percent or more of recent usage. Press d or w to switch between the last 24 hours and the last 7 days.

03

If it names Opus, switch model

/model sonnet

The Opus allowance is separate. One keystroke and you are working again, on a model that handles most coding work well and consumes the shared allowance far more slowly.

04

If it is session or weekly, decide between waiting and paying

/usage-credits

Usage credits let you continue past the plan allowance at standard API rates. On Pro and Max it opens your billing settings. On Team and Enterprise with billing access it opens organization usage settings; without billing access it sends a request to your admins after you confirm.

Why usage climbs faster than your activity

People hit these limits after what feels like an hour of light work, then conclude the allowance shrank. It did not. A long-running session draws usage in ways that are invisible from the transcript.

  • Every turn re-sends the whole conversation. A one-line question in a session that has been open all day still carries the entire history. Prompt caching makes that cheap, not free.
  • Cache misses after a break reprocess everything. The cache lifetime is an hour on a subscription and drops to five minutes once you are drawing on usage credits, or on an API key. Lunch costs you a full re-read.
  • Every tool result is another request. A thirty-step task is thirty requests, each carrying the growing prefix.
  • Scheduled tasks fire while the session is idle, sending your full context each time.
  • Agent teammates keep consuming until they exit. Agent teams use roughly 7x the tokens of a standard session when teammates run in plan mode.
  • /compact is itself a large request, because it reads the conversation it summarises. /clear costs nothing.

Stopping it recurring

Ranked by effect on a typical week.
ChangeEffect
Default to Sonnet, reserve Opus for hard reasoningLargest single lever
/clear between unrelated tasksStops paying for stale context on every turn
Lower the effort level with /effort for simple workThinking tokens bill as output
Name files in your promptFewer speculative reads before it can act
Scope test commands so they print lessCommand output is next turn input
Delegate verbose reads to a subagentThe log stays in the subagent context
Disable MCP servers you no longer useStanding overhead removed from every session
Fewer parallel sessions and teammatesConsumption scales with them, roughly linearly
Buy usage creditsWorks, and is the answer that costs money

A 429 is not a usage limit

These two get filed under "rate limit" and neither is a plan allowance. Both resolve without you changing anything about how you work.

429529 Overloaded
CauseThroughput cap on your API key or cloud projectService capacity across all users
ScopeYour organizationEveryone on that model
TimescaleSeconds to minutesMinutes
First check/status for the active credentialstatus.claude.com
LeverLower concurrency, request a higher tier/model, since capacity is tracked per model
The concurrency lever, for a 429 you keep hitting.
# fewer simultaneous tool calls per session
export CLAUDE_CODE_MAX_TOOL_USE_CONCURRENCY=2

# and check which credential is actually billing
# (inside a session)
/status

Claude reached its tool-use limit for this turn.

This one is not a limit on your account, and the wording invites the confusion. It caps how many tool calls Claude makes inside a single reply, and it stops the reply rather than your access. Nothing is exhausted, nothing resets, and there is no wait.

Claude reached its tool-use limit for this turn.

The ceiling is on Anthropic's side. When Claude runs server-executed tools such as web search, the sampling loop that drives them has an iteration limit, and the documented default is 10 iterations per request. On the API that arrives as stop_reason: "pause_turn", and the documented handling is to send the assistant content straight back so the model carries on from where it stopped. The apps do the same thing behind a button.

One ceiling, worn differently by each surface.
WhereWhat you getWhat to do
claude.ai and the Claude desktop appThe banner, with a Continue buttonClick Continue. It resumes in the same conversation.
Your own API integrationstop_reason: "pause_turn"Send the assistant content back, under a continuation cap
Claude CodeNot this message.The CLI drives its own client-side tool loop

The reason the string started showing up in searches appears to be that the ceiling moved. The report that documented it describes Desktop runs that previously chained 60 to 80 tool calls in a turn now stopping at roughly 20, which is the difference between a run you leave alone and a run you sit with. Anthropic has not published the figure, so treat the numbers as one user's measurement rather than a specification.

  • Click Continue. It picks up at the same point, in the same conversation, with the same context. Nothing is lost and nothing is re-done.
  • Split the work into named steps and send them as separate turns. Three prompts of eight tool calls never touch a per-turn ceiling that one prompt of twenty-four does.
  • Turn off connectors you are not using. Fewer available tools means fewer speculative lookups before Claude has what it needs.
  • Name the file, repository, or document. A large share of a burned turn is spent finding the thing whose location you already knew.
  • Calling the API yourself? Treat pause_turn as an ordinary continuation, and cap the number of continuations so a loop cannot run away.
The continuation pattern, with a ceiling of your own.
messages = [{"role": "user", "content": prompt}]

for _ in range(5):                       # your own continuation cap
    response = client.messages.create(model=MODEL, messages=messages, tools=tools)
    if response.stop_reason != "pause_turn":
        break
    messages.append({"role": "assistant", "content": response.content})

API Error: 400 due to tool use concurrency issues.

Despite the name, this is not a throughput problem, and turning concurrency down is usually not the fix. A 400 means the API rejected the request as malformed before doing any work, and what it rejected is the shape of the conversation you sent.

API Error: 400 due to tool use concurrency issues. Run /rewind to recover the conversation.

Two unrelated causes produce that one string, and they want opposite responses. The tell is when it started.

Same message, two mechanisms.
CauseTellFix
Tool-use settings the model does not support, typically extended thinkingIt fails from the first turn, or right after a model change/model to a compatible model, then /clear
A broken tool_use and tool_result pairing in the transcriptIt started mid-session and now every message fails/rewind past the break, or /clear

The second is the one people describe as the conversation being permanently broken, and the mechanism explains why. The API enforces a strict message shape: every tool_use block must be answered by a tool_result carrying the same id, and two messages from the same role cannot sit next to each other. Claude Code resends the whole transcript on every turn, so a single malformed pair poisons every later request. It is not intermittent, it does not age out, and retrying is guaranteed to fail in exactly the same way.

  • A turn interrupted while tool calls were still in flight, which leaves a tool_use with no matching result.
  • Automatic compaction rebuilding the history. The reconstruction merges adjacent same-role messages, and it cannot merge a plain-string content field next to an array content field, so the merge is skipped and two same-role messages survive into the request.
  • Dictated messages, which are the reported trigger for that merge. Voice input has been stored with content as a bare string while typed input is stored as an array of blocks.
  • Compaction subagents inheriting the broken shape and failing in turn, which is why the failure can outlive the turn that caused it.
01

If it failed on the first turn, change model and clear

/model
/clear

This is the documented cause: a request combining tool use with a feature the selected model does not support. If you set a default model in settings, check that one too, because a fresh session inherits it and fails identically.

02

If it started mid-session, rewind past the break

/rewind

Step back to a point before the first failure. This works when one interrupted turn broke the pairing, which is the common case.

03

If rewinding does not clear it, clear the conversation

/clear

The rejected thing is the history, so discarding the history always works. Commit your tree first: the code is untouched, but the reasoning behind it is about to be.

Questions people ask

What does "You have hit your session limit" mean in Claude Code?

You exhausted the rolling five-hour usage window on your plan. It is shared with Claude chat and Cowork on the same account, and it clears on its own; the message shows the reset time. Because the window covers every model, switching model does not restore access.

How long until my Claude Code limit resets?

The message states the reset time, and /usage shows the same thing with bars for each window. The five-hour window is rolling, so headroom returns gradually as older usage ages out rather than all at once. The weekly window resets at a fixed time assigned to your account.

Why did switching to Sonnet not fix my limit?

Session and weekly windows are shared across all models, so changing model does not give you back allowance you already spent. The one exception is the Opus limit, which is model-specific: after that message, /model to another model keeps you working immediately.

Can I keep working after hitting the limit?

Yes, with usage credits. Run /usage-credits to turn them on or request them from an admin. They bill past the plan allowance at standard API rates. Without them, a subscription blocks until the window reopens.

Why do I keep hitting the weekly limit?

Usually Opus left as the default, a session left open for hours so every turn re-sends the whole conversation, scheduled tasks firing while you are idle, or agent teammates still running. Run /usage and read the breakdown, which flags any behaviour accounting for 10 percent or more of recent usage.

What is the difference between a 429 and a usage limit?

A 429 is throughput on your API key or cloud project, measured per minute, and it resolves in seconds. A usage limit is a plan allowance measured over five hours or a week. A 529 is neither: it means the service is at capacity, which you can sometimes route around with /model.

Does upgrading my plan fix it immediately?

A plan change raises the allowance, but the durable fix is usually model choice and session hygiene, which cost nothing. If you are on Opus by default and never clear between tasks, the next tier buys weeks rather than months.

Does Claude Code usage count against Claude chat usage?

Yes. On Pro, Max, Team, and Enterprise the allowance is shared across Claude Code, Claude chat, and Cowork on the same account, which is why a heavy chat morning can shorten your coding afternoon.

What does "Claude reached its tool-use limit for this turn." mean?

Claude hit the cap on how many tool calls it can make inside one reply. It is not a usage limit and nothing resets: the server-side loop that runs tools such as web search has an iteration ceiling, documented at 10 iterations per request, which the API reports as stop_reason "pause_turn". Click Continue and it resumes in the same conversation with the same context.

How do I stop hitting the tool-use limit for this turn?

Split the work into separate turns with named steps, turn off connectors you are not using, and name the file or repository instead of describing it so Claude spends fewer calls finding it. Three prompts of eight tool calls never reach a ceiling that one prompt of twenty-four does.

Does Claude Code show "Claude reached its tool-use limit for this turn"?

No. That message comes from claude.ai and the Claude desktop app, which run tools inside a server-side loop with an iteration ceiling. The Claude Code CLI drives its own client-side tool loop and does not stop at a fixed tool count per turn. The report that made the string well known was filed against the Claude Code repository and closed as unrelated to Claude Code.

How do I fix "API Error: 400 due to tool use concurrency issues"?

The tell is when it started. If it fails from the first turn or right after a model change, the selected model does not support a tool-use feature you have enabled, usually extended thinking: run /model to switch, then /clear. If it started mid-session and now every message fails, the tool_use and tool_result pairing in the transcript is broken, so run /rewind to a point before the first failure, and /clear if that does not help.

Why does /rewind not fix the tool use concurrency error?

Because /rewind only helps when a single interrupted turn broke the message pairing. When automatic compaction rebuilt the history and left two same-role messages adjacent, the damage is spread through the reconstructed transcript, and reporters have found rewinding 40 messages back still fails while /compact runs the same faulty reconstruction again. Clearing the conversation works, because the rejected thing is the history itself.

Does lowering CLAUDE_CODE_MAX_TOOL_USE_CONCURRENCY fix the 400 error?

Almost never. The variable caps how many read-only tools and subagents execute in parallel, defaulting to 10, and it genuinely helps with a 429 on your key or cloud project. The 400 is a malformed request rather than a throughput problem, so lowering concurrency changes nothing about the message shape the API is rejecting.

Sources

Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.

  1. Claude Code error reference
  2. Claude Code: manage costs effectively
  3. Anthropic: handling stop reasons (pause_turn)
  4. claude-code #33969: per-turn tool call limit
  5. claude-code #8763: 400 tool use concurrency
  6. claude-code #37452: 400 after context compaction
  7. Anthropic help: usage limits
  8. Anthropic status
Try it

Thirty minutes
of warning.

Continuum shows live quota and time-to-limit per account, so you finish and commit instead of being stopped mid-edit.

free app · your subscriptions · local-first