Read the message before doing anything else. "You have hit your session limit" is the rolling 5-hour window and clears on its own. "You have hit your weekly limit" needs a change, not patience. "You have hit your Opus limit" is model-specific and switching model with /model keeps you working immediately. A 429 is not a plan limit at all: it is API throughput on your key or cloud project. Two more messages get filed here and belong to neither family. "Claude reached its tool-use limit for this turn." caps the tool calls in one reply and clears when you click Continue. "API Error: 400 due to tool use concurrency issues." is a malformed request, not a throughput problem, and it is fixed with /model, /rewind, or /clear.
- The message names the limit: session, weekly, Opus, or a 429. Four problems.
- A weekly or session limit is shared across models, so
/modeldoes not restore access. - An Opus limit is the exception:
/modelgets you working again in one keystroke. /usageshows the bars and reset times./usage-creditsbuys headroom past them.- Commit first. Being cut off mid-edit is the expensive part.
- Usage is shared with Claude chat and Cowork on the same account.
- Two messages here are not limits at all: the per-turn tool-use cap, and a 400 that names concurrency but means a malformed request.
Which limit did you hit
Claude Code tells you exactly which ceiling you reached and when it resets. The wording matters more than it looks, because three of these four are unrelated mechanisms.
| You see | It is | Wait? | Do |
|---|---|---|---|
You've hit your session limit · resets 3:45pm | The rolling 5-hour window | Hours | Take the break. Review the diff. |
You've hit your weekly limit · resets Mon 12:00am | The weekly plan allowance | Days | /usage-credits, or change how you work |
You've hit your Opus limit · resets 3:45pm | A model-specific allowance | No | /model to Sonnet and keep going |
API Error: Request rejected (429) | API throughput on your key or cloud project | Seconds | Lower concurrency; check the provider console |
API Error: Repeated 529 Overloaded errors | Service capacity, not you | Minutes | Check the status page; try another model |
The 5-hour window is rolling, which people misread constantly. Capacity returns continuously as older usage ages past the five-hour mark rather than all at once at a round hour. The message gives you the point at which you are back in business, not the start of a fresh allowance.
The next five minutes
Commit whatever is in the tree
git add -A && git commit -m "wip: agent progress before usage limit"
An interrupted agent leaves a partial change. Committing it means you can reason about it later instead of reconstructing what happened from memory.
Look at the actual numbers
/usage
You get plan usage bars with reset times, plus a breakdown attributing recent usage to skills, subagents, plugins, and individual MCP servers, and flags for behaviours such as long context or cache misses that account for 10 percent or more of recent usage. Press d or w to switch between the last 24 hours and the last 7 days.
If it names Opus, switch model
/model sonnet
The Opus allowance is separate. One keystroke and you are working again, on a model that handles most coding work well and consumes the shared allowance far more slowly.
If it is session or weekly, decide between waiting and paying
/usage-credits
Usage credits let you continue past the plan allowance at standard API rates. On Pro and Max it opens your billing settings. On Team and Enterprise with billing access it opens organization usage settings; without billing access it sends a request to your admins after you confirm.
Why usage climbs faster than your activity
People hit these limits after what feels like an hour of light work, then conclude the allowance shrank. It did not. A long-running session draws usage in ways that are invisible from the transcript.
- Every turn re-sends the whole conversation. A one-line question in a session that has been open all day still carries the entire history. Prompt caching makes that cheap, not free.
- Cache misses after a break reprocess everything. The cache lifetime is an hour on a subscription and drops to five minutes once you are drawing on usage credits, or on an API key. Lunch costs you a full re-read.
- Every tool result is another request. A thirty-step task is thirty requests, each carrying the growing prefix.
- Scheduled tasks fire while the session is idle, sending your full context each time.
- Agent teammates keep consuming until they exit. Agent teams use roughly 7x the tokens of a standard session when teammates run in plan mode.
/compactis itself a large request, because it reads the conversation it summarises./clearcosts nothing.
Stopping it recurring
| Change | Effect |
|---|---|
| Default to Sonnet, reserve Opus for hard reasoning | Largest single lever |
/clear between unrelated tasks | Stops paying for stale context on every turn |
Lower the effort level with /effort for simple work | Thinking tokens bill as output |
| Name files in your prompt | Fewer speculative reads before it can act |
| Scope test commands so they print less | Command output is next turn input |
| Delegate verbose reads to a subagent | The log stays in the subagent context |
| Disable MCP servers you no longer use | Standing overhead removed from every session |
| Fewer parallel sessions and teammates | Consumption scales with them, roughly linearly |
| Buy usage credits | Works, and is the answer that costs money |
A 429 is not a usage limit
These two get filed under "rate limit" and neither is a plan allowance. Both resolve without you changing anything about how you work.
| 429 | 529 Overloaded | |
|---|---|---|
| Cause | Throughput cap on your API key or cloud project | Service capacity across all users |
| Scope | Your organization | Everyone on that model |
| Timescale | Seconds to minutes | Minutes |
| First check | /status for the active credential | status.claude.com |
| Lever | Lower concurrency, request a higher tier | /model, since capacity is tracked per model |
# fewer simultaneous tool calls per session
export CLAUDE_CODE_MAX_TOOL_USE_CONCURRENCY=2
# and check which credential is actually billing
# (inside a session)
/status
Claude reached its tool-use limit for this turn.
This one is not a limit on your account, and the wording invites the confusion. It caps how many tool calls Claude makes inside a single reply, and it stops the reply rather than your access. Nothing is exhausted, nothing resets, and there is no wait.
Claude reached its tool-use limit for this turn.
The ceiling is on Anthropic's side. When Claude runs server-executed tools such as web search, the sampling loop that drives them has an iteration limit, and the documented default is 10 iterations per request. On the API that arrives as stop_reason: "pause_turn", and the documented handling is to send the assistant content straight back so the model carries on from where it stopped. The apps do the same thing behind a button.
| Where | What you get | What to do |
|---|---|---|
| claude.ai and the Claude desktop app | The banner, with a Continue button | Click Continue. It resumes in the same conversation. |
| Your own API integration | stop_reason: "pause_turn" | Send the assistant content back, under a continuation cap |
| Claude Code | Not this message. | The CLI drives its own client-side tool loop |
The reason the string started showing up in searches appears to be that the ceiling moved. The report that documented it describes Desktop runs that previously chained 60 to 80 tool calls in a turn now stopping at roughly 20, which is the difference between a run you leave alone and a run you sit with. Anthropic has not published the figure, so treat the numbers as one user's measurement rather than a specification.
- Click Continue. It picks up at the same point, in the same conversation, with the same context. Nothing is lost and nothing is re-done.
- Split the work into named steps and send them as separate turns. Three prompts of eight tool calls never touch a per-turn ceiling that one prompt of twenty-four does.
- Turn off connectors you are not using. Fewer available tools means fewer speculative lookups before Claude has what it needs.
- Name the file, repository, or document. A large share of a burned turn is spent finding the thing whose location you already knew.
- Calling the API yourself? Treat
pause_turnas an ordinary continuation, and cap the number of continuations so a loop cannot run away.
messages = [{"role": "user", "content": prompt}]
for _ in range(5): # your own continuation cap
response = client.messages.create(model=MODEL, messages=messages, tools=tools)
if response.stop_reason != "pause_turn":
break
messages.append({"role": "assistant", "content": response.content})
API Error: 400 due to tool use concurrency issues.
Despite the name, this is not a throughput problem, and turning concurrency down is usually not the fix. A 400 means the API rejected the request as malformed before doing any work, and what it rejected is the shape of the conversation you sent.
API Error: 400 due to tool use concurrency issues. Run /rewind to recover the conversation.
Two unrelated causes produce that one string, and they want opposite responses. The tell is when it started.
| Cause | Tell | Fix |
|---|---|---|
| Tool-use settings the model does not support, typically extended thinking | It fails from the first turn, or right after a model change | /model to a compatible model, then /clear |
A broken tool_use and tool_result pairing in the transcript | It started mid-session and now every message fails | /rewind past the break, or /clear |
The second is the one people describe as the conversation being permanently broken, and the mechanism explains why. The API enforces a strict message shape: every tool_use block must be answered by a tool_result carrying the same id, and two messages from the same role cannot sit next to each other. Claude Code resends the whole transcript on every turn, so a single malformed pair poisons every later request. It is not intermittent, it does not age out, and retrying is guaranteed to fail in exactly the same way.
- A turn interrupted while tool calls were still in flight, which leaves a
tool_usewith no matching result. - Automatic compaction rebuilding the history. The reconstruction merges adjacent same-role messages, and it cannot merge a plain-string content field next to an array content field, so the merge is skipped and two same-role messages survive into the request.
- Dictated messages, which are the reported trigger for that merge. Voice input has been stored with
contentas a bare string while typed input is stored as an array of blocks. - Compaction subagents inheriting the broken shape and failing in turn, which is why the failure can outlive the turn that caused it.
If it failed on the first turn, change model and clear
/model
/clear
This is the documented cause: a request combining tool use with a feature the selected model does not support. If you set a default model in settings, check that one too, because a fresh session inherits it and fails identically.
If it started mid-session, rewind past the break
/rewind
Step back to a point before the first failure. This works when one interrupted turn broke the pairing, which is the common case.
If rewinding does not clear it, clear the conversation
/clear
The rejected thing is the history, so discarding the history always works. Commit your tree first: the code is untouched, but the reasoning behind it is about to be.
Questions people ask
What does "You have hit your session limit" mean in Claude Code?
You exhausted the rolling five-hour usage window on your plan. It is shared with Claude chat and Cowork on the same account, and it clears on its own; the message shows the reset time. Because the window covers every model, switching model does not restore access.
How long until my Claude Code limit resets?
The message states the reset time, and /usage shows the same thing with bars for each window. The five-hour window is rolling, so headroom returns gradually as older usage ages out rather than all at once. The weekly window resets at a fixed time assigned to your account.
Why did switching to Sonnet not fix my limit?
Session and weekly windows are shared across all models, so changing model does not give you back allowance you already spent. The one exception is the Opus limit, which is model-specific: after that message, /model to another model keeps you working immediately.
Can I keep working after hitting the limit?
Yes, with usage credits. Run /usage-credits to turn them on or request them from an admin. They bill past the plan allowance at standard API rates. Without them, a subscription blocks until the window reopens.
Why do I keep hitting the weekly limit?
Usually Opus left as the default, a session left open for hours so every turn re-sends the whole conversation, scheduled tasks firing while you are idle, or agent teammates still running. Run /usage and read the breakdown, which flags any behaviour accounting for 10 percent or more of recent usage.
What is the difference between a 429 and a usage limit?
A 429 is throughput on your API key or cloud project, measured per minute, and it resolves in seconds. A usage limit is a plan allowance measured over five hours or a week. A 529 is neither: it means the service is at capacity, which you can sometimes route around with /model.
Does upgrading my plan fix it immediately?
A plan change raises the allowance, but the durable fix is usually model choice and session hygiene, which cost nothing. If you are on Opus by default and never clear between tasks, the next tier buys weeks rather than months.
Does Claude Code usage count against Claude chat usage?
Yes. On Pro, Max, Team, and Enterprise the allowance is shared across Claude Code, Claude chat, and Cowork on the same account, which is why a heavy chat morning can shorten your coding afternoon.
What does "Claude reached its tool-use limit for this turn." mean?
Claude hit the cap on how many tool calls it can make inside one reply. It is not a usage limit and nothing resets: the server-side loop that runs tools such as web search has an iteration ceiling, documented at 10 iterations per request, which the API reports as stop_reason "pause_turn". Click Continue and it resumes in the same conversation with the same context.
How do I stop hitting the tool-use limit for this turn?
Split the work into separate turns with named steps, turn off connectors you are not using, and name the file or repository instead of describing it so Claude spends fewer calls finding it. Three prompts of eight tool calls never reach a ceiling that one prompt of twenty-four does.
Does Claude Code show "Claude reached its tool-use limit for this turn"?
No. That message comes from claude.ai and the Claude desktop app, which run tools inside a server-side loop with an iteration ceiling. The Claude Code CLI drives its own client-side tool loop and does not stop at a fixed tool count per turn. The report that made the string well known was filed against the Claude Code repository and closed as unrelated to Claude Code.
How do I fix "API Error: 400 due to tool use concurrency issues"?
The tell is when it started. If it fails from the first turn or right after a model change, the selected model does not support a tool-use feature you have enabled, usually extended thinking: run /model to switch, then /clear. If it started mid-session and now every message fails, the tool_use and tool_result pairing in the transcript is broken, so run /rewind to a point before the first failure, and /clear if that does not help.
Why does /rewind not fix the tool use concurrency error?
Because /rewind only helps when a single interrupted turn broke the message pairing. When automatic compaction rebuilt the history and left two same-role messages adjacent, the damage is spread through the reconstructed transcript, and reporters have found rewinding 40 messages back still fails while /compact runs the same faulty reconstruction again. Clearing the conversation works, because the rejected thing is the history itself.
Does lowering CLAUDE_CODE_MAX_TOOL_USE_CONCURRENCY fix the 400 error?
Almost never. The variable caps how many read-only tools and subagents execute in parallel, defaulting to 10, and it genuinely helps with a 429 on your key or cloud project. The 400 is a malformed request rather than a throughput problem, so lowering concurrency changes nothing about the message shape the API is rejecting.
Sources
Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.