Free tool · no signup · nothing leaves this page

If you cap AI spend per seat, what stops working?

Every AI budget conversation ends at a number, and nobody in the room can say what that number blocks. This simulates it. Set a cap, and see which developers hit it, on which day and at what hour, how much of the team's week goes dark behind it, and what it actually saves. Then read the trade-off curve and pick the cap that holds.

01 · The simulator

Your team, and the cap

Cap shape

per seat per week

Usage mix
Heavy 15% Regular 55% Light 30%

Heavy runs agents most of the day, often in parallel: 7.5 agent-hours a day. Regular is 3, light is 1. Regular takes whatever is left.

Model mix
Frontier 30% Mid 60% Small 10%

Frontier is Opus-class at $5 in and $25 out per million, mid is Sonnet-class at $2 and $10, small is Haiku-class at $1 and $5. This is the biggest single lever on the bill, and a token cap cannot see it at all.

The cap that holds
$95 per seat, per week

$0uncapped, per week
$0under your cap
0%agent hours dark
$0per agent-hour
Spend saved against work blocked, as the cap moves.rates checked aug 2026

team agent hours blocked share of the uncapped bill spent

What your cap does to each profile.
ProfileSeatsUncapped Under capHit itGoes dark

Per-seat figures. The highlighted row is the profile losing the most hours to your current cap.

Assumptions, and which numbers are derived

The model is deliberately small enough to check by hand. Each seat has a demand rate. That demand accrues evenly across a working week of five days and eight hours. A cap is a ceiling on the period's total, so the hour it fires is simply the share of the period the seat could afford, read off a Monday-to-Friday, 09:00 to 17:00 clock.

  • Published. The model rates. Frontier is Claude Opus 5 at $5 per million input and $25 per million output, mid is Claude Sonnet 5 at $2 and $10, small is Claude Haiku 4.5 at $1 and $5. Cache writes bill at 1.25x input and cache reads at 0.1x. Read from Anthropic's pricing and prompt-caching documentation, August 2026.
  • Published. That spend inside one team varies by roughly 10x between two developers doing similar work. That fact, not the average, is why caps are hard, and it is why this tool models a distribution rather than three point estimates.
  • Derived. Token throughput of about 2.74 million tokens per agent-hour with a warm prompt cache: 230k cache writes, 2.48M cache reads and 32k output, taken from a worked twenty-turn agent session at two sessions an hour. It is the same figure the Claude Code plan picker runs on, so the two tools cannot disagree with each other.
  • Derived. Dollars per agent-hour, which is that token shape priced at the rates above: $3.48 frontier, $1.39 mid, $0.70 small. The ratio between them is exactly 2.5 to 1 to 0.5, which is the same ratio as the list prices, because a fixed token shape priced at three rate cards can be nothing else.
  • Derived. The three archetypes: light at 1 agent-hour a day, regular at 3, heavy at 7.5. Heavy exceeds a working day because parallel sessions count separately. At a regular profile on mid-tier models this lands near $4 per developer per active day, which sits under the $13 figure Anthropic reports across enterprise deployments, as it should: that average includes frontier-heavy usage.
  • Derived. The spread inside one archetype. Each profile is five equal-weight buckets at 0.64, 0.80, 0.94, 1.13, 1.49 times its mean, so the 90th percentile seat spends 2.33x the 10th. Combined with the archetype spread that is roughly 17x between the lightest and heaviest seat on the team. There is no randomness anywhere: the same inputs always give the same answer.
  • Judgement call. The recommended cap is the smallest one that leaves at most 0.5% of the team's agent hours dark. Not zero. A cap that can never fire is not a cap, it is a number in a document, and the whole value of the control is that it bounds your worst week rather than your average one.
  • Simplification. Demand is spread evenly across the week. A real Monday-heavy week fires earlier than the hour shown, and a team that ships on Thursdays fires later. Treat the day and hour as the shape of the answer, not the minute.
  • Simplification. Blocked work is counted as blocked, full stop. In practice a developer whose cap fires does not stop working, they stop working with an agent, which is a productivity cost rather than an outage. The dark-hours figure is the size of that cost, not a claim that the person went home.
  • Not modelled. Prepaid subscription seats. Everything here is metered spend, because a cap on a prepaid plan does not change an invoice, it only rations capacity you already bought. If most of your bill is subscriptions, the first move is not a cap, it is splitting the bill into prepaid and metered.
  • Not modelled. Batch discounts, long-context surcharges, cache-miss storms on cold repos, and the raise-request path that most real cap policies need in order to survive contact with a deadline.

Rates checked August 2026. Providers reprice; check before you commit a budget line to one.

02 · How caps behave in practice

A cap is a schedule, not a number.

The budget meeting produces a dollar figure. What the team experiences is a specific Thursday afternoon when the tool stops answering. Those are the same decision, and only one of them ever gets discussed.

The shape every cap has

Plot savings against a rising cap and you get the same curve every time, because the underlying spend is a long tail. Over most of its range a cap saves almost nothing: it sits above everybody, and a limit nobody reaches is a limit that costs nothing and returns nothing. Then over a narrow band it begins clipping the tail, and this is where a cap earns its keep, because the seats it touches are touched on a Friday afternoon at the end of an unusual week. Push lower and the curve turns: you start taking whole afternoons off developers who were working normally, and the savings that arrive are paid for in other people's time.

The practical consequence is that the interesting range is narrow and it is not where intuition puts it. Teams reliably guess the cap by taking the average spend and adding a margin, which lands well inside the band where regular developers get blocked, because the average seat is nowhere near the middle of the distribution. Spend inside a single team varies by around 10x between two developers doing similar work. Set the cap for the tail, not for the mean.

Caps and alerts are not two strengths of the same control

An alert tells you money was spent. A cap declines the request. That is not a configuration difference, it is a question of where the control physically sits: only something in the request path can refuse a call. A FinOps platform reading invoices, a dashboard reading usage exports, a weekly spreadsheet, all of these are alerts no matter what the settings screen calls them. So is a provider console limit that a developer holding their own key can raise. If the vendor is not between the developer and the model, the cap is advisory.

This is worth saying plainly because the two controls fail in opposite directions. An alert never blocks work and never bounds the bill. A hard cap bounds the bill exactly and will, eventually, block someone mid-task. A serious policy runs both: an alert at a threshold people can act on, and a hard stop far enough above it that reaching the stop is genuinely unusual.

Per seat, or pooled

A per-seat cap is fair and wasteful. It stops precisely the person who overspent, leaves everyone else untouched, and strands the unused headroom of every light user on the team. A pooled cap recovers that headroom, which is why it buys more work per dollar, and it fails as a group: three heavy developers spend the month's pool by the eighteenth and seventeen light ones are blocked for the rest of it, having done nothing unusual. Pooled without a per-seat sub-limit converts one person's heavy week into a team-wide outage, so pool the budget but keep a ceiling underneath it.

A token cap is the third shape and it trades fairness for predictability. It is stable against repricing and it is the same number for everyone, which makes it easy to write down and easy to defend. It is also blind to which model you chose, so a developer on a small model burns it at exactly the rate of one on a frontier model costing five times as much. A token cap therefore taxes the cheap behaviour and subsidises the expensive one. Cap dollars when you want a bounded invoice; cap tokens when the tokens are already paid for and what you want is bounded throughput.

What makes a cap survivable

Every cap that survives contact with a deadline has a way out of it. A hard stop with no raise path does not produce discipline, it produces a developer expensing a personal subscription, which is the same money leaving the company with none of the visibility. The controls that hold are the boring ones: a warning threshold people see coming, a request that reaches an admin in one click, and a partial approval so the answer can be sixty dollars rather than yes or no.

03 · Questions

The ones that decide the number.

Longer versions live in AI spend management and AI cost allocation.

Set it above your heaviest developer's normal week, not at the team average. Spend inside a single team varies by around 10x between two developers doing similar work, so a cap set at the mean blocks the top third of the team by Wednesday while saving very little, because the mean seat was never going to reach it. The useful rule is to cap the tail: pick the smallest number that leaves your heaviest seat with a full week and clips only the outlier weeks above it. On the worked model above, a 20 person team at a regular-heavy mix on mid-tier models lands somewhere between $50 and $80 per seat per week.

An alert is a notification after the money is gone. A cap refuses the request. The distinction is not a matter of configuration, it is a matter of where the control sits: only a vendor in the request path can decline a call, so a FinOps tool reading invoices or a dashboard reading usage exports can never do more than tell you afterwards. A developer holding their own API key can be observed but not capped by anyone except the provider that issued the key.

Per seat is fairer and pooled is cheaper, and the failure modes are opposite. A per-seat cap stops exactly the person who overspent and leaves everyone else working, but it wastes the headroom of light users who never approach their limit. A pooled cap uses that headroom, which is why it buys more work per dollar, but when it fires it fires for everyone: the three heavy developers spend the pool and the seventeen light ones get blocked for the rest of the month through no action of their own. Pooled caps need a per-seat sub-limit underneath them or a raise-request path, or they convert one person's heavy week into a team-wide outage.

They are more predictable and less fair. A token cap is stable against price changes and it is the same number for everyone, which makes it easy to explain. It is also blind to model choice: a developer running a small model burns the cap at the same rate as one running a frontier model that costs five times as much, so a token cap punishes the cheap behaviour and subsidises the expensive one. If the goal is a bounded invoice, cap dollars. If the goal is bounded throughput on a prepaid plan where the marginal token is already paid for, cap tokens.

Only if you set it below what the team actually spends, which the trade-off curve above makes visible before you commit. Caps have a distinctive shape: for most of their range they save almost nothing, because nobody is near them, then over a narrow band they start clipping the tail, then they start taking whole afternoons. The band worth living in is the one where the blocked-hours line has just lifted off zero. Below it you are buying savings with other people's Thursdays.

04 · The cap, for real

This simulated a cap.
Continuum enforces one.

Weekly caps per person, per team, and org-wide, on a rolling seven days rather than a calendar week. Hosted requests stop at the cap with a 429, because our gateway is in the request path and that is what makes a limit a limit instead of an email. People ask for more from inside the tool, and an admin approves all of it or part of it in one click.

one command · reads local history · sends nothing on its own