Every AI budget conversation ends at a number, and nobody in the room can say what that number blocks. This simulates it. Set a cap, and see which developers hit it, on which day and at what hour, how much of the team's week goes dark behind it, and what it actually saves. Then read the trade-off curve and pick the cap that holds.
per seat per week
team agent hours blocked share of the uncapped bill spent
| Profile | Seats | Uncapped | Under cap | Hit it | Goes dark |
|---|
Per-seat figures. The highlighted row is the profile losing the most hours to your current cap.
The model is deliberately small enough to check by hand. Each seat has a demand rate. That demand accrues evenly across a working week of five days and eight hours. A cap is a ceiling on the period's total, so the hour it fires is simply the share of the period the seat could afford, read off a Monday-to-Friday, 09:00 to 17:00 clock.
Rates checked August 2026. Providers reprice; check before you commit a budget line to one.
The budget meeting produces a dollar figure. What the team experiences is a specific Thursday afternoon when the tool stops answering. Those are the same decision, and only one of them ever gets discussed.
Plot savings against a rising cap and you get the same curve every time, because the underlying spend is a long tail. Over most of its range a cap saves almost nothing: it sits above everybody, and a limit nobody reaches is a limit that costs nothing and returns nothing. Then over a narrow band it begins clipping the tail, and this is where a cap earns its keep, because the seats it touches are touched on a Friday afternoon at the end of an unusual week. Push lower and the curve turns: you start taking whole afternoons off developers who were working normally, and the savings that arrive are paid for in other people's time.
The practical consequence is that the interesting range is narrow and it is not where intuition puts it. Teams reliably guess the cap by taking the average spend and adding a margin, which lands well inside the band where regular developers get blocked, because the average seat is nowhere near the middle of the distribution. Spend inside a single team varies by around 10x between two developers doing similar work. Set the cap for the tail, not for the mean.
An alert tells you money was spent. A cap declines the request. That is not a configuration difference, it is a question of where the control physically sits: only something in the request path can refuse a call. A FinOps platform reading invoices, a dashboard reading usage exports, a weekly spreadsheet, all of these are alerts no matter what the settings screen calls them. So is a provider console limit that a developer holding their own key can raise. If the vendor is not between the developer and the model, the cap is advisory.
This is worth saying plainly because the two controls fail in opposite directions. An alert never blocks work and never bounds the bill. A hard cap bounds the bill exactly and will, eventually, block someone mid-task. A serious policy runs both: an alert at a threshold people can act on, and a hard stop far enough above it that reaching the stop is genuinely unusual.
A per-seat cap is fair and wasteful. It stops precisely the person who overspent, leaves everyone else untouched, and strands the unused headroom of every light user on the team. A pooled cap recovers that headroom, which is why it buys more work per dollar, and it fails as a group: three heavy developers spend the month's pool by the eighteenth and seventeen light ones are blocked for the rest of it, having done nothing unusual. Pooled without a per-seat sub-limit converts one person's heavy week into a team-wide outage, so pool the budget but keep a ceiling underneath it.
A token cap is the third shape and it trades fairness for predictability. It is stable against repricing and it is the same number for everyone, which makes it easy to write down and easy to defend. It is also blind to which model you chose, so a developer on a small model burns it at exactly the rate of one on a frontier model costing five times as much. A token cap therefore taxes the cheap behaviour and subsidises the expensive one. Cap dollars when you want a bounded invoice; cap tokens when the tokens are already paid for and what you want is bounded throughput.
Every cap that survives contact with a deadline has a way out of it. A hard stop with no raise path does not produce discipline, it produces a developer expensing a personal subscription, which is the same money leaving the company with none of the visibility. The controls that hold are the boring ones: a warning threshold people see coming, a request that reaches an admin in one click, and a partial approval so the answer can be sixty dollars rather than yes or no.
Longer versions live in AI spend management and AI cost allocation.
Set it above your heaviest developer's normal week, not at the team average. Spend inside a single team varies by around 10x between two developers doing similar work, so a cap set at the mean blocks the top third of the team by Wednesday while saving very little, because the mean seat was never going to reach it. The useful rule is to cap the tail: pick the smallest number that leaves your heaviest seat with a full week and clips only the outlier weeks above it. On the worked model above, a 20 person team at a regular-heavy mix on mid-tier models lands somewhere between $50 and $80 per seat per week.
An alert is a notification after the money is gone. A cap refuses the request. The distinction is not a matter of configuration, it is a matter of where the control sits: only a vendor in the request path can decline a call, so a FinOps tool reading invoices or a dashboard reading usage exports can never do more than tell you afterwards. A developer holding their own API key can be observed but not capped by anyone except the provider that issued the key.
Per seat is fairer and pooled is cheaper, and the failure modes are opposite. A per-seat cap stops exactly the person who overspent and leaves everyone else working, but it wastes the headroom of light users who never approach their limit. A pooled cap uses that headroom, which is why it buys more work per dollar, but when it fires it fires for everyone: the three heavy developers spend the pool and the seventeen light ones get blocked for the rest of the month through no action of their own. Pooled caps need a per-seat sub-limit underneath them or a raise-request path, or they convert one person's heavy week into a team-wide outage.
They are more predictable and less fair. A token cap is stable against price changes and it is the same number for everyone, which makes it easy to explain. It is also blind to model choice: a developer running a small model burns the cap at the same rate as one running a frontier model that costs five times as much, so a token cap punishes the cheap behaviour and subsidises the expensive one. If the goal is a bounded invoice, cap dollars. If the goal is bounded throughput on a prepaid plan where the marginal token is already paid for, cap tokens.
Only if you set it below what the team actually spends, which the trade-off curve above makes visible before you commit. Caps have a distinctive shape: for most of their range they save almost nothing, because nobody is near them, then over a narrow band they start clipping the tail, then they start taking whole afternoons. The band worth living in is the one where the blocked-hours line has just lifted off zero. Below it you are buying savings with other people's Thursdays.
Weekly caps per person, per team, and org-wide, on a rolling seven days rather than a calendar week. Hosted requests stop at the cap with a 429, because our gateway is in the request path and that is what makes a limit a limit instead of an email. People ask for more from inside the tool, and an admin approves all of it or part of it in one click.
one command · reads local history · sends nothing on its own