Serve it yourself, or buy the tokens. Pick a model class and a monthly volume, and see the break-even against a real per-token rate, with the GPU count, the idle hours, and the engineering time all on the page instead of in a footnote.
The cheapest reserved H100 on the board. You pay all 730 hours whether or not you send it anything.
Below about 9.8 billion tokens a month the API wins; 2 replicas on together h100 cluster, reserved 91 to 180 days break even at 9.8 billion with 0.25 of an SRE.
At 5.0 billion tokens a month, buying tokens from Llama 3.3 70B (Together) costs $4,862 a month less than the smallest self-hosted deployment that would serve this load. The floor is 2 replicas, because one replica means a deploy is an outage.
effective load 1.0B out + 4.0B in / 10 = 1.4B decode-equivalent peak 541 tok/s mean x 1.15 = 623 tok/s replicas max(2, ceil(623 / 1,850)) = 2 x 1 GPU self-host 1,460 replica-hr x $3.19 + $5,417 eng = $10,074 api 4.0B x $1.04 + 1.0B x $1.04 per 1M = $5,212
| Option | Unit rate | Replica-hours | Per month |
|---|---|---|---|
| Self-host: Mid, 30 to 70B dense or small MoE2 replicas × 1 GPU on Together H100 cluster, reserved 91 to 180 days | $3.19 / replica-hr | 1,460 | $10,074 |
| API: Llama 3.3 70B (Together)$1.04 in / $1.04 out per million | per token | – | $5,212 |
Replica-hours are billed hours, not wall-clock hours. Reserved capacity bills all 730 hours in the month; per-second platforms bill only while a container runs.
| Model | Input / 1M | Output / 1M | Per month |
|---|---|---|---|
| GLM-4.7-FlashX | $0.070 | $0.40 | $682 |
| DeepSeek V4 Flash | $0.44 | $1.32 | $3,087 |
| GLM-4.7 | $0.60 | $2.20 | $4,611 |
| Qwen3.5 397B (Together) | $0.60 | $3.60 | $6,014 |
| Kimi K2.7 Code | $0.95 | $4.00 | $7,819 |
| DeepSeek V4 Pro | $1.32 | $3.96 | $9,262 |
These are different models with different capability, not discounts on the same one. Only the two Together rows are marked open weights by the shared rate card; the rest are first-party APIs from labs that publish open models, so read them as the price of comparable capability rather than proof of an identical licence.
GPU rates read 2026-08-20. Token rates read 2026-08-19.
US dollars per GPU per hour, read off each provider's own live pricing page on 20 August 2026. Modal publishes per second; the hourly figure is that number multiplied by 3,600 and is ours. Two things move on a date: Together's reserved rates depend on the term you commit to, and Fireworks raises every on-demand rate on 1 September 2026.
| Source | HBM | $ / GPU-hour | Billing |
|---|---|---|---|
| Together H100 cluster, reserved 91 to 180 days | 80 GB | $3.19 | Reserved, all 730 hours |
| Modal H100 SXM5 | 80 GB | $3.95 | Per second, scales to zero |
| Together H100 cluster, on demand | 80 GB | $3.99 | Per hour |
| Together H200 cluster, reserved 91 to 180 days | 141 GB | $3.99 | Reserved, all 730 hours |
| Together dedicated inference H100 | 80 GB | $5.49 | Per hour, Together runs the server |
| Together H200 cluster, on demand | 141 GB | $5.99 | Per hour |
| Fireworks H100, on demand | 80 GB | $7.00 | Per second, rises to $8.00 on 1 Sep 2026 |
Sources: Together pricing, Modal pricing, Fireworks pricing. Modal bills CPU and memory separately at $0.0000131 per core-second and $0.00000222 per GiB-second, which is $0.89 an hour for eight cores and 64 GiB and is included in the calculator. Longer context on all three is in serverless inference.
Five things reliably turn a self-hosting proposal that looked cheap into an invoice that is not. Each one is in the calculator above, which is the only reason its numbers are less flattering than the ones in most business cases.
A reserved GPU costs exactly as much at 3am as it does at your busiest minute. Together's cheapest reserved H100 is $3.19 an hour, which is $2,329 a month per card, and that number does not move when your traffic does. So the only figure that matters is what fraction of those 730 hours you actually use. At full utilization a reserved H100 is very cheap compute. At 20 percent it is five times its sticker price, and at 20 percent almost nothing beats a per-token API where idle costs zero because the provider amortizes one warm model across every customer at once. That is a structural advantage of the token API, not a discount it chose to give you, and no single-tenant deployment can reproduce it.
You provision for the peak and you pay for the trough. A workload with an eight-to-one peak-to-mean ratio needs eight times the capacity of its average load, and on reserved hardware roughly 88 percent of those GPU hours are idle. The per-token API bills the mean because it never sees your shape. This is the single largest swing in the calculator: switch the traffic shape from steady to bursty and watch the break-even move by most of an order of magnitude without touching the volume. A per-second platform like Modal is the honest answer to a bursty shape, which is why its premium over a reservation is not really a premium at all below about 80 percent utilization.
The calculator refuses to quote fewer than two replicas, and that floor is the difference between a benchmark and a service. With one replica every deploy is an outage, every crash is an outage, and every model upgrade is a maintenance window announced to your customers. Two is the minimum honest shape, and it doubles the GPU line before any traffic exists. Business cases that quote a single card are comparing a laptop to a product.
A quarter of an engineer at a fully loaded $260,000 a year is $5,417 a month, which is more than two reserved H100s. That quarter is not optional: somebody owns the vLLM upgrade, the autoscaler, the KV-cache tuning that made the 70B fit on one card, the pager, and the incident review after the first outage. Turn the engineering toggle off in the calculator and the break-even roughly halves, which is precisely how a business case gets written and precisely why it does not survive contact with the second quarter.
CPU and memory bill separately on a serverless GPU platform: Modal adds about $0.89 per replica-hour for eight cores and 64 GiB, roughly 23 percent on top of the H100 itself. Then there is egress, the load balancer, observability, model storage, and the concurrency quota you did not know you needed until peak demanded 144 containers and your plan allowed 50. None of these make self-hosting cheaper. Every unmodelled cost in this category pushes the break-even in the same direction, so treat any number the calculator gives you as the optimistic end of a range.
If your token volume comes from engineers running coding agents rather than from product traffic, this whole page is the wrong question. Nobody self-hosts an open model to run Claude Code, and per-token billing on a frontier model is not what the calculator above is comparing. That case is a per-seat decision: a flat plan prepays a volume that would cost considerably more metered, and on a plan one extra turn costs nothing at the margin. The pricing calculator prices that comparison, and the coding-model leaderboard is where the model choice actually gets made.
The answers that change what the spreadsheet should say.
Below a few billion tokens a month, almost never. A reserved H100 at Together is $3.19 an hour, or $2,329 a month whether or not you send it anything, and a production deployment needs at least two replicas so a deploy is not an outage. That is $4,657 a month of GPU before a single token moves, against $5,200 for five billion tokens of Llama 3.3 70B at Together's serverless rate. Add a quarter of an engineer at a loaded $260,000 a year and the break-even moves out to roughly ten billion tokens a month.
Two constraints, and the larger one wins. Weights must fit: at FP8 a 70B model is about 70 GB and fits one H100 80GB with careful tuning, while a 397B mixture-of-experts model is about 397 GB and needs eight H100s or four H200s. Then throughput must cover your peak: a measured vLLM deployment of Llama 3.3 70B at FP8 on one H100 reached 1,850 output tokens per second at 50 concurrent requests, so peak load divided by that figure gives the replica count. Round up, round up again for the availability floor of two, and remember that H200's 141 GB can halve the cards per replica on a large model.
For a 70B model at a steady load on two reserved H100s with a quarter of an SRE, it lands near ten billion tokens a month. Three things move it a long way: how bursty your traffic is, whether your billing scales to zero, and how much engineering time you count. A bursty eight-to-one workload on reserved capacity pushes it much higher because you provision the peak and pay the trough. Turning the engineering toggle off roughly halves it, which is why so many business cases quote a number half the size of the real one.
Provider leaderboards measure single-stream speed: one request, one stream, how fast it comes back. A served deployment runs continuous batching, so aggregate throughput is many times that. The same measured vLLM run gives 120 tokens per second at one concurrent request and 1,850 at fifty, on the same card. Sizing a fleet on the single-stream number overstates the GPU count by an order of magnitude, and sizing it on a vendor's marketing number does something worse. Only the mid class in this tool uses a measured figure; the small and large classes are derived from it and labelled so on every render.
The second replica, the CPU and memory a serverless GPU platform bills separately (about $0.89 per replica-hour on Modal for eight cores and 64 GiB), the platform fee, the load balancer, observability, model storage, and a real share of an engineer to own upgrades and the pager. If you also put a gateway in front of it, that is another tier-one service in the hot path of every model call: the self-hosted gateway guide is honest about what that costs. Every one of these pushes the break-even up, so the calculator's answer is the optimistic end of the range.
The bill that starts this conversation usually comes from coding agents, not from serving models, and that bill is fixed by caching and model choice rather than by infrastructure. Continuum reads the Claude Code and Codex session files already on your machine and reports cost per repo, per model, and per day, with nothing in the request path to operate or secure. If flat billing is what you actually want, hosted inference starts at $25 a month with a weekly allowance.
free app · your own keys · nothing leaves your machine