GPU rates run from $0.000164 per second on a T4 to $0.001972 on a B300. CPU is $0.0000131 per physical core-second and memory is $0.00000222 per GiB-second, billed on top. Volumes cost $0.09 per GiB-month with the first 1 TiB free. Plans: Starter at $0 with $30 monthly credits, 10 concurrent GPUs and 100 containers; Team at $250 per month with $100 credits, 50 concurrent GPUs and 5,000 containers; Enterprise custom. Per-second billing beats a reserved H100 until roughly 80 percent utilization.
- Everything is per second: GPU, CPU core, and GiB of memory, billed separately.
- H100 SXM5 at $0.001097/sec is about $3.95/hour of actual runtime.
- The Starter plan is $0/month with $30 of credits, renewed monthly.
- Team is $250/month including $100 of credits, so the net platform fee is $150.
- Volumes are $0.09/GiB-month with 1 TiB free.
- Break-even against a reserved H100 sits near 80 percent utilization.
The GPU rate card
Read from the Modal pricing page in August 2026. The per-second column is Modal's published figure; the hourly column is that number times 3,600 and is ours, provided because every competitor quotes hourly and comparison is otherwise impossible.
| GPU | Per second | Per hour (derived) |
|---|---|---|
| Nvidia B300 | $0.001972 | $7.10 |
| Nvidia B200 | $0.001736 | $6.25 |
| Nvidia H200 SXM | $0.001261 | $4.54 |
| Nvidia H100 SXM5 | $0.001097 | $3.95 |
| Nvidia RTX PRO 6000 | $0.000842 | $3.03 |
| Nvidia A100, 80 GB | $0.000694 | $2.50 |
| Nvidia A100, 40 GB | $0.000583 | $2.10 |
| Nvidia L40S | $0.000542 | $1.95 |
| Nvidia A10 | $0.000306 | $1.10 |
| Nvidia L4 | $0.000222 | $0.80 |
| Nvidia T4 | $0.000164 | $0.59 |
CPU, memory, and storage
These are billed in addition to the GPU, and forgetting them is the most common way a Modal estimate comes in low.
| Resource | Rate | Per hour (derived) |
|---|---|---|
| CPU, per physical core (2 vCPU equivalent) | $0.0000131 per second | $0.047 |
| Memory, per GiB | $0.00000222 per second | $0.008 |
| Volumes | $0.09 per GiB-month | First 1 TiB per month free |
Plans and free credits
| Starter | Team | Enterprise | |
|---|---|---|---|
| Monthly fee | $0 | $250 | Custom |
| Included credits | $30 per month | $100 per month | Custom |
| Concurrent GPUs | 10 | 50 | Custom |
| Containers | 100 | 5,000 | Custom |
The Starter allowance is the standout number in this whole comparison: $30 renewed every month, on a plan that costs nothing. Fireworks gives $1 once. Together publishes nothing. What $30 a month buys, at August 2026 rates:
| GPU | Hours per month on $30 |
|---|---|
| T4 | about 50.8 |
| L4 | about 37.5 |
| A10 | about 27.2 |
| A100 80GB | about 12.0 |
| H100 SXM5 | about 7.6 |
| B200 | about 4.8 |
The Team plan is worth reading carefully. It costs $250 and includes $100 of credits, so the true platform fee is $150 per month. What you buy for that is concurrency: 50 simultaneous GPUs against 10, and 5,000 containers against 100. If your workload is a large fan-out (embedding a corpus, running an eval sweep) the concurrency limit binds long before the credit does.
Worked example: a fine-tuning run
A LoRA fine-tune on a single H100, four hours of wall clock, with 8 CPU cores and 64 GiB of memory attached.
Runtime: 4 hours = 14,400 seconds
GPU 14,400 s x $0.001097 = $15.80
CPU 14,400 s x 8 cores x $0.0000131 = $1.51
Memory 14,400 s x 64 GiB x $0.00000222= $2.05
------------
TOTAL $19.35
Two comparisons make that number meaningful. Fireworks managed training would bill the same job per million training tokens instead: $0.50 per million for LoRA SFT on a model up to 16B, so a 30-million-token run is $15.00 with no infrastructure to think about. And a reserved H100 from Together at $3.19 per hour is $12.76 for four hours, but only if you already have the reservation and have amortized the other 726 hours in the month.
Worked example: an embedding batch
Embedding a corpus is the workload Modal is best at, because it is bursty, parallel, and finished when it is finished.
Shape: 20 containers x 540 s each = 10,800 GPU-seconds on L4
Each container: 4 CPU cores, 16 GiB
GPU 10,800 s x $0.000222 = $2.40
CPU 10,800 s x 4 cores x $0.0000131 = $0.57
Memory 10,800 s x 16 GiB x $0.00000222 = $0.38
-----------
TOTAL $3.35
Wall clock: 9 minutes (the containers ran in parallel)
Covered by the free $30 monthly credit, about 9 times over.
The same job on a single L4 you rented would take three hours of wall clock and cost about the same, because per-second billing means parallelism is free. That is the actual pitch: you are buying the scheduler, not the discount.
Modal against renting raw GPUs
The comparison people want is Modal versus just renting a GPU. Here it is with real numbers.
| Provider | Effective $/hour | Billing granularity |
|---|---|---|
| DeepInfra | $2.20 | Per hour |
| Together, reserved 91 to 180 days | $3.19 | Reserved |
| Modal | $3.95 | Per second, scales to zero |
| Together, on-demand cluster | $3.99 | Per hour |
| Together, dedicated inference | $5.49 | Per hour |
| Baseten, dedicated | $6.50 | Per minute |
| Fireworks, on-demand | $7.00 (rises to $8.00 on 1 Sep 2026) | Per second, quoted hourly |
Reserved H100 (Together, 181+ days): $3.19/hr x 730 hr = $2,328.70/month
Modal H100: $3.95/hr of ACTUAL runtime
Break-even runtime = $2,328.70 / $3.95 = ~590 hours/month
= ~81% of a 730-hour month
Below ~80% utilization, Modal is cheaper.
Above it, buy the reservation.
Questions people ask
How much does Modal cost per hour?
Modal quotes per second, so the hourly figure is the per-second rate times 3,600 and applies only while a container runs. At August 2026 rates that is about $3.95 for an H100 SXM5, $4.54 for an H200 SXM, $6.25 for a B200, $7.10 for a B300, $2.50 for an A100 80GB, $0.80 for an L4, and $0.59 for a T4. CPU at $0.047 per core-hour and memory at $0.008 per GiB-hour are billed on top.
What are Modal GPU prices per second?
B300 $0.001972, B200 $0.001736, H200 SXM $0.001261, H100 SXM5 $0.001097, RTX PRO 6000 $0.000842, A100 80GB $0.000694, A100 40GB $0.000583, L40S $0.000542, A10 $0.000306, L4 $0.000222, and T4 $0.000164. CPU is $0.0000131 per physical core-second and memory is $0.00000222 per GiB-second.
Does Modal give free credits?
Yes. The Starter plan costs $0 per month and includes $30 of compute credits every month, renewed, with 10 concurrent GPUs and 100 containers. That is roughly 50 hours of T4 time, 12 hours of A100 80GB time, or 7.6 hours of H100 time each month. The Team plan is $250 per month and includes $100 of credits with 50 concurrent GPUs and 5,000 containers.
Is Modal cheaper than renting a GPU?
Below about 80 percent utilization, yes. A reserved H100 at $3.19 per hour costs $2,329 a month whether you use it or not; Modal at an effective $3.95 per hour only bills while your container runs, so the break-even is around 590 hours a month. Above that, the reservation wins on price, though you then also own the model server and the on-call rotation.
Does Modal charge for storage?
Volumes are $0.09 per GiB-month, and the first 1 TiB per month is free. For most inference and fine-tuning workloads that free tier covers the model weights and datasets entirely, so storage rarely appears on the bill.
Is Modal cheaper than a per-token API?
For serving a standard open model, almost never. A token API amortizes one warm model across every customer, which no single-tenant deployment can match at low volume. Modal wins when there is no token API for what you are doing: a private model, a fine-tune, non-LLM GPU work, or a batch job where you want 50 containers for nine minutes and then nothing.
Sources
Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.
- Modal pricing per-second GPU, CPU and memory rates, plans, credits, volume pricing
- Modal GPU guide GPU types and multi-GPU syntax
- Together AI pricing reserved and on-demand H100 rates for the break-even
- Baseten pricing per-minute H100 rate for comparison
- Fireworks pricing on-demand GPU rates for comparison