Serverless models are priced individually, with Priority at roughly 1.25x Standard and Fast at roughly 1.5x. GLM 5.2 is $1.40 in / $4.40 out per million on Standard; Kimi K3 is $3.00 in / $15.00 out; DeepSeek V4 Flash is $0.14 / $0.28. Models without an individual rate are priced by parameter count from $0.10 to $1.20 per million. Batch is 50 percent of serverless. On-demand H100s are $7.00 per hour, rising to $8.00 from 1 September 2026. The free tier is $1 in credits.
- Three serving tiers per model. Priority is about 1.25x Standard, Fast about 1.5x.
- Cached input is heavily discounted: GLM 5.2 reads at $0.14 against $1.40 uncached.
- Unlisted models price by size: $0.10 under 4B, $0.20 for 4 to 16B, $0.90 above 16B.
- Batch inference is 50 percent of serverless on both input and output.
- On-demand GPU rates increase on 1 September 2026. H100 goes $7.00 to $8.00, B200 $10.00 to $13.00.
- Free tier: $1 in credits, then postpaid billing.
Serverless per-token rates
Every figure below is per million tokens and was read from the Fireworks serverless pricing docs in August 2026. Columns are input, cached input, and output. Fireworks reprices frequently, so treat this as a snapshot and check the live page before you build a budget on it.
| Model | Input | Cached input | Output |
|---|---|---|---|
| Kimi K3 | $3.00 | $0.30 | $15.00 |
| Kimi K3 US | $3.30 | $0.33 | $16.50 |
| Kimi K2.7 Code | $0.95 | $0.19 | $4.00 |
| Kimi K2.6 | $0.95 | $0.16 | $4.00 |
| DeepSeek V4 Pro | $1.74 | $0.145 | $3.48 |
| DeepSeek V4 Flash | $0.14 | $0.028 | $0.28 |
| GLM 5.2 | $1.40 | $0.14 | $4.40 |
| GLM 5.1 | $1.40 | $0.26 | $4.40 |
| Qwen 3.8 Max | $2.00 | $0.25 | $6.00 |
| Qwen 3.7 Plus | $0.40 | $0.08 | $1.60 |
| MiniMax M3 and M2.7 | $0.30 | $0.06 | $1.20 |
| gpt-oss 120B | $0.15 | $0.015 | $0.60 |
| gpt-oss 20B | $0.07 | $0.035 | $0.30 |
| NVIDIA Nemotron range | $0.05 to $0.60 | $0.01 to $0.12 | $0.20 to $2.40 |
The Priority and Fast multipliers
The same model costs more on lower-contention capacity. The multipliers are not perfectly uniform, so read the per-model row rather than assuming a flat factor.
| Model | Standard | Priority | Fast |
|---|---|---|---|
| Kimi K3 | $3.00 | $3.75 | $4.50 |
| Kimi K2.6 | $0.95 | $1.50 | $2.00 |
| GLM 5.2 | $1.40 | $1.75 | $2.10 |
| GLM 5.1 | $1.40 | $2.10 | $2.80 |
| DeepSeek V4 Pro | $1.74 | $2.61 | Not offered |
| DeepSeek V4 Flash | $0.14 | $0.21 | Not offered |
| MiniMax M3 | $0.30 | $0.45 | Not offered |
Models without an individual rate
| Model size | Per million tokens |
|---|---|
| Under 4B parameters | $0.10 |
| 4B to 16B | $0.20 |
| Above 16B | $0.90 |
| Mixture-of-experts variants | $0.50 to $1.20 |
GPU, fine-tuning, and embedding rates
On-demand GPUs
Fireworks bills GPUs per second and quotes them hourly. The page currently carries two rate cards because a price increase lands on 1 September 2026.
| GPU | Through 31 August 2026 | From 1 September 2026 |
|---|---|---|
| H100 80 GB | $7.00 | $8.00 |
| H200 141 GB | $7.00 | $8.00 |
| B200 180 GB | $10.00 | $13.00 |
| B300 288 GB | $12.00 | $15.00 |
| GB300 288 GB | $18.00 | $20.00 |
Managed fine-tuning
Priced per million training tokens, which is a friendlier unit than GPU hours for anyone who has not sized a training run before.
| Base model size | LoRA SFT | LoRA DPO | Full-param SFT | Full-param DPO |
|---|---|---|---|---|
| Up to 16B | $0.50 | $1.00 | $1.00 | $2.00 |
| 16.1B to 80B | $3.00 | $6.00 | $6.00 | $12.00 |
| 80B to 300B | $6.00 | $12.00 | $12.00 | $24.00 |
| Above 300B | $10.00 | $20.00 | $20.00 | $40.00 |
Embeddings
| Model size | Per million input tokens |
|---|---|
| Up to 150M parameters | $0.008 |
| 150M to 350M parameters | $0.016 |
| Qwen3 8B | $0.10 |
Free tier and rate limits
Fireworks gives $1 in free credits on signup and then moves you to postpaid per-token billing once a payment method is attached. There is no renewing free allowance.
In practice $1 buys roughly 7 million input tokens on DeepSeek V4 Flash, or about 330,000 input tokens on Kimi K3. That is enough to confirm your client authenticates and streams correctly. It is not enough to run an evaluation set, and you should budget $20 to $50 for a real comparison against your current provider.
What a coding-agent month actually costs
Headline rates are useless without a workload. Here is a realistic one: a single developer running an agentic coding assistant most of a working day, which in our own usage data looks like roughly 40 agent turns per day across 22 working days, with about 25,000 input tokens and 1,500 output tokens per turn once the system prompt, repo context, and transcript are counted.
- Turns per month: 40 x 22 = 880
- Input tokens: 880 x 25,000 = 22M
- Output tokens: 880 x 1,500 = 1.32M
| Model | No caching | With 80% cache hits |
|---|---|---|
| DeepSeek V4 Flash | $3.45 | $1.48 |
| MiniMax M3 | $8.18 | $3.96 |
| Kimi K2.7 Code | $26.18 | $12.80 |
| GLM 5.2 | $36.61 | $14.43 |
| Kimi K3 | $85.80 | $38.28 |
Two more levers on the same workload. Moving anything asynchronous to batch halves it again. And running the mechanical turns on DeepSeek V4 Flash while reserving GLM 5.2 or a frontier model for genuinely hard edits is worth more than any provider switch: the spread between the cheapest and most expensive row above is 25x.
Questions people ask
How much does Fireworks AI cost?
It depends on the model and the serving tier. On the Standard tier in August 2026, DeepSeek V4 Flash is $0.14 in / $0.28 out per million tokens, GLM 5.2 is $1.40 / $4.40, and Kimi K3 is $3.00 in / $15.00 out. Priority costs about 1.25x Standard and Fast about 1.5x. Models without an individual rate are priced by parameter count from $0.10 to $1.20 per million.
What is the Fireworks AI pricing model?
Four separate rate cards. Serverless inference is per million tokens with Standard, Priority, and Fast tiers and a discounted cached-input rate. On-demand GPUs bill per second, quoted hourly. Managed fine-tuning bills per million training tokens. Embeddings bill input tokens only. Batch inference is 50 percent of the serverless rate on both input and output.
Does Fireworks AI have a free tier?
Not an ongoing one. New accounts get $1 in free credits, then billing is postpaid per token once a payment method is added. One dollar buys roughly 3.3 million input tokens on DeepSeek V4 Flash, enough to verify an integration but not to run an evaluation.
How much are Fireworks GPUs per hour?
Through 31 August 2026: H100 80GB and H200 141GB at $7.00, B200 180GB at $10.00, B300 288GB at $12.00, and GB300 288GB at $18.00. From 1 September 2026 those rise to $8.00, $8.00, $13.00, $15.00, and $20.00 respectively. These are well above pure compute markets, where H100s list around $2.20 to $4.00 per hour.
Is Fireworks cheaper than Together AI?
Not meaningfully. In August 2026 both list DeepSeek V4 Flash at $0.14 / $0.28, GLM 5.2 at $1.40 / $4.40, and Kimi K3 at $3.00 / $15.00 per million tokens. The open-model serving market has converged on near-identical list prices, so the choice comes down to catalog, latency tiers, and support rather than headline rate.
Does Fireworks discount batch or cached tokens?
Both. Batch inference is billed at 50 percent of the serverless rate on input and output. Cached input is discounted per model and the discount is large: GLM 5.2 reads cached input at $0.14 against $1.40 uncached, and DeepSeek V4 Pro at $0.145 against $1.74. For agent workloads that resend the same context each turn, the cache rate matters more than the base rate.
Sources
Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.
- Fireworks serverless pricing Standard, Priority and Fast per-model rates, size brackets, batch discount
- Fireworks pricing GPU hourly rates for both cards, fine-tuning table, embeddings, $1 credit
- DeepInfra pricing comparison GPU hourly rates
- Modal pricing comparison per-second GPU rates