Fireworks AI pricing: every rate, tier, and free credit

Fireworks AI pricing has four separate rate cards: per-token serverless (in three speed tiers), per-hour on-demand GPUs, per-million-training-token fine-tuning, and per-token embeddings. This page reproduces all four as they stood in August 2026 and works through what an agentic coding workload actually costs.

By the Continuum team. We build a workbench that runs Claude Code, Codex, and their peers, so the model rates quoted here are the ones our own cost analytics ship with.

The short version

Serverless models are priced individually, with Priority at roughly 1.25x Standard and Fast at roughly 1.5x. GLM 5.2 is $1.40 in / $4.40 out per million on Standard; Kimi K3 is $3.00 in / $15.00 out; DeepSeek V4 Flash is $0.14 / $0.28. Models without an individual rate are priced by parameter count from $0.10 to $1.20 per million. Batch is 50 percent of serverless. On-demand H100s are $7.00 per hour, rising to $8.00 from 1 September 2026. The free tier is $1 in credits.

What you need to know
  • Three serving tiers per model. Priority is about 1.25x Standard, Fast about 1.5x.
  • Cached input is heavily discounted: GLM 5.2 reads at $0.14 against $1.40 uncached.
  • Unlisted models price by size: $0.10 under 4B, $0.20 for 4 to 16B, $0.90 above 16B.
  • Batch inference is 50 percent of serverless on both input and output.
  • On-demand GPU rates increase on 1 September 2026. H100 goes $7.00 to $8.00, B200 $10.00 to $13.00.
  • Free tier: $1 in credits, then postpaid billing.

Serverless per-token rates

Every figure below is per million tokens and was read from the Fireworks serverless pricing docs in August 2026. Columns are input, cached input, and output. Fireworks reprices frequently, so treat this as a snapshot and check the live page before you build a budget on it.

Fireworks Standard tier, per million tokens.checked aug 2026
ModelInputCached inputOutput
Kimi K3$3.00$0.30$15.00
Kimi K3 US$3.30$0.33$16.50
Kimi K2.7 Code$0.95$0.19$4.00
Kimi K2.6$0.95$0.16$4.00
DeepSeek V4 Pro$1.74$0.145$3.48
DeepSeek V4 Flash$0.14$0.028$0.28
GLM 5.2$1.40$0.14$4.40
GLM 5.1$1.40$0.26$4.40
Qwen 3.8 Max$2.00$0.25$6.00
Qwen 3.7 Plus$0.40$0.08$1.60
MiniMax M3 and M2.7$0.30$0.06$1.20
gpt-oss 120B$0.15$0.015$0.60
gpt-oss 20B$0.07$0.035$0.30
NVIDIA Nemotron range$0.05 to $0.60$0.01 to $0.12$0.20 to $2.40

The Priority and Fast multipliers

The same model costs more on lower-contention capacity. The multipliers are not perfectly uniform, so read the per-model row rather than assuming a flat factor.

Input price per million by tier.checked aug 2026
ModelStandardPriorityFast
Kimi K3$3.00$3.75$4.50
Kimi K2.6$0.95$1.50$2.00
GLM 5.2$1.40$1.75$2.10
GLM 5.1$1.40$2.10$2.80
DeepSeek V4 Pro$1.74$2.61Not offered
DeepSeek V4 Flash$0.14$0.21Not offered
MiniMax M3$0.30$0.45Not offered

Models without an individual rate

Size-bracket pricing for models not on the individual card.
Model sizePer million tokens
Under 4B parameters$0.10
4B to 16B$0.20
Above 16B$0.90
Mixture-of-experts variants$0.50 to $1.20

GPU, fine-tuning, and embedding rates

On-demand GPUs

Fireworks bills GPUs per second and quotes them hourly. The page currently carries two rate cards because a price increase lands on 1 September 2026.

On-demand GPU rates per hour. Both cards read from the pricing page, August 2026.
GPUThrough 31 August 2026From 1 September 2026
H100 80 GB$7.00$8.00
H200 141 GB$7.00$8.00
B200 180 GB$10.00$13.00
B300 288 GB$12.00$15.00
GB300 288 GB$18.00$20.00

Managed fine-tuning

Priced per million training tokens, which is a friendlier unit than GPU hours for anyone who has not sized a training run before.

Fine-tuning, per million training tokens.checked aug 2026
Base model sizeLoRA SFTLoRA DPOFull-param SFTFull-param DPO
Up to 16B$0.50$1.00$1.00$2.00
16.1B to 80B$3.00$6.00$6.00$12.00
80B to 300B$6.00$12.00$12.00$24.00
Above 300B$10.00$20.00$20.00$40.00

Embeddings

Embedding rates, input tokens only.
Model sizePer million input tokens
Up to 150M parameters$0.008
150M to 350M parameters$0.016
Qwen3 8B$0.10

Free tier and rate limits

Fireworks gives $1 in free credits on signup and then moves you to postpaid per-token billing once a payment method is attached. There is no renewing free allowance.

In practice $1 buys roughly 7 million input tokens on DeepSeek V4 Flash, or about 330,000 input tokens on Kimi K3. That is enough to confirm your client authenticates and streams correctly. It is not enough to run an evaluation set, and you should budget $20 to $50 for a real comparison against your current provider.

What a coding-agent month actually costs

Headline rates are useless without a workload. Here is a realistic one: a single developer running an agentic coding assistant most of a working day, which in our own usage data looks like roughly 40 agent turns per day across 22 working days, with about 25,000 input tokens and 1,500 output tokens per turn once the system prompt, repo context, and transcript are counted.

  • Turns per month: 40 x 22 = 880
  • Input tokens: 880 x 25,000 = 22M
  • Output tokens: 880 x 1,500 = 1.32M
Monthly cost for that workload at Fireworks Standard rates, August 2026. The cached column assumes 80 percent of input hits the prompt cache.
ModelNo cachingWith 80% cache hits
DeepSeek V4 Flash$3.45$1.48
MiniMax M3$8.18$3.96
Kimi K2.7 Code$26.18$12.80
GLM 5.2$36.61$14.43
Kimi K3$85.80$38.28

Two more levers on the same workload. Moving anything asynchronous to batch halves it again. And running the mechanical turns on DeepSeek V4 Flash while reserving GLM 5.2 or a frontier model for genuinely hard edits is worth more than any provider switch: the spread between the cheapest and most expensive row above is 25x.

Questions people ask

How much does Fireworks AI cost?

It depends on the model and the serving tier. On the Standard tier in August 2026, DeepSeek V4 Flash is $0.14 in / $0.28 out per million tokens, GLM 5.2 is $1.40 / $4.40, and Kimi K3 is $3.00 in / $15.00 out. Priority costs about 1.25x Standard and Fast about 1.5x. Models without an individual rate are priced by parameter count from $0.10 to $1.20 per million.

What is the Fireworks AI pricing model?

Four separate rate cards. Serverless inference is per million tokens with Standard, Priority, and Fast tiers and a discounted cached-input rate. On-demand GPUs bill per second, quoted hourly. Managed fine-tuning bills per million training tokens. Embeddings bill input tokens only. Batch inference is 50 percent of the serverless rate on both input and output.

Does Fireworks AI have a free tier?

Not an ongoing one. New accounts get $1 in free credits, then billing is postpaid per token once a payment method is added. One dollar buys roughly 3.3 million input tokens on DeepSeek V4 Flash, enough to verify an integration but not to run an evaluation.

How much are Fireworks GPUs per hour?

Through 31 August 2026: H100 80GB and H200 141GB at $7.00, B200 180GB at $10.00, B300 288GB at $12.00, and GB300 288GB at $18.00. From 1 September 2026 those rise to $8.00, $8.00, $13.00, $15.00, and $20.00 respectively. These are well above pure compute markets, where H100s list around $2.20 to $4.00 per hour.

Is Fireworks cheaper than Together AI?

Not meaningfully. In August 2026 both list DeepSeek V4 Flash at $0.14 / $0.28, GLM 5.2 at $1.40 / $4.40, and Kimi K3 at $3.00 / $15.00 per million tokens. The open-model serving market has converged on near-identical list prices, so the choice comes down to catalog, latency tiers, and support rather than headline rate.

Does Fireworks discount batch or cached tokens?

Both. Batch inference is billed at 50 percent of the serverless rate on input and output. Cached input is discounted per model and the discount is large: GLM 5.2 reads cached input at $0.14 against $1.40 uncached, and DeepSeek V4 Pro at $0.145 against $1.74. For agent workloads that resend the same context each turn, the cache rate matters more than the base rate.

Sources

Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.

  1. Fireworks serverless pricing Standard, Priority and Fast per-model rates, size brackets, batch discount
  2. Fireworks pricing GPU hourly rates for both cards, fine-tuning table, embeddings, $1 credit
  3. DeepInfra pricing comparison GPU hourly rates
  4. Modal pricing comparison per-second GPU rates
Try it

Run every agent
from one place.

Continuum drives Claude Code, Codex, and peers under your own subscriptions, with live quota gauges and spend by repo. The app is free. Mac is stable; Windows and Linux desktop are beta.

free app · your subscriptions · local-first