Together AI pricing: tokens, GPUs, batch, and the free tier

Together AI bills on five different units: per token for text, per image or megapixel for images, per second for video, per audio minute for transcription, and per GPU hour for compute. This page reproduces the current rate card for each, explains where the batch and reservation discounts apply, and answers the free-tier question honestly.

By the Continuum team. We build a workbench that runs Claude Code, Codex, and their peers, so the model rates quoted here are the ones our own cost analytics ship with.

The short version

Serverless text starts at $0.14 in / $0.28 out per million tokens for DeepSeek V4 Flash and runs to $3.00 / $15.00 for Kimi K3, with cached input discounted heavily on most models. Batch inference is up to 50 percent cheaper on a 24-hour window. GPU clusters are $3.99 per H100 hour on demand and $3.19 reserved for 91 to 180 days (181+ day terms are quote-only); dedicated inference is $5.49 per H100 hour. Fine-tuning starts at $0.48 per million training tokens. There is no published ongoing free tier as of August 2026.

What you need to know
  • Text rates sit at market: GLM 5.2 at $1.40 / $4.40, identical to Fireworks and Baseten.
  • Cached input is the biggest per-token lever: DeepSeek V4 Flash reads at $0.03 against $0.14 uncached.
  • Batch is up to 50 percent off on a 24-hour window and scales to 30 billion tokens per job.
  • Reserving GPUs for 91 to 180 days cuts H200 from $5.99 to $3.99 per hour.
  • Dedicated inference costs more than raw cluster compute on the same GPU. That gap is the managed serving stack.
  • No published free tier. Plan to spend on evaluation.

Serverless text rates

All figures per million tokens, read from the Together pricing page in August 2026. Together reprices as models are added and retired, so verify before budgeting.

Together serverless text pricing per million tokens.checked aug 2026
ModelInputCached inputOutput
Kimi K3$3.00$0.30$15.00
Qwen3.8-2.4T-A95B$2.50$0.50$6.25
GLM-5.2$1.40$0.26$4.40
DeepSeek V4 Pro 0813$1.32$0.13$3.96
Llama 3.3 70B$1.04Not listed$1.04
Gemma 4 31B$0.39Not listed$0.97
MiniMax M3$0.30$0.06$1.20
DeepSeek V4 Flash 0731$0.14$0.03$0.28

That is the one genuine price difference we found in the open-model tier. On GLM 5.2 and Kimi K3, Together, Fireworks, and Baseten quote identical rates, so choosing between GLM 5.2 and Kimi K3 is a model-quality question rather than a provider one.

Batch, media, and embeddings

Batch inference

Together documents batch as up to 50 percent lower cost for asynchronous workloads with a 24-hour processing window, scaling to 30 billion tokens in a single job. That is the largest available discount on the platform and applies without changing model or prompt. Anything that does not need an answer this second belongs here.

Non-text rates

Media and embedding rates.checked aug 2026
WhatModelRate
ImageFLUX.2 [pro]$0.03 per image
ImageFLUX.1 [schnell], 4 steps$0.0027 per image
ImageIdeogram 4.0$0.06 per image
ImageFLUX.2 [max]$0.070 per megapixel
VideoByteDance Seedance 2.5$0.115 per video
VideoFLUX 3$0.17 per video
VideoGoogle Veo 3.0$1.60 per video
TranscriptionWhisper Large v3$0.0015 per audio minute
TranscriptionNVIDIA Nemotron 3.5 ASR$0.0045 per audio minute
EmbeddingsMultilingual e5 large instruct$0.02 per million tokens

Sandbox and storage

Sandbox and storage rates.
ItemRate
Code Sandbox, vCPU$0.0446 per hour
Code Sandbox, RAM$0.0149 per GiB-hour
Code Interpreter session, 60 minutes$0.03
Shared filesystem$0.16 per GiB-month

GPU clusters, dedicated endpoints, and fine-tuning

GPU rates per GPU hour.checked aug 2026
GPUCluster, on demandCluster, reserved 91 to 180 daysDedicated inference
NVIDIA HGX H100$3.99$3.19$5.49
NVIDIA HGX H200$5.99$3.99Not listed
NVIDIA HGX B200$8.19$6.79$8.99

The dedicated-inference premium is the number people miss. A dedicated H100 endpoint is $5.49 per hour where the raw cluster GPU is $3.99: you are paying $1.50 per GPU hour, about 38 percent, for Together to run and scale the inference server. That is a reasonable price for not operating it yourself, but it must be in the model.

When a dedicated endpoint beats serverless, worked at GLM 5.2 rates.
Dedicated H100:  $5.49/hour  = $4,013/month running 24/7
Serverless GLM 5.2 input: $1.40 per million tokens

Break-even (input only): 4,013 / 1.40 = ~2.87 BILLION input tokens/month
                                     = ~1,100 input tokens/second, sustained

Verdict: a single dedicated H100 almost never beats serverless on
cost alone. Buy dedicated for isolation, a private model, predictable
latency, or data-residency reasons. Not to save money.

Fine-tuning

Fine-tuning per million training tokens.checked aug 2026
Base modelSupervisedDPOFull-parameter
Standard, up to 16B$0.48$0.54$1.20 to $1.35
Specialized (DeepSeek V4 Flash class), LoRA$6.00 to $40$15.00 to $100Contact sales

What a coding-agent month costs at these rates

The same workload we use across this cluster, so the numbers are comparable: one developer running an agentic coding assistant, roughly 40 agent turns a day over 22 working days, with about 25,000 input and 1,500 output tokens per turn once the system prompt, repo context, and transcript are counted. That is 880 turns, 22M input tokens, and 1.32M output tokens a month.

Monthly cost at Together rates, August 2026. The cached column assumes 80 percent of input hits the prompt cache.
ModelNo cachingWith 80% cache hits
DeepSeek V4 Flash$3.45$1.51
MiniMax M3$8.18$3.96
DeepSeek V4 Pro$34.27$13.32
GLM-5.2$36.61$16.54
Kimi K3$85.80$38.28

Two levers dwarf the provider choice on this workload. Moving anything asynchronous to batch halves it. And routing mechanical turns to DeepSeek V4 Flash while reserving GLM-5.2 or a frontier model for genuinely hard edits is worth 25x, which is the whole spread of the table above.

The free tier question

The honest answer, checked in August 2026: Together publishes no ongoing free tier or self-serve signup credit on its pricing page. The earlier signup credit was retired. What remains is an application-only startup accelerator offering credits, which is a business-development program rather than a trial.

That puts Together at the strict end of the market. For comparison:

Free allowances across the category.checked aug 2026
ProviderWhat you get free
Modal$30 of compute credits per month, renewed, on the $0 Starter plan
GroqA free tier with rate limits, no card required
OpenRouter50 free-model requests/day, 1,000/day after $10 in credits
Fireworks$1 in credits, once
BasetenCredits on signup, amount not published
TogetherNothing published

Questions people ask

How much does Together AI cost per million tokens?

It depends on the model. In August 2026: DeepSeek V4 Flash at $0.14 in / $0.28 out, MiniMax M3 at $0.30 / $1.20, DeepSeek V4 Pro at $1.32 / $3.96, GLM-5.2 at $1.40 / $4.40, Qwen3.8 at $2.50 / $6.25, and Kimi K3 at $3.00 / $15.00. Cached input is discounted heavily on most models, down to $0.03 on DeepSeek V4 Flash.

Does Together AI have a free tier?

No ongoing free tier is published on its pricing page as of August 2026. The earlier signup credit was retired, and what remains is an application-only startup credits program. If you want to evaluate open models at zero cost, Modal renews $30 of credits monthly on its free plan and Groq offers a rate-limited free tier.

How big is the Together AI batch discount?

Up to 50 percent off serverless rates for asynchronous workloads on a 24-hour processing window, and Together advertises single jobs scaling to 30 billion tokens. It is the largest discount on the platform and needs no model or prompt change, so any workload that can tolerate a delay should use it.

How does Together AI dedicated endpoint pricing work?

You pay per GPU hour for reserved instances running your model, at $5.49 for an H100 and $8.99 for a B200. That is a premium over raw cluster compute on the same silicon ($3.99 and $8.19 respectively) because Together operates the serving stack. A single dedicated H100 running 24/7 costs about $4,013 per month, which needs roughly 2.9 billion input tokens a month at GLM 5.2 rates, about 1,100 input tokens per second sustained, to beat serverless on cost alone.

What are Together AI GPU cluster prices?

On demand per GPU hour: HGX H100 at $3.99, H200 at $5.99, B200 at $8.19. Reserved for 91 to 180 days: $3.19, $3.99, and $6.79. The reservation discount is 20 percent on H100 and 33 percent on H200, which is the largest saving available on the compute side.

Is Together AI cheaper than Fireworks or Baseten?

On most models, no: all three quote identical prices for GLM 5.2 and Kimi K3. The one real difference we found in August 2026 is DeepSeek V4 Pro, where Together is $1.32 / $3.96 against $1.74 / $3.48 at both rivals. That makes Together cheaper for input-heavy workloads like coding agents and slightly more expensive for output-heavy ones.

Sources

Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.

  1. Together AI pricing all per-model, GPU, fine-tuning, sandbox, and storage rates
  2. Together AI pricing docs billing units and the batch discount window
  3. Modal pricing comparison free-tier credits
  4. Fireworks serverless pricing comparison per-model rates
  5. Baseten pricing comparison per-model rates
Try it

Run every agent
from one place.

Continuum drives Claude Code, Codex, and peers under your own subscriptions, with live quota gauges and spend by repo. The app is free. Mac is stable; Windows and Linux desktop are beta.

free app · your subscriptions · local-first