Catalog breadth and GPU hours on one side, the supervised session that spends them on the other.
Together AI is a full-stack open-model platform. Its catalog advertises 200+ models spanning chat, code, vision, image, video, audio, embeddings, rerank, and moderation, served through serverless per-token endpoints, a Batch API at a 50% discount, provisioned throughput, and dedicated endpoints. Underneath that sits a real GPU cloud: on-demand clusters at $3.99 per H100 hour, $5.99 per H200 hour, and $8.19 per B200 hour, with reserved terms bringing B200 to $6.79. It also fine-tunes, and it does its own research, with the FlashAttention line, speculative decoding, and the Together Kernel Collection behind claims of 2x faster inference and 90% faster pre-training. Customers named on its site include Cursor, Cohere, DeepMind, ElevenLabs, Mozilla, and Salesforce. Continuum is not selling any of that. It is the free workbench where a developer runs Claude Code, Codex, Cursor, Grok, and OpenCode under their own subscriptions, with worktree isolation per session, plan gates, per-hunk diff, PR merge, live quota gauges, and cost by repo. Together is a supplier you can consume; Continuum is where a coding agent consumes it.
Updated 2026-08-03 · Mac stable · Win/Linux desktop beta
Pick Continuum when the consumer of inference is a coding agent and you want the workbench, review loop, and per-repo cost around it, with flat-fee hosted inference instead of a per-token and per-GPU-hour meter.
Pick Together when you need catalog breadth across modalities, cheap batch throughput, dedicated endpoints, GPU clusters you scale to thousands of cards, or fine-tuning on your own data.
Open Claude Code and Codex in separate worktrees on the same repo, each with its own branch and transcript.
Approve the better plan from the phone and interrupt the weaker run without returning to the Mac.
Review the winning diff hunk by hunk, revert one change, open the pull request, and watch its checks.
Read the quota gauge and the day's cost by repo in Usage analytics before the next batch.
Pick candidates from a 200+ model catalog spanning chat, code, vision, image, video, and audio, and benchmark them serverless.
Push everything that can wait through the Batch API at half price, up to 50,000 requests per file with 24-hour best-effort completion.
Fine-tune the winner with LoRA from $0.48 per million training tokens, subject to a $4.00 minimum per job.
Move steady production traffic onto dedicated endpoints, or rent an on-demand GPU cluster and scale it toward thousands of cards.
Install, connect, first session - steps you can run the same day.
Name the consumer first. A product serving models at scale is a Together question. Engineers running coding agents is a workbench question.
For the capacity path, price serverless against your actual prompts. Published rates run from Gemma 3n E4B at $0.06 in and $0.12 out per million to Llama 3.3 70B at $1.04 both ways, so model selection dominates the bill.
If any of that workload can wait, test the Batch API. It is a 50% discount on selected serverless models, accepts up to 50,000 requests or 100MB per batch file, and targets best-effort completion within 24 hours.
If throughput is steady, compare GPU clusters instead: $3.99 per H100 hour, $5.99 per H200 hour, and $8.19 per B200 hour on demand, with reserved terms of 181 days or more bringing B200 to $6.79.
For the coding path, install Continuum, open one worktree session per ticket, and let Plan mode hold the agent read-only until someone approves the plan.
Add a Together key to the OpenCode connector if you want those models in the same workbench. The provider id is togetherai and the variable is TOGETHER_API_KEY.
Compare the flat-fee alternative honestly. Continuum hosted inference is $25, $100, $200, or $500 a month with weekly allowances of $25, $100, $200, and $1,000, which prices risk differently than a meter does.
Score each on its own axis: tokens per second, catalog fit, and cost per million for Together; sessions supervised safely and cost per repository for Continuum.
Cells use product nouns on both sides. Continuum is the multi-provider workbench; they win where their loop is the product.
A GPU hour and a completion are inputs. A workbench decides which agent runs, in which worktree, on which branch, under whose approval, and whether the diff merges. Neither replaces the other, and most teams that need both simply buy both. The useful question is which one is missing today.
Together spans chat, code, vision, image, video, audio, embeddings, rerank, and moderation, which matters enormously if your product is multimodal. Continuum's reach is deliberately narrow: the coding agents developers actually run, with the session, review, quota, and cost surfaces around them. Breadth of catalog and depth of operating loop are orthogonal purchases. If you want the coding scores behind a name on that catalog, the Qwen model hub is one slice of it with a card per checkpoint.
Together's headline lever is the Batch API at 50% off selected models for work that tolerates a 24-hour window, plus reserved GPU terms that bring B200 from $8.19 to $6.79 per hour at 181 days or more. Those are excellent for planned, deferrable load. A coding agent is the opposite: interactive, bursty, and impossible to defer. That is why Continuum's hosted inference prices a flat monthly fee with a weekly allowance instead.
Together's security page did not resolve during this comparison, so this page makes no claim about SOC 2, HIPAA, GDPR, or ISO certification for Together AI in either direction. If certification is a gate on your purchase, ask their team directly and get it in writing rather than trusting any comparison page, including this one. We would rather flag a gap than fill it with an assumption.
Continuum's OpenCode connector carries Together as provider id togetherai with a TOGETHER_API_KEY and 36 models under that id today. So a Together key genuinely runs inside the workbench beside your Claude and Codex sessions, and the same per-repo ledger prices the result. Bring the key; keep the review loop.
Use both, and the Together half works today. Continuum reaches Together through its OpenCode connector: the models.dev provider id is togetherai, the credential is a TOGETHER_API_KEY, and 36 models are carried under that id right now. Add the key once and Together models appear alongside your Claude, Codex, Cursor, and Grok sessions in one workbench, metered by Together and attributed per repo by Continuum. If you would rather not meter at all, Continuum's own hosted inference is the flat-fee path.
$0 for the workbench on Mac, web, iPhone, and Watch, running under the subscriptions and provider keys you already hold, with no cut taken on those sessions. Optional hosted inference is Plus at $25, Max 100 at $100, Max 200 at $200, and Ultra at $500 per month, carrying weekly allowances of $25, $100, $200, and $1,000, reachable at an OpenAI-compatible endpoint and an Anthropic-compatible bare origin with cont_sk_ keys.
Together AI is usage based. Serverless examples published today include Gemma 3n E4B at $0.06 in and $0.12 out per million tokens, Qwen3.5 9B at $0.17 and $0.25, DeepSeek V4 Flash at $0.14 and $0.28, MiniMax M3 at $0.30 and $1.20, and Llama 3.3 70B at $1.04 both ways. The Batch API is a 50% discount on selected serverless models. Dedicated inference runs $5.49 an hour for HGX H100 and $8.99 for HGX B200. GPU clusters are $3.99 per H100 hour, $5.99 per H200 hour, and $8.19 per B200 hour on demand, with reserved terms of 181 days or more at $6.79 for B200. Fine-tuning starts at $0.48 per million training tokens for LoRA SFT up to 16B, with a $4.00 minimum per job.
Different meters, not different prices for the same thing. Continuum charges nothing for BYOK sessions and sells a flat monthly inference allowance; Together sells tokens, batch throughput, GPU hours, and training capacity. A team that needs both pays each for what it does.
Together plans checked 3 August 2026 against their pricing page. Vendors reprice often - if you spot a stale figure, tell us and we will fix it. Continuum never marks up a session it runs under your own login.
Continuum ladder: Free app · Plus $25/mo · Max 100 · Max 200 · Ultra - full pricing. Also analytics, multi-account, devices.
Together AI is a full-stack platform for open models. It offers serverless per-token inference, a Batch API, provisioned throughput, dedicated endpoints and containers, fine-tuning, GPU clusters, sandboxes, and managed storage. Its catalog advertises 200+ models across chat, code, vision, image, video, audio, embeddings, rerank, and moderation, and its research team publishes the FlashAttention line and the Together Kernel Collection.
It is usage based. Serverless ranges from about $0.06 in and $0.12 out per million tokens on Gemma 3n E4B up to $1.04 both ways on Llama 3.3 70B. The Batch API takes 50% off selected serverless models. GPU clusters are $3.99 per H100 hour, $5.99 per H200 hour, and $8.19 per B200 hour on demand, with B200 at $6.79 on 181-day-plus reserved terms. Fine-tuning starts at $0.48 per million training tokens with a $4.00 minimum per job.
Yes. Continuum reaches Together through its OpenCode connector using the models.dev provider id togetherai and a TOGETHER_API_KEY, which carries 36 models today. Add the key once and those models appear alongside your Claude, Codex, Cursor, and Grok sessions in the same workbench, with the same per-repo cost ledger applied.
Only if what you actually wanted was a workbench. Continuum does not serve 200+ models, rent GPU clusters, run batch jobs, or fine-tune. It runs and supervises coding agents. Where the two touch is hosted inference for coding work, and there Continuum's answer is a flat monthly fee with a weekly allowance rather than a per-token meter.
Not really, and that is a property of the workload rather than a criticism. Batch trades latency for a 50% discount with best-effort completion inside 24 hours, which suits evaluation runs, bulk embedding, and offline generation. A coding agent is interactive and bursty, so it cannot use that window. Save batch for the deferrable work and price the interactive work separately.
We could not verify that during this comparison; the security page we checked returned a 404, so this page makes no claim in either direction. Ask Together directly and get the current certification list in writing if compliance gates your purchase.
Free app. Your subscriptions. Optional hosted inference. Mac stable - Windows and Linux desktop are beta.
vendor-neutral · local-first · multi-device