Compare·AI routing gateway
Continuum
Multi-agent workbench
VS
Requesty
AI routing gateway

Continuum vs Requesty

A 5% markup on every token routed, versus a free workbench whose hosted tier is a flat weekly allowance.

Requesty is an AI gateway with an unusually clean pricing story: change your base URL, reach 600+ models across 30+ providers through an OpenAI-compatible API, and pay a 5% markup on base model cost with every feature included. Its site claims 70,000+ developers and 90+ billion tokens per day. The free tier gives free models, 200 requests per day, and routing, caching, fallbacks, spend tracking, and EU data residency at no cost, with $10 in credits on signup. Pay-as-you-go adds bring-your-own-keys, routing policies, spend limits, an MCP gateway, and advanced observability with P50, P90, P95, and P99 latency, TTFT, error rates, and cache impact. Enterprise adds SSO, full RBAC, audit logs, model whitelists, guardrails, and PII detection. The savings claims deserve a careful read: the homepage says prompt caching reduces token costs by up to 90%, while the documentation says up to 80% on repeated prompts, so treat both as ceilings on a favorable workload rather than an expected result. Continuum overlaps only on the inference bill. It is the free workbench where Claude Code, Codex, Cursor, Grok, and OpenCode run under your own subscriptions, each in its own git worktree and branch, with plan gates, per-hunk diff review, and pull requests in session, and optional flat hosted inference from $25 to $500 per month.

Updated 2026-08-03 · Mac stable · Win/Linux desktop beta

Choose Continuum when

Pick Continuum when the workload is coding agents you supervise: worktree isolation, plan gates, diff and PR review, live quota gauges, cost by repo, BYOK-first with no markup, and flat hosted tiers if you want them.

Choose Requesty when

Pick Requesty when an application needs 600+ models behind one endpoint with cost, latency, and availability routing, automatic failover, prompt caching, an MCP gateway, and observability, at a transparent 5% markup.

Snapshot direct alternative
Dimension Continuum Requesty
Primary job Run and supervise coding agents Route application traffic to many models
Model reach Provider agents plus optional hosted router 600+ models across 30+ providers
Billing shape Free app · flat monthly hosted tiers 5% markup on base model cost
Free entry Whole app free with your subscriptions Free models, 200 requests per day, $10 credit
Git isolation Worktree + branch per session Not a repository tool
Review loop Plan gate · diff · PR create/review/merge No review surface
Governance Model policy · weekly caps · member approvals RBAC, model whitelist, guardrails, PII detection
Mac stable · web · iPhone · Watch · Win/Linux desktop beta · free app
01

Route tokens vs run agents

In Continuum

Ship a ticket with two agents and no markup

09:30

Two worktree sessions start on the same repo, one Claude Code and one Codex, each on its own branch under an existing subscription.

13:15

Plan mode holds both read-only. The better plan is approved from the phone and the weaker session is interrupted.

16:00

The diff is reviewed hunk by hunk, a bad turn is rolled back to a checkpoint, and the pull request opens from the session.

17:40

Cost by repo and the live quota gauges decide whether tomorrow's batch runs now or after the window resets.

In Requesty

Cut an application's model bill without rewriting it

09:30

Swap the base URL to Requesty and reach 600+ models through the same OpenAI-compatible request shape.

13:15

Enable auto-caching and a fallback policy, then measure the real hit rate against the up-to-80% and up-to-90% claims.

16:00

Add cost or latency routing and read P95 and TTFT in the dashboard rather than the average response time.

17:40

Reconcile the invoice at 5% over base model cost, and scope Enterprise if SSO, RBAC, audit logs, or PII detection are required.

02

Monday path

Monday with a markup gateway and a flat allowance in the same week

Install, connect, first session - steps you can run the same day.

01

Decide what is calling the model. An application is a Requesty case. A developer supervising coding agents is a Continuum case.

02

If it is an application, start on Requesty's free tier: free models, 200 requests per day, plus the $10 signup credit, and confirm the OpenAI-compatible swap is really one base-URL change.

03

Turn on caching and a fallback policy, then measure the cache hit rate on your own prompts. The up-to-90% and up-to-80% claims are ceilings, not forecasts.

04

Compare cost, latency, and availability routing against a fixed model, and read the P95 and TTFT numbers rather than the average.

05

If it is coding agents, install Continuum and start two worktree sessions on a bounded ticket under the subscriptions already on the machine. That path has no markup because nothing is proxied.

06

Approve a plan, review the diff hunk by hunk, and open the pull request from the same session.

07

Only now compare bills: a week of Requesty at 5% over base model cost against a Continuum hosted tier with a fixed weekly allowance of $25, $100, $200, or $1,000.

08

Check cost by repo, provider, model, and day before committing, and set a weekly spend cap if the answer surprises you.

How Continuum runs and prices agent work

How Continuum keeps agent work supervised and attributable

  • Each managed session launches the provider's real agent under the developer's own subscription, so BYOK work carries no gateway markup.
  • A git worktree and branch are created before the first edit, which is what makes running several agents at once safe.
  • Plan mode stays read-only until approved, then the same session respawns with write permission and the accepted plan intact.
  • The diff pane reads the repository index directly and supports per-hunk stage, revert, and commit before a pull request leaves the worktree.
  • Live gauges read each provider's own rate-limit window, including subscription windows that no metered gateway can see.
  • Local agent logs are priced into tokens and dollars by repo, provider, model, and day, with no proxy in the path.
  • Optional hosted inference adds an OpenAI-compatible endpoint and an Anthropic-compatible bare origin with cont_sk_ keys and an auto model router, priced as a flat monthly tier.
03

Gateway and workbench matrix

Cells use product nouns on both sides. Continuum is the multi-provider workbench; they win where their loop is the product.

Capability Continuum Requesty
Finished coding-agent workbench Mac, iPhone, Watch, web, Windows and Linux beta A gateway and dashboard, not a workbench
Run Claude Code, Codex, Cursor, Grok, OpenCode First-class sessions under your own subscriptions Docs include a Claude Code integration guide; it serves the model, not the session
Model catalog breadth Provider agents plus optional hosted auto router 600+ models across 30+ providers per requesty.ai
Per-session git worktree and branch Isolation before the first edit No repository layer
Plan approval, diff review, PR merge Plan gate, per-hunk review, PR in session Out of scope for a gateway
No markup on your existing subscriptions BYOK sessions are direct to the provider BYOK is supported on pay-as-you-go; the 5% markup applies to model usage cost
Cost, latency, and availability routing Auto model router on optional hosted inference Cost, latency, and availability routing with weighted load balancing
Automatic failover between providers A failing provider surfaces its own error with an explicit retry path Fallback policies reroute failed requests to backup models
Prompt caching Provider-native caching inside each agent session Auto-caching, claimed at up to 80% to 90% on repeated prompts
Cost by repository Repo, provider, model, and day, computed locally Spend by key, user, model, or project, not by git repository
Live provider quota gauges Reads each provider's own rate-limit window Real-time cost, latency, TTFT, and error dashboards for routed traffic
Enterprise access controls Model policy, weekly spend caps, member approvals Enterprise adds SSO, full RBAC, audit logs, model whitelist, guardrails, PII detection
EU data residency BYOK sessions stay with the provider you chose; verify hosted residency Listed on both the free and pay-as-you-go plans
04

Where markup and allowance separate

01 · What is being sold

Routed tokens versus a supervised session.

Requesty sells one endpoint over 600+ models with routing, caching, and failover, priced at 5% over base model cost. Continuum sells the place where coding agents run: worktree, branch, plan gate, diff, pull request, quota gauge, and a cost ledger. Only the inference bill overlaps.

02 · Markup versus allowance

5% of a moving number versus a fixed weekly ceiling.

A percentage markup is honest and easy to model, and it scales exactly with usage, which is the point and also the catch: an agent loop that triples its tokens triples the fee. Continuum hosted tiers are $25, $100, $200, and $500 per month with weekly allowances of $25, $100, $200, and $1,000, so the ceiling is known before the week starts. Compare them on your own usage curve, not on a list price.

03 · The savings claims

Up to 80% and up to 90% are the same claim, stated twice.

Requesty's homepage says prompt caching reduces token costs by up to 90%; the documentation says up to 80% on repeated prompts. Both are ceilings on a workload with heavy prompt repetition. Coding agents do repeat context, so the mechanism is real, but the realized number depends on your prompts and should be measured rather than assumed.

04 · What a gateway cannot gate

A router sees tokens; it never sees the repository.

Routing, caching, and failover are decisions about a request. Whether an agent may write to a file before a human approves the plan, whether its edits land on an isolated branch, and which repository consumed the budget are decisions about a workspace. Requesty is not built to answer those, and Continuum does not claim to beat it on routing.

05 · BYOK

Both support it; only one has no fee at all.

Requesty supports bring-your-own-keys on pay-as-you-go with the 5% markup applied to model usage cost. Continuum's BYOK sessions talk to the provider directly with no Continuum fee, because there is no proxy in the path. That difference matters most for teams whose agent spend is mostly subscription seats rather than metered API calls.

05

Coexistence

Stack recipe

How people run both.

The competition is narrow. Requesty and Continuum hosted inference are two answers to the same question about who supplies and meters tokens; the workbench above them is not contested, and Continuum does not ship a Requesty adapter. If Requesty is already your company endpoint, the honest comparison is Requesty against Continuum hosted tiers, not against the free workbench.

06

Pick by what is calling the model

Continuum

Choose Continuum for supervised coding agents

  • The tokens are being spent by coding agents you supervise against real repositories.
  • You want worktree isolation, plan approval, per-hunk diff review, and pull requests in the same session.
  • Your agent spend is mostly on subscriptions you already pay for, which no gateway markup should touch.
  • A fixed weekly hosted allowance is easier to defend than a percentage of an unpredictable token curve.
  • You need cost attributed to a repository and provider, and weekly caps an org admin can set.
  • You want the same live session on Mac, iPhone, and web without rebuilding it on each device.
Requesty

Choose Requesty for application routing

  • An application, not a person, is calling the model.
  • You want 600+ models across 30+ providers behind one OpenAI-compatible base URL.
  • Automatic failover and cost or latency routing are the features you are actually buying.
  • Prompt repetition is high enough that auto-caching should pay for the 5% markup.
  • You need P50 through P99 latency, TTFT, error rate, and cache impact in one dashboard.
  • Enterprise needs SSO, full RBAC, audit logs, model whitelists, guardrails, or PII detection at the gateway.
07

What you pay

Continuum

App + your labs

$0 for the app on Mac, iPhone, Watch, web, and the Windows and Linux desktop beta, running under the provider subscriptions you already pay for. Optional hosted inference is $25 per month for Plus, $100 for Max 100, $200 for Max 200, and $500 for Ultra, with weekly hosted-usage allowances of $25, $100, $200, and $1,000.

Requesty

Their bill

A free tier with free models, 200 requests per day, routing, caching, fallbacks, spend tracking, and EU data residency, plus $10 in credits at signup. Pay-as-you-go is a 5% markup on base model cost with no subscription, seat fee, or minimum spend, illustrated on the pricing page as a $10 per million token model costing $10.50 through Requesty; it includes 600+ models, bring-your-own-keys, routing policies, caching, fallbacks, spend limits, an MCP gateway, EU data residency, and advanced observability. Enterprise is custom priced and adds SSO, full RBAC, audit logs, an approved-model whitelist, team controls, guardrails, PII detection, and custom SLAs.

How to compare

Total cost of work

Requesty's markup is transparent and scales with usage; Continuum's hosted tiers are flat with a weekly ceiling, and its BYOK sessions cost nothing because nothing is proxied. Run one real week on each before choosing, because the crossover point depends entirely on your token curve.

Requesty plans checked 3 August 2026 against their pricing page. Vendors reprice often - if you spot a stale figure, tell us and we will fix it. Continuum never marks up a session it runs under your own login.

Continuum ladder: Free app · Plus $25/mo · Max 100 · Max 200 · Ultra - full pricing. Also analytics, multi-account, devices.

08

Questions

Continuum vs Requesty.

Deep dives: docs, providers, sessions.

Pay-as-you-go is a 5% markup on base model cost with no subscription, seat fee, or minimum spend; the pricing page illustrates it as a $10 per million token model costing $10.50. There is a free tier with free models and 200 requests per day plus $10 in signup credits, and a custom-priced Enterprise plan.

For the inference bill, yes: Continuum hosted inference is a flat monthly tier from $25 to $500 with a weekly allowance, competing with a per-token markup. For routing, caching, and failover across 600+ models, no. Continuum is the workbench where coding agents run, and it does not claim to replace a routing gateway.

That is a ceiling, and the two sources differ: the homepage claims up to 90% from prompt caching while the documentation says up to 80% on repeated prompts. The mechanism is real for workloads with heavy prompt repetition, but the realized figure depends on your cache hit rate and should be measured on your own traffic.

Both are OpenAI-compatible aggregation gateways. Requesty advertises 600+ models across 30+ providers with a flat 5% markup on model cost. OpenRouter advertises 500+ models across 80+ providers with no subscription and fees on credit purchases: 5.5% on Stripe with a $0.80 minimum, 5% on crypto, and 5% on bring-your-own-key usage above a $25,000 monthly allowance. Compare provider coverage for the specific models you need, then compare the fee shape.

Requesty's documentation includes a Claude Code integration guide, so it can supply the model. It does not supply the session layer: worktree isolation, plan approval, diff review, and pull requests come from a workbench like Continuum, which runs Claude Code under your own subscription with no markup.

They measure different axes. Requesty reports spend and latency by key, user, model, or project for traffic it routes. Continuum prices local agent logs by repo, provider, model, and day, including subscription work that never crosses a gateway. If the question is which repository burned the budget, that is Continuum's axis.

Begin

Run your agents
in Continuum.

Free app. Your subscriptions. Optional hosted inference. Mac stable - Windows and Linux desktop are beta.

vendor-neutral · local-first · multi-device