A 5% markup on every token routed, versus a free workbench whose hosted tier is a flat weekly allowance.
Requesty is an AI gateway with an unusually clean pricing story: change your base URL, reach 600+ models across 30+ providers through an OpenAI-compatible API, and pay a 5% markup on base model cost with every feature included. Its site claims 70,000+ developers and 90+ billion tokens per day. The free tier gives free models, 200 requests per day, and routing, caching, fallbacks, spend tracking, and EU data residency at no cost, with $10 in credits on signup. Pay-as-you-go adds bring-your-own-keys, routing policies, spend limits, an MCP gateway, and advanced observability with P50, P90, P95, and P99 latency, TTFT, error rates, and cache impact. Enterprise adds SSO, full RBAC, audit logs, model whitelists, guardrails, and PII detection. The savings claims deserve a careful read: the homepage says prompt caching reduces token costs by up to 90%, while the documentation says up to 80% on repeated prompts, so treat both as ceilings on a favorable workload rather than an expected result. Continuum overlaps only on the inference bill. It is the free workbench where Claude Code, Codex, Cursor, Grok, and OpenCode run under your own subscriptions, each in its own git worktree and branch, with plan gates, per-hunk diff review, and pull requests in session, and optional flat hosted inference from $25 to $500 per month.
Updated 2026-08-03 · Mac stable · Win/Linux desktop beta
Pick Continuum when the workload is coding agents you supervise: worktree isolation, plan gates, diff and PR review, live quota gauges, cost by repo, BYOK-first with no markup, and flat hosted tiers if you want them.
Pick Requesty when an application needs 600+ models behind one endpoint with cost, latency, and availability routing, automatic failover, prompt caching, an MCP gateway, and observability, at a transparent 5% markup.
Two worktree sessions start on the same repo, one Claude Code and one Codex, each on its own branch under an existing subscription.
Plan mode holds both read-only. The better plan is approved from the phone and the weaker session is interrupted.
The diff is reviewed hunk by hunk, a bad turn is rolled back to a checkpoint, and the pull request opens from the session.
Cost by repo and the live quota gauges decide whether tomorrow's batch runs now or after the window resets.
Swap the base URL to Requesty and reach 600+ models through the same OpenAI-compatible request shape.
Enable auto-caching and a fallback policy, then measure the real hit rate against the up-to-80% and up-to-90% claims.
Add cost or latency routing and read P95 and TTFT in the dashboard rather than the average response time.
Reconcile the invoice at 5% over base model cost, and scope Enterprise if SSO, RBAC, audit logs, or PII detection are required.
Install, connect, first session - steps you can run the same day.
Decide what is calling the model. An application is a Requesty case. A developer supervising coding agents is a Continuum case.
If it is an application, start on Requesty's free tier: free models, 200 requests per day, plus the $10 signup credit, and confirm the OpenAI-compatible swap is really one base-URL change.
Turn on caching and a fallback policy, then measure the cache hit rate on your own prompts. The up-to-90% and up-to-80% claims are ceilings, not forecasts.
Compare cost, latency, and availability routing against a fixed model, and read the P95 and TTFT numbers rather than the average.
If it is coding agents, install Continuum and start two worktree sessions on a bounded ticket under the subscriptions already on the machine. That path has no markup because nothing is proxied.
Approve a plan, review the diff hunk by hunk, and open the pull request from the same session.
Only now compare bills: a week of Requesty at 5% over base model cost against a Continuum hosted tier with a fixed weekly allowance of $25, $100, $200, or $1,000.
Check cost by repo, provider, model, and day before committing, and set a weekly spend cap if the answer surprises you.
Cells use product nouns on both sides. Continuum is the multi-provider workbench; they win where their loop is the product.
Requesty sells one endpoint over 600+ models with routing, caching, and failover, priced at 5% over base model cost. Continuum sells the place where coding agents run: worktree, branch, plan gate, diff, pull request, quota gauge, and a cost ledger. Only the inference bill overlaps.
A percentage markup is honest and easy to model, and it scales exactly with usage, which is the point and also the catch: an agent loop that triples its tokens triples the fee. Continuum hosted tiers are $25, $100, $200, and $500 per month with weekly allowances of $25, $100, $200, and $1,000, so the ceiling is known before the week starts. Compare them on your own usage curve, not on a list price.
Requesty's homepage says prompt caching reduces token costs by up to 90%; the documentation says up to 80% on repeated prompts. Both are ceilings on a workload with heavy prompt repetition. Coding agents do repeat context, so the mechanism is real, but the realized number depends on your prompts and should be measured rather than assumed.
Routing, caching, and failover are decisions about a request. Whether an agent may write to a file before a human approves the plan, whether its edits land on an isolated branch, and which repository consumed the budget are decisions about a workspace. Requesty is not built to answer those, and Continuum does not claim to beat it on routing.
Requesty supports bring-your-own-keys on pay-as-you-go with the 5% markup applied to model usage cost. Continuum's BYOK sessions talk to the provider directly with no Continuum fee, because there is no proxy in the path. That difference matters most for teams whose agent spend is mostly subscription seats rather than metered API calls.
The competition is narrow. Requesty and Continuum hosted inference are two answers to the same question about who supplies and meters tokens; the workbench above them is not contested, and Continuum does not ship a Requesty adapter. If Requesty is already your company endpoint, the honest comparison is Requesty against Continuum hosted tiers, not against the free workbench.
$0 for the app on Mac, iPhone, Watch, web, and the Windows and Linux desktop beta, running under the provider subscriptions you already pay for. Optional hosted inference is $25 per month for Plus, $100 for Max 100, $200 for Max 200, and $500 for Ultra, with weekly hosted-usage allowances of $25, $100, $200, and $1,000.
A free tier with free models, 200 requests per day, routing, caching, fallbacks, spend tracking, and EU data residency, plus $10 in credits at signup. Pay-as-you-go is a 5% markup on base model cost with no subscription, seat fee, or minimum spend, illustrated on the pricing page as a $10 per million token model costing $10.50 through Requesty; it includes 600+ models, bring-your-own-keys, routing policies, caching, fallbacks, spend limits, an MCP gateway, EU data residency, and advanced observability. Enterprise is custom priced and adds SSO, full RBAC, audit logs, an approved-model whitelist, team controls, guardrails, PII detection, and custom SLAs.
Requesty's markup is transparent and scales with usage; Continuum's hosted tiers are flat with a weekly ceiling, and its BYOK sessions cost nothing because nothing is proxied. Run one real week on each before choosing, because the crossover point depends entirely on your token curve.
Requesty plans checked 3 August 2026 against their pricing page. Vendors reprice often - if you spot a stale figure, tell us and we will fix it. Continuum never marks up a session it runs under your own login.
Continuum ladder: Free app · Plus $25/mo · Max 100 · Max 200 · Ultra - full pricing. Also analytics, multi-account, devices.
Pay-as-you-go is a 5% markup on base model cost with no subscription, seat fee, or minimum spend; the pricing page illustrates it as a $10 per million token model costing $10.50. There is a free tier with free models and 200 requests per day plus $10 in signup credits, and a custom-priced Enterprise plan.
For the inference bill, yes: Continuum hosted inference is a flat monthly tier from $25 to $500 with a weekly allowance, competing with a per-token markup. For routing, caching, and failover across 600+ models, no. Continuum is the workbench where coding agents run, and it does not claim to replace a routing gateway.
That is a ceiling, and the two sources differ: the homepage claims up to 90% from prompt caching while the documentation says up to 80% on repeated prompts. The mechanism is real for workloads with heavy prompt repetition, but the realized figure depends on your cache hit rate and should be measured on your own traffic.
Both are OpenAI-compatible aggregation gateways. Requesty advertises 600+ models across 30+ providers with a flat 5% markup on model cost. OpenRouter advertises 500+ models across 80+ providers with no subscription and fees on credit purchases: 5.5% on Stripe with a $0.80 minimum, 5% on crypto, and 5% on bring-your-own-key usage above a $25,000 monthly allowance. Compare provider coverage for the specific models you need, then compare the fee shape.
Requesty's documentation includes a Claude Code integration guide, so it can supply the model. It does not supply the session layer: worktree isolation, plan approval, diff review, and pull requests come from a workbench like Continuum, which runs Claude Code under your own subscription with no markup.
They measure different axes. Requesty reports spend and latency by key, user, model, or project for traffic it routes. Continuum prices local agent logs by repo, provider, model, and day, including subscription work that never crosses a gateway. If the question is which repository burned the budget, that is Continuum's axis.
Free app. Your subscriptions. Optional hosted inference. Mac stable - Windows and Linux desktop are beta.
vendor-neutral · local-first · multi-device