Grok Bot review (August 2026): is the $40 to $300 agent worth it?

Grok Bot launched on 11 August 2026 and this review was written on 22 August, which is eleven days of public evidence rather than a year of production use. That framing matters, because the failure modes people report are launch-week failure modes and the objections people raise are not. This page separates the two: what may be fixed, and what is a design decision you either accept or do not.

By the Continuum team. We build a workbench that runs Claude Code, Codex, and their peers, so the model rates quoted here are the ones our own cost analytics ship with.

The short version

Grok Bot is the smoothest onboarding in the always-on agent category and one of the few products where desktop and mobile are genuinely the same thing. Teaching a routine by demonstration works and is the most compelling part of the product. Against that: launch-week reports describe Bots looping in conversation with each other, a usage meter that showed 48% consumption as 0%, and GitHub authentication failing on iOS. Underneath the bugs sit three decisions that will not be patched away: one shared computer with no boundary between Bots, an undisclosed model router with no picker, and no queryable organisation-wide audit view beyond per-Bot transcripts and the last twenty routine records reported by Daily Dose of Data Science and unconfirmed by xAI. Worth trying if you already hold an eligible plan and have no compliance obligation. Not yet worth buying a plan for, and not usable at all if your organisation runs Cursor Legacy Privacy Mode.

What you need to know
  • Setup is genuinely excellent. One developer who had been hand-rolling the same thing called it "about as far as I imagine is possible" and "a pretty slick experience".
  • Teach-by-demonstration is the standout feature, and a better interaction model than writing a prompt that describes a procedure.
  • Launch-week bugs are real: Bots looping at each other, a usage display showing 48% as 0%, GitHub auth failing on iOS.
  • Three problems are structural, not bugs: one shared computer, an undisclosed model router with no picker, and no real audit log.
  • The trust objection is the loudest thing in the launch thread, and it is about the vendor rather than the software.
  • Verdict: try it if you already hold an eligible plan. Do not buy a plan for it, and do not point it at anything regulated.

The verdict, up front

Grok Bot is the most polished first-run experience in a category where first-run experiences are usually terrible, and it is the first mainstream product to make teaching an agent by demonstration feel like an ordinary thing to do. Those are real achievements and the reason the launch got the attention it did.

It is also eleven days old, in beta, with a usage meter that did not work at launch, no queryable organisation-wide audit view, no dry-run mode, no model picker, no published allowance, no spend cap, and one shared machine holding every credential you give it. Some of that may be fixed quickly. The shared machine and the model router are architecture, and the compliance gaps are a company decision rather than a backlog item.

Two disclosures about this review. It is written from the public record: xAI’s documentation, the launch announcements, the Hacker News launch thread, and the published teardowns, all read on 22 August 2026. And we build a competing product, which is stated here rather than buried, because the honest concessions below are worth more if you know what we sell.

What genuinely works

Four things stand out, and the first one is the reason this product got traction rather than a shrug.

01

The setup is the best in the category

Install, sign in with Cursor, pick a suggested teammate or create your own, name it and describe its job. That is the whole thing. The most credible praise in the launch thread came from a developer who had been assembling the same capability by hand: Grok Bot "simplifies the setup for this about as far as I imagine is possible, and frankly it’s a pretty slick experience." Coming from someone who had already done it the hard way, that is worth more than any benchmark.

02

Teaching a routine by demonstration

Instead of writing a prompt that describes a procedure and hoping the description is complete, you perform the procedure once while the Bot watches, then ask it to save the path as a routine. The routine can be re-run on demand or on a schedule. This is a better interaction model than prompt engineering for anything repetitive, and it is the part of the product that will still look good in a year.

03

Desktop and mobile are the same product

Not a full app and a viewer. You can watch a live session from the phone and take over the machine from either surface. In a category where the phone client is usually a read-only afterthought, this is a real differentiator and it is well built.

04

Group chats with real handoffs

Several Bots in one thread, reported working with two to six, sharing context and passing ownership of a task. Because they share the computer, a handoff carries the files and the browser sessions with it, so the second Bot does not start by logging in again. It is the shared-machine design paying off rather than costing.

Underneath all four sits the persistence, and it is the thing chat-based assistants genuinely cannot do. A named Bot that remembers your invoice format, still has last week’s spreadsheet on its filesystem, and is still logged into the portal it needs is qualitatively different from a chat window that starts empty every time.

What does not work yet

These are the reported failures from the first eleven days. They are bugs rather than design, which means they are the part of this review most likely to be out of date soonest. Check the date at the top of this page before treating any of them as current.

Reported problems in the launch build, with the source for each. Compiled 22 August 2026.
ProblemWhat was reportedSource
Bots loop at each otherAgents "go in loops talking to each other before finally shutting up" in group threadsA practitioner first-impressions thread; second-hand, we could not load the original post
Metering display broken48% of the allowance consumed while the interface showed 0%eesel teardown, 12 Aug
GitHub auth fails on iOSAuthentication to GitHub failing specifically on the iPhone appeesel teardown, 12 Aug
Router quality at launchA tester found the automatic model routing "wasn’t great" early onVentureBeat
Docs against download pageDocumentation says no Linux app; a user reported the download page offered oneHacker News launch thread

The metering bug deserves separate weight, because it is not merely cosmetic. This is a product with an unpublished allowance, no spend cap, and overage at raw model rates. The usage display is the only instrument on the dashboard, and at launch it read zero while the tank was half empty. Anyone who ran Bots hard in week one had no way to know what they were spending, which is precisely the situation the beta tester quoted in that teardown described: "I’ve used more tokens this month than not this month. That’s not a typo."

The looping problem is the more interesting one for the product’s future. Multi-agent group chat is the headline feature, and agents talking past each other in circles is the classic failure mode of exactly that design. It is fixable, and how well it is fixed will decide whether the group chat is the product or a demo.

The three problems that are not bugs

Separate these from the list above, because no release note is going to close them without changing the product.

01

One computer, no boundary

The documentation states that all of your Bots share one persistent cloud machine, that they share files, browser sessions, and app logins, and that each Bot having its own screen does not give it a separate security boundary. It then tells you directly to "Treat a login or file placed on the computer as available to all of your Bots." That rules out the pattern most people reach for on day two, which is a throwaway Bot for risky work and a trusted Bot for the accounts that matter.

02

A router you cannot see or steer

xAI names no default model for Grok Bot anywhere. There is no picker. Grok 4.6 is available in the product and is the plausible engine, but you cannot verify what ran your task, cannot send cheap work to a cheaper model, and cannot predict a cost. On a metered product with no cap, opacity in the router is a billing problem as much as a transparency one.

03

No audit trail worth the name

The documentation describes an audit view as coming, twice. What exists today is per-Bot chat transcripts, not queryable across an organisation, plus the twenty most recent routine records, reported by Daily Dose of Data Science and unconfirmed by xAI. For an agent that logs into your accounts and acts as you, twenty records is not an audit trail, it is a recent-items list.

What the launch thread actually objected to

The Hacker News launch thread reached 350 points, and the striking thing about it is how little of the criticism was about the software. A second submission the same day was flagged off the front page, and its comments are entirely about the flagging. That is useful context on how this audience received the product, separately from what the product does.

The objections cluster into four, and they are worth quoting rather than paraphrasing because the wording carries the temperature.

  • Vendor trust. "Musk’s personal brand is so poisonous that I would never let him anywhere near my data." (basisword) And, more damning because it concedes the field: "Somehow the American AI industry managed to create a product I trust less than existing commercial offerings." (jknoepfler)
  • Geography. "I’m not sure how many companies are comfortable with giving SpaceXAI access to all your files and data. Outside of America this is, most likely, not going to fly." (LaurensBER) With no data residency option published, this one has no counter-argument available.
  • Prompt injection. "Are you all comfortable with the idea of agents running non stop with access to all your accounts?… get hijacked via prompt injection." (dgellow) A long subthread rejected the suggestion that injection is largely solved: "having a system that fails every 40 attempts is outright catastrophic." (stymaar)
  • Accountability. "Grok bot snags your credentials and then pretends to be you whilst surfing the web. If it uses your account to post on X, is it adhering to X terms of service?" (SilverBirch) The question has no answer in the documentation.

There was product criticism too, and it was sharper than the trust complaints because it went at the shape of the thing: "It seems like Grok Bot is just a personal agent swarm… surprising that’s all it offers because it does so in a group chat app." (thenbrent) That is the fair version of the "is this a product or a demo" question, and eleven days is too early to answer it.

How it compares with the rest of the category

Grok Bot did not invent always-on agents, it made them easy to start. Judged against the products it will actually be compared with, it wins on onboarding and loses on control, and the pattern is consistent enough to be worth stating as a rule.

How Grok Bot lands against the products it is most often compared with. Positioning from each product’s own materials, August 2026.
Compared withGrok Bot wins onGrok Bot loses on
OpenClawNothing to host, nothing to maintain, no engineering timeSource you can read, a machine you control, no vendor meter
Hermes AgentSetup speed and mobile paritySandboxing between agents
ManusPersistent machine and taught routinesA visible credit meter you can reason about
ChatGPT WorkA genuinely persistent computer rather than a sessionVendor trust, compliance posture, and enterprise readiness
DevinBreadth beyond code, and a far lower entry priceGovernance, audit, and repository-shaped work
A workbench on your own hardwareZero setup, nothing to keep awakeModel choice, cost visibility, isolation, and auditability

The row that matters most for anyone reading this after a comparison listicle is the OpenClaw one, because OpenClaw occupies the "your own hardware" slot in effectively every alternatives list and occupies it badly. It is an MIT framework rather than a product, with a real bill underneath it: a VPS to host it and model API spend to feed it, plus the maintenance. The launch thread carried the honest version of that from someone who tried: OpenClaw "broke every update for a month, so I gave up and moved to Hermes." So the choice presented in most listicles, between a slick rented machine and a framework you maintain yourself, is a false one. A finished workbench that runs on the machine already on your desk is a third option that those lists do not carry.

Verdict by who you are

The recommendation, by situation. As of 22 August 2026.
If you areVerdict
On Cursor Ultra or SuperGrok Heavy alreadyTry it today. It costs you nothing extra and the setup takes minutes.
On Cursor Pro at $20Not worth the upgrade to Pro+ yet. Wait for the audit view and a working meter.
Blocked by software with no APIThe strongest case for this product. Computer use on a persistent machine is the hard thing it does well.
A developer wanting repository workWrong product. Grok Build, or a workbench that runs several coding agents.
A team with any compliance obligationNo. No published SOC 2, ISO 27001, GDPR, or HIPAA claim found; no published residency option; no queryable organisation-wide audit view.
Running Cursor Legacy Privacy ModeNot available at any price. That mode blocks Grok Bot entirely.
Outside the United StatesRead the residency question first. There is no published answer.
Cost sensitiveWait. Unpublished allowance, no cap, no model picker, and the meter was broken at launch.

The row worth arguing with is the first one, because "try it today" is not the same as "adopt it". A trial account that holds one throwaway login tells you what you need to know: whether computer use survives contact with your specific software, how often an auth wall pulls you back in, and how much of an unpublished allowance a real week eats. None of those are answerable from a feature list, and all three are answerable in an afternoon.

What we could not test, and what would change this review

Stating this plainly is the difference between a review and a summary. This assessment is built on the public record as of 22 August 2026: xAI’s documentation and announcements, Cursor’s pricing page, the launch thread, and the published teardowns. It is not a long-term production trial, because the product is eleven days old and nobody has one.

Specifically unverified: the exact free-trial duration, which xAI does not publish and third parties disagree about; SuperGrok Heavy and Plus prices, which every article reports and xAI publishes nowhere we could reach; the looping report, which came from a practitioner thread we could not load directly; and the current state of the Linux download, which the documentation and a user report describe differently.

  • What would move this review up: the audit view shipping, a published weekly allowance, a spend cap, a model picker, and a dry-run mode. Any three of those would make it recommendable to a team.
  • What would move it down: a prompt-injection incident against the shared computer, or the group-chat looping turning out to be hard to fix rather than a launch bug.
  • What will not change: the shared-machine design and the Cursor dependency. Those are decisions, and they are the ones a buyer has to actually accept.

Questions people ask

Is Grok Bot worth it?

If you already hold Cursor Pro+, Cursor Ultra, Cursor Teams Standard, or Cursor Teams Premium, or SuperGrok Plus or Heavy, yes, it is worth an afternoon: the setup is the best in the category and teaching a routine by demonstration genuinely works. If you would need to buy a plan for it, wait. The usage meter was broken at launch, there is no queryable organisation-wide audit view, no dry-run mode, no model picker, and no spend cap.

What are the biggest problems with Grok Bot?

Three that may be patched: Bots looping in group chats, a usage display that showed 48% consumption as 0%, and GitHub authentication failing on iOS. Three that will not: all Bots share one computer with no security boundary between them, the model router is undisclosed with no picker, and there is no queryable organisation-wide audit view beyond per-Bot transcripts and the last twenty routine records reported by Daily Dose of Data Science and unconfirmed by xAI.

Is Grok Bot safe?

No published SOC 2, ISO 27001, GDPR, or HIPAA claim was found in its documentation; there is no published retention period or data residency option, no queryable organisation-wide audit view, and no dry-run mode, and every Bot on your account can reach every credential on the shared machine. For personal use with logins you can afford to lose, that may be acceptable. For anything regulated it is not, today.

Do Grok Bots really talk to each other?

Yes, that is a designed feature: Bots share context in threads and group chats, reported working with two to six in one thread, and can pass ownership of a task. Launch-week reports also describe them going in loops talking to each other before stopping, which is the classic failure mode of multi-agent chat and the thing most worth watching in the next few releases.

Does Grok Bot work on a phone?

Yes, and this is one of its better features. The iPhone app is not a viewer: desktop and mobile are the same product, you can watch a live session from the phone and take over the shared machine from either surface. iOS 18 or later. There is no iPad app and Android is listed as coming.

Which model does Grok Bot use, and can I change it?

You cannot change it, and xAI does not say what it is. Grok 4.6 is live in Grok Bot and is the plausible default, but no default is documented and there is no picker. Since overage bills at raw model rates, that means you cannot predict what a task will cost before running it.

Is Grok Bot better than running agents on my own machine?

It is better at two specific things: there is nothing to set up, and there is no machine of yours to keep awake. It is worse at cost visibility, model choice, isolation between agents, auditability, and vendor concentration. If your blocker is not wanting to babysit hardware, it wins. If your blocker is wanting to reach an agent from your phone, that is a reachability problem with cheaper answers.

Should my company use Grok Bot?

Not yet, unless it has no compliance obligations at all. There is no audit trail, no retention policy, no data residency option, and no certification claim, compliance defers to Cursor’s terms rather than an xAI agreement, and enterprise access is a waitlist rather than a plan. If your organisation runs Cursor Legacy Privacy Mode, Grok Bot is blocked outright.

Sources

Every figure above was read from these pages on August 2026. Vendors reprice without notice; if you find a stale number, tell us.

  1. Grok Bot overview, docs.x.ai the shared cloud computer, per-Bot screens, memory, routines, connectors, credential handover
  2. Grok Bot get started, docs.x.ai eligible plans, Cursor sign-in requirement, platform downloads, first-run steps
  3. Hacker News: the Grok Bot launch thread launch-week reception, trust and prompt-injection objections, the Linux download report
  4. eesel AI: Grok Bot review metering behaviour, compliance gaps, no dry run, audit view described as coming
  5. VentureBeat on Grok Bot SpaceXAI branding, tier prices, criticism of automatic model routing
  6. xAI: Introducing Grok Bot launch date, beta status, positioning, enterprise waitlist
  7. xAI: Grok Bot on more plans 21 August 2026 access expansion to SuperGrok Plus, Cursor Pro+, and Cursor Teams
Try it

Same agents.
Your hardware.

Continuum runs Claude Code, Codex, Cursor, Grok, Gemini, and OpenCode in isolated worktrees on machines you own, with live quota gauges, spend by repository, and a model you actually choose.

free app · your subscriptions · local-first