Hold Option (or Fn) and talk - a global overlay captures speech into the Code composer on Mac, or into whatever field is focused in any other app. On iPhone, the Continuum Voice keyboard works in any app’s text field. Local models when you want offline; hands-free when your hands are full.
Lisbon · fable-5 · xhighworktree · lb-1Hold Option. Speak. Release. Text lands in the prompt.
A global Option-hold (configurable) brings up a Continuum voice overlay - waveform, partial transcript, silence countdown - without leaving the app you’re in. Release to stop, or let silence auto-stop finish the take.
There’s also an in-composer mic when you’re already in Code. Same engine, same models, same silence settings. Web and desktop composers reach the same stack through the voice bridge when a host is paired.
Partial text streams while you speak, so you can see the recognizer keeping up and stop early when it has enough.
Pick the key and the gesture in Settings → Voice. The global trigger can be left Option, right Option, or Fn / Globe, and it can fire on hold, on double-tap, or on either. Hold is push-to-talk: speak while the key is down, release to send. Double-tap is hands-free: one tap opens the take, a second closes it, and going quiet closes it for you.
The global trigger runs on a keyboard event tap, so macOS asks for Accessibility the first time you arm it. Continuum watches that trust state live and re-arms the moment you flip the switch in System Settings, with no relaunch. Leave the trigger unset and no tap gets installed at all - you still get the in-composer mic and Control-M whenever Continuum is in front.
Where the words land depends on what is frontmost. Continuum in front streams into the active composer. Another app in front, with system-wide dictation on, inserts into whatever field has focus: a commit message, a ticket, a Slack thread. System-wide off is the kill switch - the gesture then fills Continuum’s own fields only and never types into someone else’s window. A plain terminal Claude Code session has no composer to dictate into at all, which is the gap this closes.
Pick a local engine: Apple Speech, WhisperKit, or Parakeet. Download the model once; run offline after that. No cloud STT hop for the take you don’t want to leave the laptop.
Parakeet is the fast path for long dictations. WhisperKit covers the size-versus-accuracy curve. Apple Speech is zero-download on supported system locales. Switch in Settings → Voice - the overlay, the composer mic, and the iPhone keyboard all honor the choice.
Every engine sits behind one transcription protocol, so changing models never changes a hotkey, a permission, or where the text lands.
The catalog is short on purpose. Parakeet V3 is the recommended local engine at roughly 478 MB: it runs on the Apple Neural Engine, auto-detects across 25 European languages, and is the fastest thing on the list. Parakeet V2 is English-only with the highest recall. Apple Speech is built in and downloads nothing. WhisperKit fills in the rest of the curve, from tiny near 39 MB for rough notes up through medium near 769 MB. Anything past 500 MB warns you before it starts pulling, and every card shows its size, language reach, and a coarse speed-versus-quality read.
A model downloads once and every dictating surface reuses it. After that a take never touches the network: no upload, no queue, no vendor in the path. That matters more the closer your prompts sit to production, which is the same instinct behind keeping an agent’s blast radius local. If auto-detect guesses your language wrong, pin the recognition locale in the same pane.
Local is the default and the honest one. When you want a bigger model than your laptop should run, Continuum routes the take through a provider you already pay for, on the host you already trust.
Grok Speech to Text through a connected xAI account. GPT-4o Transcribe and GPT-4o mini Transcribe through a connected OpenAI account. Each row appears only when the host confirms the credential is really there.
The browser or phone records and ships bounded audio chunks over the encrypted host link. The provider call is made by the host. Credentials are never copied onto a controller device.
The Windows and Linux desktop app keeps a direct native bridge. The web app adopts a paired host only when that host advertises the upload route, so a mic button never appears where it can’t work.
Availability is host-authoritative, not a hardcoded list: the client asks the host which speech models it can actually serve and renders exactly those. That keeps the picker honest across a fleet where one Mac has an xAI key and a Linux box does not. It is the same rule Continuum follows for every provider credential, and a large part of why a client beats running the CLI on its own once you have more than one machine.
The Continuum Voice keyboard is not locked to Continuum. Enable it once, switch to it from any app - Messages, Mail, Notes, Safari - and dictate. Text inserts at the cursor; if the host field won’t accept it, Continuum falls back to the clipboard.
While you record, Dynamic Island and a Live Activity show the take so you can see status without opening the app. Watch steers sessions; it doesn’t host speech recognition.
It is deliberately one key. Nothing to relearn, no autocorrect to fight, and one globe tap back to your usual layout.
The mechanics explain the one requirement. An iOS keyboard extension cannot open the microphone. So the keyboard asks the Continuum app to record, the app publishes recognition state through a shared App Group, and the keyboard reads that state and types the result into the field you were already in. That bridge is what Full Access buys, which is why the dictation key only appears once you grant it. Without Full Access the keyboard still loads, as a plain local QWERTY, and says so on its face instead of failing silently.
A finished transcript belongs to the field first. The keyboard gets a short window of first refusal to insert it; if focus moved, iOS recycled the extension, or Full Access was revoked mid-take, the text goes to the clipboard instead. It never lands nowhere - a dictation that visibly did nothing is the one outcome the delivery path exists to rule out. Dictating into Continuum’s own composer runs the same path, which is how you open a run from a phone while you are keeping several agents moving at once. An editor-bound tool like Cursor has nothing to offer the phone in your pocket.
When your hands are on a build or a train pole, Continuum still takes the take. Every trigger below runs the same intent, toggles the same session, and delivers to the same place.
Map Toggle Continuum Voice to the Action Button, a Control Center tile, or a Lock Screen control. One press starts a take, the next one finishes it. No keyboard, no chord, no unlock.
The same action ships as an App Shortcut, so a Back Tap gesture, a Siri phrase, or any Shortcuts automation can fire it. One implementation behind four doors, not four half-working copies.
A button press cannot report a release, so silence ends the take. Pick 1.0s, 1.25s, or 2.5s of quiet - the pause that reads as “finished” for you is mid-thought for someone else.
Back Tap and Siri routinely fire at a process that isn’t running. Continuum waits briefly for its recorder instead of throwing “open the app once” at a button you just configured.
Dynamic Island and Lock Screen get phase, elapsed time, and a coarse level bar: listening, finalizing, copied, failed. Transcript text never enters the Live Activity, so nothing you said is readable from a locked phone.
A hands-free session caps at five minutes and gives up early if it hears no speech at all. A pocket press cannot leave a mic open for an hour, and a stalled recognizer resolves instead of hanging.
Recording runs in the background under the system’s audio-recording indicator, so the orange dot is on whenever Continuum is listening and off the instant it stops. This is what makes voice useful for real agentic work instead of a demo: the interesting prompts arrive while you are walking to a meeting, and a thought you can only capture at a desk is usually a thought you lose.
Short answers here, long answers in the docs.
Not on a local engine. Apple Speech, WhisperKit, and Parakeet all run on-device once the model is downloaded, and that is the default. A provider transcription model is the explicit opt-in: you pick it by name, the audio goes to your own paired host, and the host makes the call with its own credential.
Yes. Settings → Voice picks both the key and the gesture: left Option, right Option, or Fn / Globe, firing on hold, on double-tap, or either. Option-hold stays clear of normal typing and of common app shortcuts. Leaving the trigger unset installs no global tap at all.
A global hold or double-tap has to see key events while another app is frontmost, and macOS gates that behind Accessibility. It is also what lets Continuum insert into another app’s focused field. Skip the grant and you keep the in-composer mic and Control-M; only the system-wide half needs it.
Because a keyboard extension can’t use the microphone. The app has to do the recording and hand the result across a shared App Group, and that container is what Full Access unlocks. Without it the keyboard falls back to a plain local layout and says why, instead of showing a mic key that can never work.
Yes. Continuum Voice is a system keyboard - any text field in any app. Enable it under iOS Settings → General → Keyboard, then switch to it from the globe key when you need a take.
No. Watch is for session steering - approve, interrupt, glance. Speech recognition is Mac and iPhone only. That keeps the stack honest about where the mic and the models actually live rather than shipping a wrist feature that quietly leans on the phone.
Download Continuum, grant the mic, hold Option. Pair the spoken instruction with App Shots when the agent also needs to see the frontmost window.
local models · iOS keyboard · hands-free starts