/push docs

The manual for the terminal that remembers your agent sessions.

Browse the docs
iPhone

Voice capture

Hold, talk, done — a structured issue in the right folder. And the complete story of what happens to your audio.

Say the thing; Push files it. A recording becomes a structured issue — titled, written up, and landed in the right folder on its own.

Recording

One capture surface, two button styles (Settings → Voice Capture):

  • Push to Talk — "Hold the button to record, release to finish." Shows a live (blurred) transcript while you speak.
  • Tap to Talk — "Tap the button to start a long hands-free recording," for thinking out loud or capturing a conversation.

Capture is also an App Shortcut named Push to Talk, which means you can put it on the Action Button and record without opening the app.

While recording you get a live waveform, gentle haptics synced to your voice (optional), and an on-device "Speaking.. / (silence)" indicator. Recordings run up to 60 minutes, with a one-minute warning before the cap.

What comes out

  • A normal capture (up to about five minutes of speech) is transcribed in full and authored into a todo — title, body, and an automatic folder assignment ("Lands in the right folder on its own").
  • A long recording becomes a Meeting item — titled with its date, its first stretch transcribed, and marked "Still recording. The session will read the rest." The full audio is always kept and syncs to your Mac, where an agent session can work through the whole thing. Show full transcript is right on the item.

Recording on an existing item edits it — voice edits merge into the todo rather than creating a duplicate (the bar switches to "push to edit"). Not happy with an extraction? Regenerate re-runs it. And the + menu adds typed text, links, images, and files — no microphone required.

Offline

With no connection, capture still works: on a device with Apple Intelligence (and offline mode enabled), transcription and structuring run entirely on-device using a bundled speech model and Apple's on-device foundation models. No connection and no Apple Intelligence? The recording is kept and processed when you're back online.

What happens to your audio

The routing rule is simple: online → cloud, offline → on-device. Typed text, links, and images never use the cloud pipeline at all.

When your recording is processed in the cloud, two steps run:

  1. Transcription — Whisper Large v3-class speech recognition via our processing partners. If the live transcript already captured your speech as you spoke (a streaming transcriber handles that), only the text is sent — no audio bytes at all.
  2. Structuring — a Gemini-class model via OpenRouter turns the transcript into the todo. The model sees text (and an attached screenshot, if you added one) — never your raw audio.

The request can include: the transcript (or the audio when there's no transcript yet), your folders' names and descriptions (so filing happens in one step), custom vocabulary, the item being edited (for voice edits), an attached screenshot, the name of the app you were in when you started recording (context for "add this to the doc I'm reading"), and language hints. It never includes your location, contacts, or anything from your Mac.

Retention: the audio copy used for processing is deleted within about five minutes of successful extraction, with a 24-hour sweep for any stragglers. Your recording itself lives on your iPhone and your Mac — under your control, clearable anytime — as described in How Push handles your data.

The recording allowance

Push includes a monthly recording-minutes allowance at no cost — the same for everyone; see pricing for the current numbers. Settings → Transcription & Data shows this month's usage. Hit the cap and typed capture keeps working; minutes reset at the start of each month. Voice input into a live terminal session doesn't count against it.

Verified against Push 1.10.197 · iOS 0.11.0 · Updated 2026-08-08