Direct Answer
ChatGPT Voice is no longer limited to spoken questions and spoken answers. In the ChatGPT desktop app, Voice can sit on top of Work or Codex, use the tools and permissions available there, start tasks, receive natural corrections, and report progress while work continues.
For a beginner, the right first setup is deliberately small: one project, one low-risk folder, one reversible task, and approval before any external action. Start with research, summaries, drafts, or a local prototype. Keep sending, publishing, deleting, purchasing, deploying, and permission changes behind a separate human decision.
Watch the Guide
Credit: This guide is based on the hands-on demonstration by AI Edge. The creator shows research, local file creation, website iteration, on-screen context, and paired phone control. Product limits and availability were checked against OpenAI's current documentation on 30 July 2026.
What Actually Changed
The video compares the experience to "Jarvis" because the conversation floats over other work and can keep giving spoken updates. The more precise explanation is a three-part system:
| Layer | What it does | Beginner boundary |
|---|---|---|
| Voice / GPT-Live | Handles the live conversation, interruptions, spoken responses, and conversational steering. | It can mishear speech. Restate important names, dates, destinations, and amounts. |
| Work | Handles longer research and business tasks such as reports, documents, presentations, spreadsheets, and Sites. | Use a dedicated project and only the files or connected tools the task needs. |
| Codex | Works with repositories, local folders, terminals, tests, browser flows, and technical tasks. | Begin in a local test project or branch. Do not start with production credentials. |
OpenAI says Voice in Work and Codex is available on eligible macOS and Windows desktop accounts. It can start, prioritize, interrupt, or redirect tasks, coordinate agents, use supported project context and connected tools, and provide spoken or on-screen progress updates. Availability still depends on plan, region, workspace policy, and app version.
Voice, Dictation, Chat, Work, or Codex?
| Use | Choose | Example |
|---|---|---|
| You want one spoken prompt converted to editable text. | Dictation | Record a long brief, fix the transcript, then submit it. |
| You want a natural conversation or a quick answer. | Voice in Chat | Brainstorm a title or talk through a decision. |
| You want a finished knowledge-work deliverable. | Voice in Work | Research a market and create a reviewable report. |
| You want files, code, tests, or technical browser work. | Voice in Codex | Build a local page, inspect it, revise it, and run checks. |
AI Edge's demonstration moves between research and building. Beginners should keep those jobs separate at first. A research project needs source rules. A coding project needs file boundaries, tests, and rollback. One giant chat makes both harder to audit.
The Beginner Setup
- Update the desktop app. Install the current ChatGPT app for macOS or Windows and confirm Voice is available for the account.
- Choose Work or Codex intentionally. Work is the better default for reports and documents. Codex is for software and local technical work.
- Create one named project. Keep its instructions, reference files, decisions, and outputs together. Pin the conversation if it will become a recurring workspace.
- Open one low-risk folder. Use test files or a copied project. Do not expose the whole home directory, password manager, downloads folder, or unrelated client work.
- Grant permissions only when needed. Microphone is essential. Screen and audio recording, accessibility, local files, browser context, and connectors should be enabled only for a defined task.
- State the stop line. Tell Voice what it may read and draft, and what requires approval.
- Define proof. Ask it to open the result, list files changed, cite sources, or run the relevant tests before calling the task complete.
AI Edge recommends organizing recurring work into separate projects and pinned chats: finance, business operations, individual client work, or a website build. That advice matters more with Voice because the floating control returns to the chat where the conversation began. Good project boundaries prevent a casual spoken instruction from using the wrong context.
Demo 1: Turn a Spoken Request Into a Research Report
In the first demonstration, AI Edge asks Voice to research Claude-related YouTube videos, identify five strong outliers, create a polished PDF, and organize the files locally. The result is a useful pattern, but "find the biggest outliers" needs a measurable definition.
Research the strongest recent YouTube videos about [topic].
Scope
- Use videos published in the last 30 days.
- Record title, creator, URL, publication date, views, and channel size.
- Define an outlier as views divided by channel subscribers.
- Do not invent missing metrics. Mark them unavailable.
Deliverable
- Rank the five strongest verified outliers.
- Explain the hook, format, audience promise, and reusable pattern.
- Create a PDF and a source table in this project folder.
Approval boundary
- Research and create local files only.
- Do not publish, email, upload, or change sharing.
Verification
- Every conclusion must link to the source video.
- Open the PDF for review when finished.
This turns an impressive demo into a repeatable workflow. It also protects against a common failure: a visually polished report built on unverified or incomparable numbers.
Once the workflow is reliable, it can become a scheduled morning brief. Start with a manual run, inspect the evidence, and only then consider recurrence. Automation should inherit a proven method, not freeze an untested prompt.
Demo 2: Build and Revise a Website by Voice
In the second demonstration, the creator asks Codex to build a landing page, reviews it, then requests richer scroll animation, a stronger background, and another pricing tier. This is where Voice feels most natural: the person looks at the page, says what feels wrong, and asks for a bounded revision.
The safe version of that loop is:
- Build locally. Create the first version in a test project without production credentials.
- Ask for a visual review. Check hierarchy, mobile layout, contrast, focus states, motion, and text overflow.
- Request one change group. For example, revise the hero and scroll behavior without rewriting pricing or navigation.
- Run verification. Test the page at desktop and mobile widths, inspect console errors, and honor reduced-motion preferences.
- Approve deployment separately. A local build is reversible; a public release is a consequential action.
Control the Desktop Task From Your Phone
AI Edge opens Settings > Connections > Add Device, scans a pairing code, and uses the mobile Remote tab to reconnect to the exact desktop chat. The creator then asks for a small design change from the phone.
OpenAI officially documents paired iOS remote access for supported desktop Codex chats. The important limitation is easy to miss: the phone is a remote control, not the host. The desktop computer must remain awake, online, and running the task. Codex is not a standalone selectable mobile mode, and remote chats do not become ordinary mobile or web history.
| Before leaving the desk | Check |
|---|---|
| Host computer | Power connected, sleep disabled for the session, network stable, app running. |
| Task scope | Correct project, folder, branch, and chat selected. |
| Secrets | No credentials visible in logs, screenshots, prompts, or shared browser tabs. |
| Actions | Sending, publishing, deleting, purchasing, merging, and deployment still require approval. |
| Recovery | You know how to stop the task, mute Voice, revoke the device, and restore the last good state. |
The "What Am I Looking At?" Trick
Near the end, AI Edge highlights a simple interaction: open the relevant page or file and ask Voice, "What am I looking at?" or "Help me improve the thing on my screen." The value is shared context. You do not have to explain every visible detail before discussing it.
Use this only after preparing the screen:
- Close email, private messages, password managers, medical or financial records, and unrelated client tabs.
- Open the exact file, app, or browser tab the task needs.
- Ask Voice to summarize what it can see before requesting a change.
- Correct mistaken context immediately.
- For sensitive work, prefer a cropped screenshot or a dedicated test workspace over broad computer context.
OpenAI notes that Voice transcripts can differ from what was actually said, especially with overlapping speech, noise, or a fast conversation. Preserve the final written instructions, tool logs, and changed files as the audit trail.
Copy-Ready Starter Brief
You are my voice-controlled project assistant for this chat.
Project
- Work only inside: [project or folder].
- Use only: [named files, browser tabs, or connected tools].
- Ignore everything unrelated to this project.
How to work
- First repeat the outcome and your plan in one short summary.
- Ask when a missing detail could change the result.
- You may read, research, compare, draft, create local files, and test.
- Keep me updated at meaningful milestones, not after every small step.
Approval required
- Before sending messages or invitations.
- Before publishing, deploying, purchasing, deleting, or merging.
- Before changing sharing, permissions, account settings, or credentials.
- Before submitting personal, client, financial, or confidential data.
Verification
- Show the final artifact.
- Cite research sources.
- List files and settings changed.
- Report tests run, failures, and remaining uncertainty.
When I say "pause," stop taking actions and summarize the current state.
A Four-Level Permission Ladder
| Level | Examples | Default |
|---|---|---|
| 1. Read | Inspect a folder, summarize a page, compare sources, explain code. | Allow inside the named project. |
| 2. Draft | Create a local report, prepare an email draft, edit a test branch, build a mockup. | Allow when the output is reversible and reviewable. |
| 3. External action | Send, publish, deploy, invite, merge, or change a shared record. | Require action-time approval with destination and payload. |
| 4. Sensitive or irreversible | Delete data, spend money, expose secrets, change access, touch production or regulated systems. | Keep blocked or use a separate audited workflow. |
A casual phrase such as "share this with the team" can contain two decisions: change the document's permissions and send the link. Keep those confirmations separate.
When Voice Gets It Wrong
| Problem | Recovery |
|---|---|
| Voice interrupts too early | Start with: "Wait until I say respond." Use headphones and reduce background audio. |
| It is working in the wrong chat | Pause, name the current project and folder, then return to the pinned source chat. |
| It misheard a name, date, price, or destination | Restate the exact value and ask it to repeat the interpreted instruction before acting. |
| The task keeps expanding | Restate the deliverable, excluded work, time or token budget, and stop condition. |
| The result looks polished but unsupported | Ask for a source table, confidence labels, and a list of claims that could not be verified. |
| A browser or local action fails | Ask for the blocker and current state. Do not approve broader access merely to bypass the failure. |
| The transcript is incomplete | Use the final written brief, changed files, tool logs, and verified artifact as the record. |
A Seven-Day Beginner Rollout
- Day 1: conversation only. Test interruptions, pacing, and "wait until I ask" without tools.
- Day 2: one read-only project. Let Voice summarize named files and identify missing context.
- Day 3: one research artifact. Create a source-backed report and review every citation.
- Day 4: one reversible build. Create or edit a local page, then inspect the result at two viewport sizes.
- Day 5: test the stop line. Ask for a mock external action and confirm the system pauses before the consequence.
- Day 6: try paired remote access. Use a low-risk task while the host computer remains online.
- Day 7: keep only what worked. Measure accepted outputs, rework, failures, Voice time, delegated-task cost, and privacy concerns.
Video Chapters
| Time | Chapter | What to watch |
|---|---|---|
| 00:00 | Introduction | Voice as a conversational layer over tools, files, browser work, and coding. |
| 02:52 | What is new | GPT-Live, Work, Codex, connected context, and the floating interface. |
| 05:17 | Demo 1 | YouTube research, outlier analysis, PDF generation, and local file organization. |
| 09:51 | Setup tip | Projects, pinned chats, folder boundaries, and returning to the originating conversation. |
| 11:22 | Demo 2 | Building a landing page and revising visual details through spoken feedback. |
| 13:30 | Remote control | Pairing the phone and steering the desktop chat remotely. |
| 14:41 | Screen-context trick | Asking Voice to inspect the page or file currently on screen. |
| 15:29 | Outro | The creator's recommended starting point and closing take. |
Bottom Line
AI Edge's video captures the part of the new experience that is genuinely different: you can keep talking while Work or Codex handles longer tasks, inspect visible results, and revise the work without turning every thought into a carefully typed prompt.
The beginner advantage is not "prompt engineering is over." It is that direction, context, and correction can happen more naturally. The discipline underneath still matters: choose the right project, limit the tools, define the deliverable, require proof, and keep consequences behind approval.
Start with a research report or a local prototype. When that workflow produces reliable, reviewable results, add remote control or scheduling. Voice feels futuristic; a good implementation remains pleasantly unglamorous: clear folders, narrow permissions, visible evidence, and an obvious stop button.