AI Tools

Apple Should Have Built VoiceOS: 10 Tools Making Agents Usable

Direct Answer

The most important launch in Andrew Warner and Corey Ganim's roundup is not a new model. It is a better interface between a person and an agent. VoiceOS turns pointing and speaking into actions across desktop apps. Grok Bot gives persistent workers their own cloud computers. OmniWork organizes creative specialists around deliverables. Annotate lets someone show a coding agent what needs changing instead of writing a brittle paragraph.

That explains the title "Apple should have built this." Siri made voice familiar, but VoiceOS demonstrates the interaction many users now expect: understand the visible context, preserve the task, act in the correct application, and ask before consequential steps. Whether VoiceOS itself becomes the winner is less important than the interface shift it represents.

Practical verdict: the agent market is moving from general chat windows toward specialized control surfaces. Test the interface that removes the biggest communication bottleneck in your workflow, then judge it on accepted output, correction time, permissions, and cost - not on how human its demo feels.

Watch the Episode

Credits: this roundup and its demonstrations come from Andrew Warner and Corey Ganim on The Next New Thing. Product descriptions below are checked against current official pages where available. Smaller creator projects are labelled as demonstrations because a social post is not the same evidence as stable product documentation.

The Shared Pattern: Reduce Translation Work

People already know what they want surprisingly often. The friction is translating that intent into the form software expects: typed prompts, tickets, timelines, tool calls, lesson plans, or operating procedures. Each useful launch in this episode removes one translation layer.

Human intentOld translationNew interface
"Handle this on my computer"Open apps, copy data, click through stepsVoiceOS voice-to-action
"Own this recurring job"Repeated prompts and manual follow-upGrok Bot durable worker and routine
"Produce this creative campaign"Coordinate specialists and move filesOmniWork expert-agent team
"Teach it in a way I understand"Search through generic tutorialsScrimba Explain personalized lesson
"Change that part of the interface"Write a long UI bug descriptionAnnotate synchronized screen and voice brief
"Track one family routine"Adapt a broad productivity appSmall purpose-built baby tracker

The model remains important, but the product advantage increasingly sits around it: capture, context, permissions, memory, orchestration, review, and a clear output. Those layers determine whether an agent fits real work.

What Is Ready to Test?

LaunchWhat it isEvidence levelBest first test
VoiceOSDesktop voice-to-action assistantPublic product, pricing, privacy, and security pagesCalendar, reminders, and low-risk document retrieval
Grok BotPersistent agents on cloud computersOfficial product and documentationRead-only research or a review-list routine
OmniWorkOrchestrated creative-agent workspacePublic product, plans, and case descriptionsOne content brief with a human review gate
Scrimba ExplainPersonalized explainer experiencePublic experience plus creator demonstrationOne concept you already know well enough to grade
AnnotateLocal screen-and-voice briefs for coding agentsPublic product and feature documentationOne visual UI correction in a test branch
SloptrimCreator-built AI writing cleanup utilityCreator post and demoCompare specificity and meaning before and after
Watermark removerTool intended to remove AI text marksCreator post; high trust sensitivityDo not operationalize for evading disclosure
Baby trackerPurpose-built family tracking appCreator-built demonstrationValidate data entry and privacy before real use
ShibadokuVibe-coded puzzle-game experimentEpisode demonstrationTest whether the mechanic is fun, not just generated
Operations.comStrategy, operating-system, and AI implementation servicePublic service and claims with disclaimerVerify one measurable operational bottleneck

1. VoiceOS: Voice Becomes an Action Layer

VoiceOS is the clearest example of a consumer interface catching up with agent capability. Its public site describes two modes. Dictation turns speech into edited text; Agent Mode uses screen context and connected capabilities to send email, schedule meetings, find files, create notes, file tickets, search, and manage reminders. The cursor can serve as context, reducing the need to explain which visible object matters.

The design opportunity is substantial. Voice is faster than typing for ambiguous, evolving thought, but traditional voice assistants force people into rigid command grammar. An agent can accept a richer instruction, inspect context, ask a question, and carry the work across apps. That is closer to how an assistant should feel.

The trust boundary is equally substantial. VoiceOS says its raw audio and transcripts are deleted after processing by default and are not used for model improvement unless the user explicitly enables its Help Improve setting. Under that opt-in, covered data may be retained for evaluation or training. Its security page says the company is working toward SOC 2 Type I; that is not the same as claiming a completed certification. Review current settings before exposing legal, health, financial, customer, or credential-bearing material.

Start with reversible work: reminders, personal notes, calendar proposals, and file retrieval. Keep sending, deletion, payments, account changes, and external commitments behind confirmation. Measure how often the assistant chooses the correct app and action, not merely transcription speed.

2. Grok Bot: Give Each Job Its Own Computer

xAI describes Grok Bot as durable AI teammates running on persistent cloud computers. Each bot can have a name, job, tools, working context, routine, and approval boundary. Bots can run in parallel, use websites and applications, learn a demonstrated task, and continue working while the user's computer is off.

The product lesson is job separation. xAI's documentation recommends creating a distinct bot when the work has a different goal, tool set, working style, approval boundary, or recurring schedule. That is much stronger than one universal assistant with access to everything. An account-health bot should not inherit the authority of an expense bot simply because both use the same model.

Cloud computers also widen the attack and error surface. A first pilot should use test accounts, narrow connectors, read-only research, a low spend limit, and a review list. Record what the bot accessed, what it changed, how much it cost, and whether a person accepted the output. Persistence is useful only when the organization can also pause, inspect, and revoke it.

3. OmniWork: A Creative Team Instead of One Creative Chat

OmniWork positions itself as an agent operating system for creative work. Its site presents specialist agents for trend monitoring, content analysis, video editing, film direction, music, voice, account growth, and other creative roles. A coordinator assigns work, gathers outputs, moves through playbooks and review gates, and retains taste, context, standards, and project history across agents.

This is the right direction for complex creative production. Research, concept, script, edit, quality control, publishing, and performance analysis have different acceptance criteria. Separating them produces a visible workflow and lets a person intervene at the expensive decisions.

The risk is ceremonial multi-agent complexity. Five agents do not automatically create better work than one well-briefed model. Test a single deliverable both ways. Compare idea quality, factual accuracy, edit minutes, visual consistency, total cost, and elapsed time. Keep the multi-agent version only if the handoffs improve the result.

4. Sloptrim: Editing Should Restore Meaning, Not Hide AI

The episode shows Sloptrim as a small utility for removing generic, robotic AI-writing patterns. The useful problem is real: model-generated drafts often contain vague intensifiers, repetitive transitions, padded conclusions, and sentences that could belong to any company.

A cleanup tool should not merely swap forbidden phrases. It should ask what specific claim, example, constraint, or opinion belongs in their place. The strongest editing loop is source draft, factual check, point-of-view pass, structural edit, and final voice review. If the output becomes shorter but loses the author's actual meaning, it is not improved.

Treat Sloptrim as a creator-built experiment unless its own product documentation establishes broader guarantees. Compare it against manual edits on a small set, and preserve a before-and-after record so stylistic cleanup cannot silently alter facts.

5. Scrimba Explain: Personalized Teaching Needs an Accuracy Test

Scrimba Explain turns a question into a customized explainer experience. The episode's appeal is that a learner can request a concept in a preferred style instead of searching for the one pre-recorded tutorial that happens to fit.

Personalization can improve pacing, examples, assumed background, and format. It cannot guarantee correctness. Test the system first on a concept you already understand. Check whether definitions are accurate, examples actually demonstrate the rule, visuals match the narration, and the explanation distinguishes simplification from fact.

The best educational agent also includes retrieval practice: ask the learner to explain the idea back, solve one transfer problem, and identify what remains unclear. A polished explainer is an input to learning, not evidence that learning happened.

6. Annotate: Show the Agent the Bug

XAnnotate's current public page uses the simpler name Annotate. It records one or more screens, keeps drawings and spoken instructions synchronized with the visible frames, stores the material locally, and exposes the session to Cursor, Claude Code, or Codex. The page says the agent reads frames and speech through MCP rather than receiving one giant video dump.

This solves a familiar problem in AI-assisted development: "move that," "the menu closes too early," and "make this state match the previous screen" are difficult to express without spatial and temporal context. A recording preserves the click path, timing, cursor position, and spoken intent.

The handoff still needs a finish line. End every session with acceptance criteria: which route, viewport, state, behavior, and test prove the change is done? Then make the agent inspect the code, propose a plan, work in a branch, and verify the rendered result. Visual context makes the ticket clearer; it does not replace review.

7. AI Watermark Removal: A Trust Problem Disguised as a Utility

The roundup includes a creator tool intended to remove invisible marks from AI-generated text. There are legitimate reasons to study watermark robustness: false positives, transformations, accessibility, research, and interoperability. There is also an obvious misuse: stripping a mark specifically to conceal AI involvement or defeat a platform's disclosure system.

I would not add watermark removal to a content pipeline. Machine-readable marks are imperfect evidence, and absence of a mark does not prove human authorship. The durable answer is editorial ownership: verify the facts, rewrite from real experience, preserve sources and revision history, disclose material AI use where required, and accept responsibility for the published result.

Removing provenance can also conflict with contracts, policies, or audience expectations. Treat any such utility as security-sensitive research, not as a routine content-polishing step.

8 and 9. The Baby Tracker and Shibadoku: Small Software Wins Through Fit

The creator-built baby tracker and Shibadoku game are modest compared with cloud agent platforms, which is precisely why they matter. AI coding tools make it economical to build software around one household routine or one playful mechanic instead of pursuing a giant horizontal market.

The baby tracker should be judged as a sensitive data product. Feeding, sleep, diaper, health, and family information require clear storage, export, deletion, access, backup, and sharing choices. A quickly generated interface is not enough for real family use. The product needs reliable timestamps, offline behavior, correction paths, and explicit limits on medical interpretation.

Shibadoku combines familiar puzzle ideas with a distinctive visual theme. Its test is simpler: does the mechanic remain fun after the novelty of AI generation wears off? Retention, difficulty progression, input quality, load time, and accessibility matter more than how quickly the first version was coded.

10. Operations.com: The Service Layer Around AI Adoption

Operations.com represents a different type of AI product: implementation as a service. Its site targets established service businesses and says it combines strategic consulting, an operating system, and hands-on AI and automation deployment. The underlying argument is sound: adding agents to a broken process can make confusion move faster.

The company publishes strong performance claims on its own site, including documented client growth, efficiency gains, and time-to-result statements. Those are first-party marketing claims, not independent evidence established by this article. Its own income disclaimer says results vary and are not guaranteed. A buyer should request definitions, baselines, case methodology, client references, scope, data access, ownership, support, and the exact metric used to declare success.

The strategic lesson is still useful: diagnosis, process ownership, decision rights, training, monitoring, and measurement are often more valuable than installing another tool. This is where many AI agencies should mature.

The Emerging Agent Usability Stack

  1. Capture: voice, screen recording, gestures, files, or events express intent.
  2. Context: the system retrieves the relevant app, project, memory, or source.
  3. Specialization: a named agent or workflow owns one bounded job.
  4. Execution: tools, browsers, cloud computers, or code perform the work.
  5. Review: approval gates, annotations, tests, and rubrics catch mistakes.
  6. Learning: accepted patterns become versioned skills or scoped memory.
  7. Measurement: cost, correction time, accepted output, and business results decide whether the system stays.

A strong product can own one layer or coordinate several. The dangerous product hides the layers behind a human metaphor while asking for broad access. The best ones make state, permissions, evidence, and approval visible without forcing users to understand agent infrastructure.

A Seven-Day Test Plan

  1. Day 1: choose one repeated task with a measurable current time and error rate.
  2. Day 2: select one interface - voice, cloud bot, creative team, or visual brief - rather than installing every launch.
  3. Day 3: create a test account or project, reduce permissions, and define actions that always need approval.
  4. Day 4: run five representative cases, including one ambiguous and one failure case.
  5. Day 5: record completion time, correction minutes, wrong actions, cost, and accepted output.
  6. Day 6: repeat the same cases with the previous workflow or a simpler agent.
  7. Day 7: keep the tool only if it removes meaningful translation work without creating more review, access, or support burden.

Video Chapters

TimeChapter
00:00VoiceOS: dictate, schedule, email, and take action
02:06Grok Bot and persistent cloud computers
03:27OmniWork creative agent teams
04:39Sloptrim and generic AI writing
06:00Scrimba Explain personalized lessons
08:06Annotate screen recordings for coding agents
10:21AI text watermark removal
11:42Vibe-coded baby tracker
12:45Shibadoku puzzle game
13:39Operations.com AI implementation service

Verdict

Andrew Warner and Corey Ganim found a useful collection because the tools do not all compete in one category. Together they reveal the product work happening around models: voice capture, durable workers, creative orchestration, personalized explanation, visual feedback, tiny vertical apps, and services that rebuild the surrounding operation.

The next winning AI products may feel less like AI products because they ask users to describe less infrastructure. They let someone point, speak, demonstrate, approve, and receive a finished artifact in the place where the work already lives. The standard should remain high: clear boundaries, inspectable evidence, reversible actions, privacy choices, and measurable improvement over the old workflow.

Sources and Credits

Common questions

What is VoiceOS?
VoiceOS is a desktop voice-to-action assistant from WakoAI Inc. It combines dictation, screen context, search, reminders, and user-directed actions in apps such as email, calendar, notes, files, and browsers. It is available for macOS and Windows; its pricing page currently lists mobile as coming soon.
Does VoiceOS keep voice data private?
VoiceOS says raw audio and transcripts are deleted after processing by default and are not used for model improvement unless the user enables its Help Improve setting. If that setting is enabled, covered content may be retained for evaluation or training. Review the current policy and settings before using it with sensitive work.
How is Grok Bot different from ordinary Grok chat?
xAI describes a Grok Bot as a durable AI teammate with a defined job, its own conversation, persistent working context, connectors, routines, approvals, and a cloud computer that keeps working when the user's laptop is closed.
What is OmniWork?
OmniWork describes itself as an agent operating system for creative work. It coordinates specialist agents for activities such as trend monitoring, content production, video editing, music, voice, and account analysis, with shared memory, playbooks, and review gates.
What is XAnnotate called now?
The current product page uses the name Annotate. It records a screen walkthrough with drawings and synchronized speech, stores the session locally, and hands it to Cursor, Claude Code, or Codex through an MCP-based workflow.
Should AI watermark-removal tools be used?
Not to misrepresent authorship, bypass disclosure, or erase provenance that another party expects to remain attached. Detection marks are imperfect, but deliberately stripping them can create trust, policy, contractual, or evidentiary problems. Edit AI-assisted work because you own and verify the final result, not merely to defeat a detector.
Which tool in the episode should a small team test first?
Annotate is the lowest-risk starting point for teams already using coding agents because it improves task handoff without granting a new cloud agent broad account access. VoiceOS is attractive for personal productivity, while Grok Bot and OmniWork need a more deliberate permissions, cost, and review pilot.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call