Peter Yang's Tabi walkthrough starts with a useful constraint: build a language coach for an actual trip, not a feature list in search of a learner. The result is a phone-friendly Japanese practice app that speaks with you in real time. The most transferable lesson is the order of work: clarify the scenario, specify a small journey, build one working lesson, then improve both the interface and the teaching.
Tabi teaches ten Japanese phrases per lesson and then asks the learner to use them in a live role-play with a voice coach named Yuki. Peter builds it in six steps: explore, specify, build, iterate on the app, iterate on the lessons, and launch. Gemini 3.8 Live supplies the conversation. The video demonstrates a promising personal prototype, not a validated language course or a production-ready public service.
Watch Peter Build Tabi
Credit and disclosure: This six-step build, Tabi concept, and demonstrations are Peter Yang's. He discloses that Google sponsored the video and says the opinions are his own. Security, accessibility, and learning-outcome checks below are our editorial additions. Peter did not claim a controlled study of Tabi's teaching effectiveness.
What Tabi Actually Teaches
Peter frames Tabi around a Japan trip: ten lessons with ten practical phrases each. The lesson screen has Japanese, romaji, and English text bubbles plus a call button. In the live demo, Yuki first introduces phrases such as a greeting, listens to Peter repeat them, and then plays people he might meet in Tokyo. That two-part structure matters: recall inside a conversation is a different task from reading a phrase list.
Peter says the initial audience is himself and a few friends, not a paying subscriber base. He also notes that the API costs money. That makes the scope coherent: a small personal tool can test whether voice practice is worth extending before taking on user accounts, billing, moderation, and support.
Steps 1-2: Explore the Idea, Then Specify It
At 2:45, Peter asks an AI assistant to research before building and to ask three clarifying questions. His answers define the lesson shape, the Japanese/romaji/English presentation, and who will use the app. A good first pass should also pin down what counts as a helpful correction: should the coach interrupt, repeat, give a hint, or let the learner finish?
At 5:01, he turns that brief into a spec and a few representative screen designs: the journey, lesson, and onboarding flow across phone and desktop. Peter uses his own paid spec skill for part of this. You do not need that purchase to write the spec yourself. His separate free human-review skill supports direct comments and edits on HTML or Markdown drafts, so the human can correct the plan before code is generated.
Steps 3-4: Build a Thin Version and Observe It
Peter starts the first build in Google Antigravity at 8:24, using the Gemini Live API for low-latency voice interaction and a key from Google AI Studio. Google's docs describe real-time audio and video sessions over WebSockets; current model access, quotas, and prices should be checked before copying his setup.
The first version is not the finish line. At 10:03, Peter tests the flow and finds a wrong microphone selection, awkward message timing, and visual placeholders. He iterates on when chat bubbles appear, adds scene illustrations, and refines Yuki's avatar. The useful habit is to test a live call early: a beautiful lesson page cannot compensate for a silent microphone or feedback that arrives too late.
Step 5: Improve the Teaching, Not Just the Screens
At 12:28, Peter revises lesson sequencing, phrases, and the coach script. He notes that this content work probably deserved attention earlier than some of the visual polish. For a tutor, a smooth call is only a delivery mechanism. The practice needs plausible situations, useful corrections, and a progression that does not jump from a beginner greeting to a task the learner cannot yet perform.
For a first pilot, test one scenario with a human speaker of the target language. Check whether romanization is accurate and consistent, whether the coach understands hesitant speech, and whether it distinguishes a harmless accent from a meaning-changing error. Let learners repeat a phrase or read the transcript if audio is unavailable. Do not present an AI pronunciation judgment as a formal proficiency score without validation.
Step 6: Launch Without Exposing the Key
Peter deploys Tabi on Vercel at 14:57, fixes an outdated-library issue, and adds a passcode to limit who can trigger paid API calls. That is sensible for a friends-and-family demo, but a shared passcode is not a public-app security model. The tutorial also shows an API key during development that Peter says he deleted afterward.
Important distinction: keeping a key in a local .env file is helpful only if the key stays on the server. If a frontend build includes that value, visitors can recover it. Google's Live API ephemeral-token guide recommends issuing short-lived client tokens from a backend that holds the long-lived key; a server proxy is another option. Store server secrets in deployment-side environment variables, rotate any exposed key, and add spending and rate limits before a public launch. Voice recordings and transcripts also need clear consent and retention rules.
A Small Test Before Turning It Into a Product
Do not use minutes of conversation as the only success measure. Give five willing testers the same practical role-play before and after practice. Record whether they can understand the prompt, produce the target phrase without reading it, recover after a misunderstanding, and finish the task with confidence. Ask what the coach corrected incorrectly. Separately log connection failure, speech latency, cost per completed lesson, and requests for a text-only route.
That evidence tells you whether the next week should go into lesson design, speech interaction, or no further build at all. Peter's video is a strong example of getting to a testable prototype quickly; it does not establish that Tabi is superior to a human tutor, a conventional course, or free conversation practice.
Turn the Workflow Into a Business Idea
This prompt borrows Peter's narrow scenario, spec, voice prototype, lesson review, and staged launch. It works in a general AI assistant and asks for a buyer and a learning outcome before choosing a model or tool.
Find one practice scenario
Three ideas, then a seven-day pilot.
Video Chapters
0:00 Why voice · 1:18 Tabi demo · 2:45 Explore · 5:01 Spec · 8:24 Build · 10:03 Iterate on the app · 12:28 Improve lessons · 14:57 Launch.
Sources and Resources
- Primary creator source: Peter Yang's sponsored tutorial, published 5 October 2026, and the supplied transcript. Follow Peter on X or LinkedIn.
- Model and API: Google's Gemini 3.8 Live announcement, Live API documentation, ephemeral-token guidance, and AI Studio. Google sponsored the video; this post is not sponsored.
- Creator workflow: Peter's free human-review skill. His spec skill is discussed in the video as part of Behind the Craft, a separate paid offering.
- Deployment: Vercel environment-variable documentation. Public access, session cost, and learner data handling require more work than the private demo shows.