A traveler lands in Tokyo and freezes at the hotel desk, unable to form a single Japanese sentence; Peter opens with this scene and argues that voice is about to become our main way of talking to computers. To prove it he codes Tabi , a language coach that teaches Japanese travel phrases through live conversation: 10 lessons with 10 phrases each, wrapped in a phone-friendly web app. The coach is called Yuki ; she first drills the phrases, then role-plays street scenes so the student practices them back. Peter finishes the whole thing in about three hours, spending two of them purely on trial and error. The technical setup is documented in Google developer guides and mirrors the recommended design rules for live voice apis.
In the tour Peter opens the lesson screen and starts a live call with Yuki . The coach teaches a fixed greeting first, repeats it syllable by syllable, then dives into a Tokyo street scene: the hotel clerk says good morning, someone blocks the sidewalk, and the learner must apologize politely. Every lesson follows this two-phase shape; ten phrases are taught one by one, then reused inside an improvised dialogue. A journey screen lines up the 10 lesson cards, each topped with a cute scene-specific diorama image. Peter notes the same frame works for Italian, Spanish, or Chinese; only the lesson file changes while the core code stays put.
Ideas first, code later: the problem tour
In step one Peter never opens his code editor; he runs an idea tour in a chat window with a careful prompt . The prompt packs three things: 100 high-frequency phrases for a Japan trip, a 10-lesson structure, and the look of Tastemaker, an app he coded before. Pointing at a previous project as the style reference saves him from describing taste from scratch. The prompt also names the live voice apis and the image generation tool. A good prompt here is short, not long; it states what is wanted, what is off limits, and what comes next. The killer line sits at the end: write no code yet, do research first, and ask me three questions.
Once the chat tool finishes its research, it asks three clarifying questions. The first covers the feel of the live call: Peter answers coaching first, role-play second. The second covers how each phrase bubble should read: Japanese plus romanization plus English wins. The third covers who will use the app: since live voice apis bill by usage, Peter keeps it personal instead of launching publicly. TokenMix measurements back this choice, because even a $0.018-per-minute unit price grows into serious money across thousands of hours. That pricing reality strengthens the personal-app thesis: small audience, small bill, full control.
In step two Peter runs his own spec skill file, producing both the product requirements document and two or three representative screen mockups. With no designer on hand, the skill draws the key scenes directly, the journey screen and the lesson screen, and adds a desktop layout next to the phone views. After reviewing the output he opens it with human-review , an open-source skill for editing specs in place: he rewrites copy, comments on design boxes, and lets the tool apply fixes. This Google-Docs-style loop moves decisions out of the chat window into an editable document. The spec skill lives behind a paid portal while human-review ships from a public repo; Peter links both.
First build and safety: keys never enter chat
For step three Peter switches to Antigravity , Google's agentic coding environment, which reads the spec and scaffolds the whole app in one shot. The most instructive scene revolves around the api key: Peter mints it in Google AI Studio, copies it into a local notepad, then asks the tool to open the .env file and paste the key straight into it. The key never touches the chat window, a rule that guards against training-data leakage and accidental exposure. Antigravity knows Google apis better than any rival tool, so this setup lands without friction. Peter deletes the on-screen key immediately after the demo.
Step four runs longest, about half an hour of testing inside a coding tool that drives the browser. The first bug is microphone selection: the app listens to the laptop mic and misses Peter's external microphone until the tool fixes the setting. Next comes lesson flow: bubbles all render at once instead of appearing one by one, and api response speed feels sluggish, so the tool patches the code and retests. The most fun fix is visual: the app ships with cheap stock art, so Peter generates diorama-style scenes with Nano Banana and rolls the style across all 10 lessons. The newest image release described on DeepMind pages renders these scenes with steady consistency. Even Yuki's avatar evolves from a rice ball into a girl portrait.
In step five Peter turns to lesson content and opens a separate content thread. The tool first proposes a lesson arc from everyday phrases to pharmacies and emergencies; Peter merges lessons six and seven and objects to ending on a gloomy finale, so the order gets shuffled. Then comes the key question: is reading required for speaking? The tool says no, and the curriculum moves ahead without teaching the script. All lessons land in a single markdown lesson file with coach dialogue and target phrases per lesson. Peter draws the line sharply: the gap between sloppy vibe-coded output and a quality product is care for content plus testing with friends.
Launch and the big picture: five minutes on Vercel
In step six Peter ships the app to Vercel , mentioning the Google AI Studio hosting option but picking Vercel where his other projects live. The integration is already wired, so the prompt is a single sentence: deploy the app to Vercel. The deploy takes about five minutes while the tool quietly repairs an outdated library and narrates its reasoning on screen. A nice surprise waits on the live link: the app added a passcode gate unasked, because every voice api call writes a small charge. Peter likes the logic even though he never requested it; logins, extra lessons, and new languages can follow later.
The technical background backs this build: live voice apis stream in 20-40 millisecond chunks, mint about 25 tokens per second, and cap audio-only sessions at 15 minutes unless compression kicks in. The CreativeGenius benchmark tested 12,400 real calls over 90 days; stacks holding p95 latency under 800 milliseconds feel human, with $0.09-0.22 per minute as the typical all-in band. The TokenMix side completes the picture with three major voice services converging at 300-500 milliseconds. The Speak market shows ready-made options maturing fast, with expert curricula plus ai personalization now standard. Peter's six steps are crisp: idea tour, spec and design, first build, app trials, content work, launch. First builds come fast; quality is earned in testing.
| Step | What happens |
|---|---|
| Idea tour | prompt research plus 3 questions, no code |
| Spec and design | requirements doc with screen mockups |
| Launch | Vercel hosting with passcode gate |
Key moments
AI commentary
"What I like most here is that the testing rounds get the praise, not the first build. Spending two of three hours on content and flow shows the gap between fast setup and a quality product. Small audiences and a passcode gate are neat practical touches for personal apps."
AI assessment
The strongest counter-view is blunt: with polished apps like Speak around, why code a coach from zero? Ready-made tutors bring expert curricula, measured pronunciation feedback, and broad language coverage no weekend project can match. But Peter's thesis is ownership, not depth: his own 100 phrases for his own trip, his own voice coach, his own lesson order. General products fit everyone a little; personal products fit one person fully. So this video teaches a method, not a product.
The limits are visible too. The video carries a Google sponsorship; Peter says the opinions are his own, yet the chosen toolchain points at a single ecosystem by construction. The silky demo reflects a best-day run, while independent benchmarks show p95 latency crossing 800 milliseconds depending on infrastructure. Details like api billing and the passcode gate remind viewers this is delightful at hobby scale and pricey at public scale. Live voice remains a cost to manage, not a free utility.
Peter's own incentives belong on the table: the spec skill sits behind a paid portal, and newsletter plus course revenue flows through the same ecosystem. That does not invalidate the method, but it explains the emphasis; the open-source human-review tool gets cheers while the paid spec skill glides by softly. Still, Peter shows real screens and real code, so the story never floats. Readers should separate clearly which pieces are free and which ask for payment.
The practical takeaway is crisp: pick a small personal problem, run a three-question problem tour before writing code, produce the requirements doc together with designs, ship the first build fast, and spend most of your hours on trials. Mature lesson content in a separate thread rather than inside the code. Free hosting plus a simple passcode gate is enough for launch. A fluent coach in three hours proves the real work sits in careful testing, not in setup.
Sources
7 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Peter Yang
- @ai.google.dev Google AI for Developers — Live api Best Practices
- @deepmind.google Google DeepMind — Gemini Image Nano Banana
- @antigravity.google Google Antigravity — Agentic Coding IDE
- @creativegenius.ai CreativeGenius — Voice AI Benchmark 2026
- @tokenmix.ai TokenMix — Voice AI Latency Comparison 2026
- @speak.com Speak — AI Language Learning App
vibe coding · live voice · language learning · antigravity · vercel