Back to feed

I Made My Terminal Code by Voice: A Pocket Jarvis with Codex Voice

The experimental voice mode in Codex CLI 0.155 promises terminal coding by talking; the presenter uses it to assemble a small voice agent that operates his computer in about five minutes.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — Bo9iotMDDLw
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Coding by talking in the terminal

I tell my computer what I want, and it builds websites, apps, and small tools in front of me inside the terminal. It sounds like science fiction, but the experimental voice mode in OpenAI Codex CLI 0.155 promises exactly that: I speak instead of typing, and the agent writes the code. The presenter sums up the experience in one line, and I followed his path to test the claim piece by piece: setup, live trial, and building my own voice agent .

First Codex is updated, then /experimental is typed in the terminal to switch on Voice conversations , Codex restarts, and the chat begins with the /voice command. I cross-checked this sequence against the official 0.155.0 notes compiled by hammerautomation, and the steps match exactly; microphone control uses /voice mute, and /voice stop exits. One caveat matters: voice support is absent from some builds, and clients without the native voice runtime hide the related menus.

From setup to first trial

The presenter first dictates a broad request: a new page for his community site. While the agent codes, he asks questions about the design and receives feedback; moving forward without touching the keyboard feels genuinely impressive. Then comes a small sports-themed app trial, with an honest limit on display: reading the screen stays visibly slow. I read this experimental label together with the warning in the gptmap review; behavior varies by build and environment, so identical results on every machine should not be expected.

The critical point must not be missed: speech arriving through the microphone is a draft instruction shown on screen, and it must be read before any files change. The same approval model as typed directions applies; a spoken sentence creates neither authority nor an exception to workspace rules. The Codex desktop-use guide at chatgptaihub states the same rule: sensitive operations such as file deletion and settings changes require explicit user approval. Names, paths, and flags can be misheard, so reviewing each request is mandatory.

My own Jarvis in five minutes

The most entertaining part of the demo starts here: by voice, the presenter produces a small voice agent running in the browser in about five minutes. He presses the start-conversation button, grants microphone permission, and asks for the search engine to be opened; the agent opens the page. This is the computer use capability described in the mindstudio roundup: moving the cursor, pressing buttons, filling in forms; a universal connector for jobs with no shortcut. The ready-made tool ships with a community link, which I note in a single sentence.

The most practical trait of this mini agent is that it asks for no separate API key; it inherits the Codex permissions directly. Voices can be swapped, the interface restyled, and extra tools plus MCP connections added. The 0.156 additions noted by ccleaks complete the picture: voice arrives switched on by default, toggled with F8, the /usage panel shows spend, and worktrees come enabled by default. In other words, the presenter's setup became the out-of-the-box behavior of the next release.

I tracked the release calendar through the journal at greenlitbooks: 0.155.0 on September 17, a small fix the next day, 0.156.0 on September 22, and 0.159.1 at month end; voice, task view, worktrees, and session resume landed piece by piece across that window. Because the presenter's recording appeared in early October, his account matches the 0.156 world. I liked the analogy in the remio analysis: the terminal is turning into a command center gathering scattered command-line work.

To sum up, voice mode does not remove the keyboard, but it gains speed on spoken work such as recipes and bug reports. Slow screen reading and missing documentation annoy, yet the core claim checks out against the release notes. In my setup, voice joined the keyboard instead of replacing it; I will run the first trial on a small, reversible task.

Visualization: nodesdaily AI
StepTip
SetupUpdate first, enable voice in /experimental, start with /voice
SafetyRead a spoken order like a typed one, never run unreviewed
AgentA browser agent in five minutes, no extra API key

Key moments

  1. Opening claim: a Jarvis in the terminal
  2. Setup sequence and voice commands
  3. Sports-app trial and slow screen reading
  4. Five-minute agent build
  5. Search-engine opening and computer control

AI commentary

"Voice coding sounds like a toy, yet the picture here is serious: a terminal that obeys speech after minutes of setup. Still, every utterance must be read on screen; the microphone deserves no blind trust. My verdict is warm but cautious: worth trying, not worth handing over blindly."

AI assessment

The strongest counter-argument is unevenness: microphone, terminal, remote connection, and corporate policy change the outcome, so critical work should first be trialed on one's own machine. The tension in the remio analysis knots here too: as use grows easier, observability and isolation matter more.

Gaps exist as well: official documentation is nearly absent, the Plus and Pro requirement rests on the presenter's guess, and screen-reading speed is still raw. Price and quota can be watched through the /usage panel, but the token cost of long voice sessions stays unclear.

The presenter's interest is visible: the free zip file and ready-made tool point toward the paid Profit Boardroom community, so I read the enthusiasm at reduced volume. The core claim nevertheless stands on verifiable release notes.

The practical takeaway for readers: voice input saves real time on short recipes and bug reports, but every order must be approved through the on-screen draft. Starting small and keeping approval in hand is the safe-use prescription for this feature.

Sources

8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

codex · voice coding · terminal · computer use · ai agent

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…