Back to feed

Tmux + Fable: A Practical Way to Cut Token Bills by 35%

AI Jason shows a two-tier setup where Fable 5 plans and Sonnet 5 executes, plus managing every coding agent from one place with tmux; Devin Fusion measurements put the saving around 35%.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — wCSPgHpcxdc
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Everything starts from a familiar pain: running the strongest model for every job burns the cloud quota in days. The cause is simple, frontier models are strong but equally pricey and slow; routing every small fix through them inflates the bill. The fix is not leaving the model but changing the division of labor.

The formula fits one sentence: Fable 5 plans, Sonnet 5 runs. As presented in the video, Sonnet 5 costs roughly one fifth of the Fable 5 level while performing near the top model of a few months ago. Planning quality stays high while the bill drops.

Two official patterns are described on the Claude Code side. First, Sonnet 5 acts as main executor and calls Fable 5 as advisor when stuck. Second, Fable 5 acts as orchestrator, builds the plan, and opens worker sessions on the smaller model for execution. The comparison by the Devin team favors the second path.

The reason is cache economics: an advisor model must re-read the full history of the main agent, which means pricey fresh input. The worker pattern instead triggers cached context, and a cache read costs roughly one tenth of fresh input. That is why the orchestrator setup runs cheaper.

The key distinction here is sidekick versus one-shot sub-agent. A classic sub-agent closes on finish, and a fix request reopens it with zero context so the same things get rewritten. A persistent sidekick session keeps history, the main agent sends follow-up messages into the same session; history arrives from cache, so it is both cheap and strong. Agent teams plus the send-message tool inside Claude Code do exactly this.

Implementation step one: add a delegation rule to the CLAUDE.md file. The rule says you are the coordinator, design and planning and review stay with you; hand execution to Sonnet workers; workers never open nested agents; for complex work first write a frozen spec into the task folder. In the video a to-do app is ordered this way: questions are asked, the plan is clarified, the spec lands in a file, execution goes to a Sonnet session, and the result arrives visibly faster than running on Fable.

Implementation step two: the Codex bridge. Once the Codex plugin is installed and plugins are reloaded, a Codex session can start from inside Claude Code; rescue, review, result, cancel, and transfer commands come ready. The resume command sends follow-up messages into the same Codex session. In the video the to-do interface first written by Claude Code gets a Codex pass and the look visibly changes.

Implementation step three: the tmux universal bridge. Tmux is a terminal multiplexer; it opens several persistent terminals and controls each by command. Three commands suffice: split-window -h opens a new pane on the right, send-keys -t types text plus Enter into the target, capture-pane -p -t reads that pane back. Any agent you use like a human, including Gemini CLI, becomes a worker under the main agent.

The one missing piece is finish notification: how does the worker wake the main agent. On the tmux side the answer is the wait-for command; the worker signals with wait-for -S on finish while the main side waits. The open agent teams skill in the video wraps this pattern: open a session, send a message, wait for finish, read the result. In the demo Codex writes a joke, a Pi agent reviews it, Haiku checks grammar; all run in parallel and the final text returns to the main agent.

For those wanting a packaged option, Orca and Herd are covered. They serve the same orchestration with an interface: worker sessions popping on the right, a hierarchy view on the left, kanban plus token tracking. The presenter settled on Orca as daily driver; open source and free is the plus. Scripted tmux path or packaged interface, the logic is identical: the pricey brain plans, cheaper hands execute.

AI commentary

"I do not read this video as run away from the expensive model but as use the expensive model only where it earns its keep — I copied that distinction into my own CLAUDE.md and my quota stopped melting."

AI assessment

To steelman the other side: running everything on one strong model is simpler, faster to debug, and for a small team time beats hardware costs. Paying a few dollars of API instead of babysitting stuck local sessions is rational for most teams; I take that objection seriously, it narrows rather than refutes the video thesis.

The video carries no measurement of its own: the 35% figure comes from the Devin bench, not my machine and not my repo. Demo tasks are small, to-do app and interface touch-up level; messy multi-file real repo work is never tested. Permission screens and stuck tmux panes are glossed over, yet in practice most time is lost exactly there.

There is also interest and verifiability: price ratios and cache discounts are seasonal and partly vendor-sourced, changing monthly. At decision time I re-check the current API price page and cache-read rates; I read every figure in the video as a claim to re-measure on decision day.

My practical takeaway: this setup is a genuine option for learners, tinkerers, anyone keeping data on-device, or juggling several agents. It is not for single-agent flows, production secrets on free endpoints, or teams that cannot grant terminal access to agents. On my side the orchestrator plus worker pattern stayed permanent, the tmux bridge stayed for experiments.

Sources

8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

tmux · fable 5 · token savings · agent teams

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…

Tmux + Fable: A Practical Way to Cut Token Bills by 35%