Back to feed

Claude Code from Scratch: Models, Effort Levels and a Sub-Agent Setup That Scales

Avenox walks newcomers through Claude Code from first install to orchestration, showing why terminal and desktop share the same core, how Sonnet, Opus and the premium tier fit different jobs, what effort levels really buy you, and how sub-agents keep long sessions light alongside limits, permission modes and a layered CLAUDE.md setup.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — cT-52mKI8I4
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

The video opens on Claude Code’s two doors, terminal and desktop, and makes a simple point: they share the same CLAUDE.md, the same settings and the same extension layers. A job started in the terminal can continue on the desktop because state lives in one configuration tree. For now the Claude team leans toward the terminal while the Codex team leans toward the desktop, which explains the small feature gaps. For a newcomer with some terminal comfort, warming up in the terminal feels natural, but the final choice is taste and flow. Setup mirrors that simplicity, a one-liner for macOS and Linux, a one-liner for Windows, and typing /desktop in the terminal hands the session to the desktop app.

The model family is the backbone. Four names are mentioned, yet Haiku is set aside for practical work, leaving Sonnet, Opus and the top planning tier on stage. Names may shift with releases such as 5.5, but the logic stays: Sonnet is economical for everyday chores, Opus is the balanced generalist, and the premium tier is the orchestrator that loves to plan. The host puts Sonnet at roughly a third of Opus on price and notes an expected Opus price trim around the video’s release, a move that makes choosing by task rather than by habit the rational default.

Opus is the host’s favourite for a reason that is not hype but fit. It handles a wide band of work without feeling overpriced, adapts quickly to new context and keeps its voice pleasant over long stretches. That makes it the everyday driver, especially on a Pro plan around the $20 mark where Opus plus Sonnet covers most needs. Its generalist nature removes the need for extra choreography on small jobs. The takeaway is to let the balanced model carry the main line and reserve the premium tier for planning rather than craft.

The premium tier is framed differently, as a chief of staff. The host calls it the smartest and the best at orchestrating, fluent in conversation and strong at shaping a plan and steering workers under it, yet not always the best hands for direct coding. The preference is to design architecture and breakdown at the top and hand execution to flocks of Opus agents. Price and quota reinforce that split: Max comes in roughly $100 and $200 steps, the premium share is about half the weekly allowance, and the boost in the five-hour window does not translate one-to-one to the weekly budget. Keeping planning at the top and production in the middle protects both budget and headroom.

A practical split shows how tiering saves real money. Designing a schema migration lives with Opus, while renaming a variable across many files, an atomic and repetitive chore, lives with Sonnet. Support roles such as gathering codebase facts, cleaning a large context or summarising search results also benefit from Sonnet’s speed and cost edge. Solving everything with a single heavy model inflates both the bill and the context for no gain. The lesson is that cost is not only a price list but a discipline of matching model to task texture.

Effort Levels: How Much Thinking to Buy

The effort ladder opened by /effort is one of the clearest explanations in the video. Low walks a single path and tries few branches, enough for simple tests. Medium is presented as the balanced middle for those sensitive to limits. High and xHigh sit near today’s defaults and branch more deeply, buying more search at higher cost. Max pushes thinking for the same job to an extreme that the host finds excessive, since thinking steps bill closer to output rates and can run about five times the price of plain input. The advice is to stay on High or xHigh by default and reach for Max only when the task truly warrants it.

The most confused label is ultracode, which the video corrects: it is not a step above Max but the same thinking depth as xHigh with a different trick, cloning the agent into a swarm. That colony runs in parallel under a special coordination scheme, which explains why limits melt quickly. The host shares a run that coordinated more than 200 agents under one ultracode umbrella, a scale that cannot be steered by hand or even by a single top agent collecting context. Writing workflow or ultracode in the prompt cues the system to enter that multiplying mode. Powerful in the right place, wasteful in the wrong one.

Limits are summarised as three counters: a five-hour session window, a weekly budget and, on Max, the premium share. The gap between the $100 and $200 steps is not fourfold on the week, about fourfold in the five-hour window and roughly two and a half fold on the weekly budget in the host’s estimate. Overuse is flagged sharply: pay-as-you-go via the API at these levels can sting fast. For tracking, /usage and the in-app meters are enough. The message is plain: adding parallelism buys speed but spends quota at the same rate.

Permissions, Queues and Undo: The Safety Dials

Permission modes are presented as a ladder of trust. Manual lets reads through and asks before edits. Accept edits lets edits through and still asks before running commands, since commands carry risk. Plan mode researches and proposes a plan for approval before any work begins. Auto mode puts a safety model behind the scenes that judges whether an action is risky and blocks only when needed, the recommended start for most users. Bypass permissions opens every lock and can, in theory, wipe the machine, which the host reserves for operators who truly know what they are doing. Newcomers are urged not to jump to the freest mode early.

The message queue and undo tools are small features with outsized impact in long sessions. While a job runs, new messages sit queued until the next tool call is read. ESC pauses, double ESC opens a rewind menu with three choices: rewind both code and conversation, rewind conversation only, or rewind code only. The pair /branch and /fork cover a common need to try variants without reteaching valuable context, one moving you to the new line and the other keeping you put while the fork opens elsewhere. For interface experiments this saves real time and keeps the history useful.

Four special run modes are layered by need. Plan, via Shift+Tab or /plan, is an exploration and approval loop that stops when the machine sleeps. Goal is a persistence mode that does not stop until a separate judge model says the objective is met, the pattern behind leave-it-overnight stories. Loop repeats the same job on a cadence — every minute, five minutes or hourly — illustrated with scanning new commits. Schedule runs routines on Anthropic’s servers whether your machine is on or off, for example checking a repo daily and merging or closing by rules. The host rarely needs Schedule because similar loops already run on a personal server, but the option remains for those who want cloud persistence.

CLAUDE.md Hierarchy and Keeping Context Lean

CLAUDE.md is corrected from a single file to a hierarchy. It reads from the home directory down through the project root and into subfolders, so an agent loads the chain up to where it works. In a monorepo that means a general file at the top and focused files in backend, frontend and database folders, loading only what the current task needs. When context swells, /context shows spend and /compact trims fat. This layering saves tokens and reduces confusion because the agent sees only the instructions relevant to its layer rather than a single bloated brief.

The strongest architectural idea is the sub-agent. One agent doing everything fills its window and slows down, while dividing work across sub-agents lets each handle a slice and return only a summary to the top. The host places the premium planner on top and flocks of Opus, with Sonnet where it helps, underneath. In ultracode the picture gets more abstract: the agent writes code that describes how the swarm should run, and that code drives the hierarchy so the top context never bloats even while hundreds of agents coordinate. That is the sustainable path at scales where hand orchestration is not feasible.

The closing section separates concepts and draws a starter map. CLAUDE.md is memory read every session, a skill is a capability loaded only when called, connectors and MCP are bridges to the outside world, hooks are enforced guarantees that fire on events, and plugins are standardised bundles of the same. Comfort features follow: visual self-check via the desktop’s browser, /voice dictation that does not consume quota, and /statusline customisation for a personal bar that can show branch and spend. Parallel agents can be requested in plain language, for example three Sonnet plus two Opus workers, and the host closes with a simple ramp: on Max keep Opus on top, on Pro keep Sonnet on top, start in auto, run /init in the first project to seed CLAUDE.md, then shape the bar and write the first skill.

Visualization: nodesdaily AI

AI commentary

"What sticks with me is the video’s blunt rule not to throw the most expensive model at every job: do planning up top and let the next tier do the craft. That split protects both budget and context window, and in long coding sessions it keeps the work architectural rather than a chain of memorised commands."

AI assessment

Steelman the counter-argument and the video looks a touch too tidy: as abstraction rises, felt control drops, and the quota math is not as transparent as it seems. The strongest version of that critique says cost is not the sticker price but the price of thinking, billed closer to output and roughly five times input, plus the way parallelism spends quota at speed. On that view the real risk is a newcomer turning effort and ultracode too high too early and burning through the allowance before learning the ropes. The clash is between an experienced power user who is at home with one main model and a newcomer who loses budget control early.

Limits keep the enthusiasm in check. The premium share on Max and the cost of thinking are overkill for small projects, and the 200-agent story is inspiring but not practical for most teams. The bypass warning is well placed, and the note that hooks and MCP once bloated context but now load on demand is reassuring, yet backwards compatibility and misconfiguration remain. Visual self-check in the desktop app is a real edge for interface work, but expecting the same smoothness on Windows or on a modest machine is not realistic. The video touches on Codex trade-offs, yet it speaks in the language of personal preference rather than a neutral benchmark.

Verifiability and incentives paint a mixed picture. The host is candid about long-term real use and price expectations, yet claims about a coming discount and larger quotas are time-bound and should be cross-checked against official announcements. Product and docs pages provide solid ground as sources, but there is no independent reproduction of performance claims. There is no obvious conflict of interest; the tone is practitioner, not sales. Still, viewers should test quota and price items against their own usage and keep the official pricing page as the source of truth.

The practical take splits by audience. For someone at home in the terminal and working with a team, the suggested ramp makes sense: Sonnet on top for Pro, Opus on top for Max, start in auto, run /init in the first project and then grow a layered CLAUDE.md tree with a first skill. Keeping planning at the top, craft in the middle and saving ultracode for large, parallelisable work builds a quota-aware discipline. For desktop-heavy visual work or low-budget experiments it is safer to stay on High or xHigh by default and to open loop and schedule only when truly needed. In short the video proposes not a single correct path but a scalable style of work.

Sources

6 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

claude code · anthropic · ai agent · terminal · opus sonnet

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…