Back to feed

Stop Building Agents, Start Building Skills: Four Rules from Anthropic Engineers

Anthropic engineers Barry Zhang and Mahesh Murag stopped building one agent per job and now attach reusable skills to a general agent, with four rules: save proven code, describe sharply, learn permanently, verify before delivery.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — HIRDzMtuWFk
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

The two engineers behind the Agent Skills format at Anthropic, Barry Zhang and Mahesh Murag, have stopped building a separate agent for every job. Their reasoning is simple: the underlying agent turned out far more general-purpose than expected, so dressing the same body with job-specific knowledge is enough. To me this is the most practical course correction of the past year.

The phone analogy sums it up in one line: the model is the processor, the agent runtime is the operating system, and the skill is the app. Nobody assembles a separate phone per app; you install the right app on the same device. Since Claude Code already reads files, writes code and calls tools, the same body serves everything from slide decks to company research, and only the attached skill changes.

The first rule is my favorite: stop making the model re-solve solved technical problems. The team saw Claude rewriting nearly the same Python script from scratch every time it styled slides, which burned tokens and produced a different result on each run. The fix was to store the working script inside the skill folder so later runs execute the proven file directly. Software's DRY principle, transplanted here. My workshop follows the same pattern: I never let a script that produced an output I liked rot inside a chat. The recipe from the video matches: save the script into the skill's scripts folder, update SKILL.md so future runs execute that file, then run the same task twice and compare the important parts. The surrounding prose may still vary, but the critical piece is no longer a fresh guess, it is tested code.

The second rule solves discovery: once dozens of skills pile up, how does the model pick the right one. The answer is staged loading called progressive disclosure: Claude first sees only each skill's name and description, the YAML front matter. When the wording of a request matches a description, the full SKILL.md of that skill is read, while large companion files wait in the folder until truly needed. Irrelevant instructions never bloat the working context.

But this mechanism lives or dies on description clarity. Two vague descriptions like content help and marketing asset production force the model into a coin toss. A strong description states the job plus the real user phrasings that should trigger it. The triple test from the video is practical: an obvious triggering request, the same request worded differently, and an unrelated request that must never trigger. A skill that cannot be found might as well not exist.

The third rule makes corrections permanent: every fix-it-then-close-the-chat loop throws a lesson away. Saying only fix it repairs the output but leaves the process broken. Instead the model is asked to trace back: where did it search, why did it miss, should the routing or the skill be updated. Wrong process means updating instructions, missing voice or examples means adding a reference file, a repeating mistake means adding an explicit blocking rule. Then the same task is re-run to verify the fix.

An honest boundary is drawn here: no skill can make a weaker model perform like a stronger one, since every model interprets instructions differently. What travels is the process. Because Agent Skills is an open format, the same folder can be tried on compatible runners, and wherever it falls apart you hunt for hidden assumptions. The open-standard announcement of December 2025 formalized this portability.

The fourth rule, billed as the most important: a skill should never hand its first attempt to the user. Finishing seventy percent of a job and dropping a brokenly formatted, unsourced file on a human is the common failure. The cure is to bake the checks you already know into the skill: render slides visually and fix overflows, match research claims against primary sources and strip what cannot be proven, read drafts through beginner and skeptic eyes. My favorite pattern knots into one sentence: define acceptance criteria first, build version one, inspect it with the fitting method, fix every issue found, run a second pass, and withhold the result until the criteria hold. Whatever cannot be verified must be stated openly. Where an objective success metric exists, agents keep working until it is hit. That way the user's first look is never the model's first look, it is its fourth or fifth.

To sum up, the picture is crisp: save proven code, make each skill discoverable with a sharp description, convert corrections into lasting instructions, verify with evidence before delivery. Together the four rules turn a general-purpose agent into an assistant that knows how I work. That is the closing line of the video as well: build not the agent, but how the agent works for you.

Some context belongs here too: Agent Skills was announced in October 2025 for Claude Code, explained through a real example like PDF editing, and shared in a public repository. The concept spread to other platforms including Codex. So the video describes not one channel's idea but a pattern that became industry practice within months, with the four rules as its distilled form.

Visualization: nodesdaily AI

AI commentary

"I find this the most actionable AI-workflow video in months: four rules I can apply to my own assistant this week."

AI assessment

Let me steelman the strongest objection generously: the praise of skills fades under realistic conditions. In a study by UC Santa Barbara and MIT researchers over 34,000 real skills, the benefit turns fragile once distractor skills and a noisy pool enter the picture, barely beating the skill-free baseline in the hardest scenarios. Current SKILLSBENCH setups flatter the picture by handing agents hand-picked skills that amount to a solution guide.

There are also fronts the video never tests: retrieving the right skill from a crowded pool, weaker models performing even worse once skills are attached, and the responsibility-boundary debate. The SRCP-titled entry in the public repository argues skills carry no semantic responsibility frame. None of this refutes the four rules, but it explains why the second and fourth rules are the places that break most often.

I will also weigh who says what: the primary basis is Anthropic's engineering post and platform docs, the concept author's own account. The critical basis is independent researchers and press. The presenting channel has a course-and-community interest, so the closing course invitation deserves a cautious read. The video's numerical claims rest on the channel's own experience and call for independent replication.

My verdict is this: for small teams with repeating work these four rules are a directly applicable package, with the first and fourth rules paying off from week one. For one-off jobs the cost of building a skill outweighs the return, and a plain prompt stays the smarter move. In my own flow I will give every repeating job its skill and keep every one-off job lean.

Sources

9 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

agent skills · anthropic · claude code · progressive disclosure · ai agents

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…