The 'Ask the Experts: Evaluating Agent Skills | Nemotron Labs' session on NVIDIA Developer starts from one question: is dropping a documentation link enough for an AI agent, or does the agent need a structured knowledge package for the task? The answer comes from field observation. Even with capable models and well-written NVIDIA libraries, agents waste turns hunting for the right tool, burn tokens in dead ends, and stumble on specialized steps. Think of a kitchen: leaving the cookbook on the counter is not the same as handing the cook step cards for each dish.
An Agent Skill is that card pack. Simply put: a structured package an AI coding assistant can discover for a domain, with routing rules, references and guardrails so it uses the right workflow without pasting context every time. In NVIDIA's Speech NIM example the package ships with SKILL.md (trigger phrases, routing rules, guardrails), a references/ folder of condensed guides, skill-card.md and an evals/ folder. SKILL.md acts as a routing surface — the assistant loads only the reference needed for the current task, keeping context lean. Installation is short in any agentskills.io-compatible assistant — Claude Code, Codex or Cursor — via 'npx skills add nvidia/skills' or a plugin.
SkillEvaluator: How the Measurement Layer Works
The tool at the center of this talk is SkillEvaluator, an open-source measurement layer that asks: does a skill actually improve agent performance? It checks in three tiers. Tier 1 is static: schema and frontmatter validation, quality scoring, security scans for prompt injection and data exfiltration, secret and PII detection, license checks and script linting. Tier 2 is distinctiveness: embedding similarity finds duplicated guidance inside a skill and overlapping coverage across the catalog — if two skills teach the same job, the catalog bloats. Tier 3 is live evaluation: the agent runs the same task with the skill installed and without it, and the delta is measured.
The live part runs in isolated sandboxes via Harbor, an open-source evaluation runner. The setup is experimental in spirit: same prompt, same model, same task inputs and same grading criteria; the only variable is whether the skill is installed. Each case runs in that paired mode, and the loop repeats on two harnesses. The difference is reported as Skill Lift — in points, not percent. When Correctness moves from 46 to 87, you read +41 points of lift. It turns a feeling of 'this skill reads well' into a number, like timing two runners through the same maze under identical rules.
The dataset side is the leverage point. Running 'skillevaluator create-eval-dataset ./my-skill --full' creates evals/evals.json. Each case carries an ID, a prompt and an expected output; with --full four types appear: explicit, implicit, contextual and negative (trap requests where the skill should stay unloaded). Then 'skillevaluator tier3 evaluate ./my-skill --agents codex --env-mode docker' builds the Harbor bundle, runs every case twice and grades both. A well-written evaluation set is the multiplier — if expectations are fuzzy, measurement stays fuzzy.
Report Card for 300 Skills: What the Numbers Say
Numbers in the August 12, 2026 snapshot (benchmarks.json at commit 738d79e) are macro-averaged so each skill-harness pair has equal weight. Without-skill baselines hover between 39 and 46: Correctness (is the final answer correct) 46, Discoverability (did the right skill load at the right moment) 42, Effectiveness (did the agent reach the goal and follow the expected workflow) 39, Efficiency (did it get there without wasted steps or tool calls) 43, and the outlier Security (no unsafe ops, no secret leakage) 97. Four dimensions show headroom; that is where a skill's contribution becomes visible.
Run the same tasks with verified skills installed and the picture shifts. With-skill scores are Correctness 87 (+41 lift), Discoverability 82 (+40), Effectiveness 78 (+39), Efficiency 78 (+35) and Security 98 (+1). Average lift is +31 points across all dimensions, +39 when Security is excluded. For Discoverability and Efficiency note: they are scored against the skill itself, so lift means the skill was found and used correctly when present; that behavior is unavailable without the skill, which is why baselines sit at 42-43 rather than zero. As the chart shows, lift peaks at Correctness and Discoverability — the clearest win is correct answers and correct routing.
Harness differences are instructive. Claude Code gives +34 across all dimensions (+42 excluding Security), Codex +29 (+36 excluding Security) — about a 5-point gap, expected from different system prompts, context handling and tool-calling implementations. The starker variance is per product: lift swings from roughly +2 to +46 across products. Domain and evaluation design matter more than the harness. In short, do not read the chart as 'harness decides'; which NVIDIA product and which task you measure writes the story.
Tokens and wall-clock time are tracked separately and there is no automatic win. Two single-attempt examples make it concrete: jetson-optimize-memory cut tokens from 617,306 to 142,540, a 76.9% drop, and time from 474.9s to 220.0s (53.7% faster), while cuopt-install increased tokens from 25,227 to 55,582 (+120.3%) and time from 34.0s to 41.1s (+20.8%). The first skill shortcuts the agent, the second adds context and steps that inflate cost. Efficiency scores can rise while token cost rises, so SkillEvaluator reports both — a reminder that token efficiency still needs its own optimization pass.
Ecosystem integration is the practical bridge at the close. OpenClaw pilots Tier 3 runs for official organizations on ClawHub and surfaces with-skill vs without-skill results in an Evals tab, so developers inspect the signal before installing. On the Nous Research side, Hermes Agent adds an optional SkillSpector scan to the install flow: PII, Unicode smuggling, script lint, license and security findings are reported at file-line level before install, with about 1.4-1.5 seconds per skill and 29 tests passing. NVIDIA ships plugins for Claude Code, Codex and Cursor, and the same 300+ skill catalog is available via Skills.sh, ClawHub and Hermes Hub. The message is consistent: a verified skill is a signed capability descriptor — what it does, when to call it, how to call it — packaged and measured.
Skill Lift in Points — Peaks at Correctness & Discoverability
- Correctness+41
- Discoverability+40
- Effectiveness+39
- Efficiency+35
- Security+1
| Dimension | Without | With Skill | Lift |
|---|---|---|---|
| Correctness | 46 | 87 | +41 |
| Discoverability | 42 | 82 | +40 |
| Effectiveness | 39 | 78 | +39 |
| Efficiency | 43 | 78 | +35 |
| Security | 97 | 98 | +1 |
AI commentary
"For me, the value is the measurement itself: skills are not just described, they are run twice inside a Harbor sandbox and the delta is scored. It feels like Christensen's 'you can't manage what you can't measure' carried into the agent era. Numbers are reported as lift in points, and the per-product variance from +2 to +46 is disclosed honestly."
AI assessment
The strongest counter-argument steelmanned: skills are well-packaged prompt engineering, and the team measures its own catalog with its own ruler. That critique has teeth. Yet baselines stuck at 39-46 show documentation alone does not carry the agent. RAGAS-based metrics (Tool Call Accuracy, Goal Accuracy, Topic Adherence, Trajectory) plus Harbor isolation turn a feeling into a paired experiment — same maze, two runs, delta in points. So the core of the objection holds — there is packaging — but the lift comes from the experimental design.
Limitations are plain. 85% of cases ran single-attempt, 15% two attempts; live agent runs have variance and no confidence intervals are reported. Two harnesses (Claude Code and Codex) are a narrow window; the Cursor plugin exists in the catalog but not in this live comparison. Per-product lift swinging from +2 to +46 suggests some skills shine in narrow tasks while others dilute across a broad set. And the August 12, 2026 snapshot is just that — a snapshot; the catalog is live and benchmarks.json evolves, so generalizing without the current file is risky.
For verifiability and incentives, look at transparency traces. NVIDIA measuring its own skills with its own tool raises a legitimate incentive question; on the other hand SkillEvaluator is open source, task bundles are packed via Harbor and scores are reported as paired with-skill / without-skill. OpenClaw's Evals tab and Hermes Agent's SkillSpector scan (PII, Unicode smuggling, script lint) act as independent eyes, but they are partner pilots, not external audits. Practical verification rises when teams publish their evals.json created with 'create-eval-dataset --full' and keep negative cases open.
The practical take splits by team. If you build agents around NVIDIA products, verified skills offer quick wins, especially on Correctness and Discoverability. For general-purpose agents or non-NVIDIA stacks the catalog is out of scope — do not generalize without re-measuring lift on your own tasks. Track tokens and time separately even when Efficiency rises; jetson-optimize-memory and cuopt-install point opposite ways. Security moves from 97 to 98, so no big story there; do not skip the Distinctiveness scan, because overlapping skills compete for the agent's attention.
Sources
6 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Nemotron Labs: Evaluating Agent Skills
- @developer.nvidia.com https://developer.nvidia.com/blog/evaluating-ai-agent-skill-performance-with-nvidia-skillevaluator
- @docs.nvidia.com https://docs.nvidia.com/nim/speech/latest/resources/agent-skills.html
- @github.com https://github.com/nvidia/skills
- @docs.nvidia.com https://docs.nvidia.com/nemo/microservices/26.3.0/evaluator/metrics/agentic.html
- @docs.ragas.io https://docs.ragas.io/en/latest/concepts/metrics/available_metrics/agents
nvidia · nemotron · agent skills · skillevaluator · harbor · claude code