Back to feed

Google's WikiSkill Framework: Teaching Agents Your Company's Rules by Themselves

Google Research's WikiSkill lets an AI agent teach itself a company's unwritten rules: it writes each mistake into a readable playbook and keeps the change only when scores rise on unseen work. A laptop demo on messy sales sheets goes from a seventy percent baseline to ten out of ten with no fine-tuning.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — ewxQLr7IxzY
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Google Research published a study called WikiSkill that is making the rounds, and the name misleads at first: it is not a single skill but a framework for agentic applications. The core metaphor is a new hire on their first week. The newcomer makes mistakes, writes down what went wrong, and stops repeating the error, all without going back to school. Here the underlying language model is never retrained and no fine-tuning happens; instead the framework keeps rewriting the agent's instruction file until the results hold up.

The problem it attacks is unwritten company rules. A finance newcomer may know spreadsheets inside out and still fail the first month, because the decisive conventions live only in colleagues' heads. In the demo, one case means six units where the usual standard would be twelve, and a returned product is entered as a negative revenue figure. No model can infer such local habits from general knowledge, however capable it is, because they were never written anywhere.

The remedy is a loop with a strict acceptance test. Whenever the worker fails, a lesson describing the failure goes into the instruction file, and the change survives only if scores improve on work the agent has never seen. Each round therefore ends with a verdict of keep or discard, and the file keeps evolving until the whole set passes. Automation replaces the human who would otherwise hand-edit the instructions after every mistake.

For the demo the presenter rebuilt this loop on his own laptop around a fictitious company's messy month-end sheet, driving everything from a chat window built on Google's agent development kit. Asking the agent to introduce itself reveals the architecture: a worker that reads the messy files, writes a small cleanup script, runs it and saves a clean file; a checker that grades the result against known-correct numbers; a note taker that records each failure; a playbook writer that converts notes into a single instruction change; and a strict manager that tests that change on ten unseen sheets and keeps it only when the score rises.

The test data itself is deliberately ugly, the way real sales exports look. One table mixes several date formats, leaves gaps, spells region names inconsistently, abbreviates units in incompatible ways, and mixes capitalization in product labels. Returns appear as negative revenue lines that a newcomer would easily misread. Anyone who has merged branch-office spreadsheets will recognize the genre instantly.

The finance team's specification adds a second difficulty: it demands sorting by date, region and order id plus standard finance conventions for dates, units, returns and duplicates, yet nobody defined what the local standard actually is. Run without any learned guidance, on day-one behavior, the agent spots missing region values and lists failures but cannot resolve them, which is exactly the point: detection without the local rulebook goes nowhere.

This is where the paper's diagram, three co-workers plus one strict manager, earns its keep, and where the comparison with older techniques lands. Retrieval needs a document to look up, but here the rules exist in no document. Fine-tuning takes days and its lessons cannot be read back. Plain memory records things but never checks whether the memory was correct. WikiSkill learns from errors, writes the lesson in readable prose, and retains it only when it demonstrably helps, which is why the presenter calls the outcome special.

Training ran on twenty files: ten for practice and ten held-out sheets for the manager's verdict. The visible scoreboard starts from a seventy percent baseline, rejects the eleventh and twelfth attempts at forty percent, and finally accepts the seventh iteration with ten out of ten. Attempts three to six collapsed for a mundane reason, exhausted usage quota on the presenter's own account, which is a setup problem rather than a framework failure but still instructive about operating cost.

Then comes the payoff scene: the same file and the same model, this time with the learned guidance loaded. The run passes, every row matches the expected ground truth, and returns are handled with the correct sign. Best of all, the resulting guidance is plain readable prose stating when each finance rule applies and when it does not, including habits the agent inferred on its own rather than receiving them ready-made.

Who is this for? Any team running repetitive work where errors come from house habits rather than world knowledge: month-end closing where each store's export has its own quirks, or invoice routing where only the payables team knows the right cost center. The presenter assembled the whole experiment by describing it in plain language to a coding assistant, uploading the paper, and building on Google's agent kit, and he points to a community implementation plus the paper link for replication. Long-running agents with stable routines gain the most; one-off tasks gain the least.

Visualization: nodesdaily AI

AI commentary

"What hooked me here is the framing: the hardest rules in a company are the ones nobody wrote down, and WikiSkill turns that problem into a loop the agent runs by itself. I find the demo more convincing than the buzzwords around it."

AI assessment

The strongest objection I can steelman is distribution lock-in: the validation sheets are siblings of the training sheets, so a perfect score may only prove the playbook memorized one family's quirks. Real store exports drift every quarter, and the paper's five controlled evaluation sets do not measure that drift either. A ten-out-of-ten result on twenty related files deserves applause, not trust.

What the video does not test is cost, latency and safety. Every candidate change is re-run over ten unseen sheets, which burns compute and time; the presenter himself ran out of account quota mid-demo, which tells its own story. A worker that writes and executes its own scripts against finance data is also a destructive-command risk with no human approval step in the loop I saw.

On verification: this is a single-person demo where the same hand built the mess, defined the correct answers and judged the result, so the checker's comparison can be circular. The paper's figures need independent replication before I would quote them in a decision, and the presenter's employer disclaimer, while honest, still leaves this as a one-source story. I would want a second implementation's numbers before committing budget.

My practical read, in the first person: for repetitive back-office work with stable formats this is a valuable learning layer, and I would pilot it read-only with human sign-off on every playbook change. For one-off analysis or creative work the setup cost does not pay back. The readable playbook is the real prize for me, because an auditor can inspect a document in a way no weight update allows.

Sources

7 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

wikiskill · ai agents · google research · adk · skill evolution · fine-tuning · enterprise automation

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…