Google Research published a study called WikiSkill that is making the rounds, and the name misleads at first: it is not a single skill but a framework for agentic applications. The core metaphor is a new hire on their first week. The newcomer makes mistakes, writes down what went wrong, and stops repeating the error, all without going back to school. Here the underlying language model is never retrained and no fine-tuning happens; instead the framework keeps rewriting the agent's instruction file until the results hold up.
The problem it attacks is unwritten company rules. A finance newcomer may know spreadsheets inside out and still fail the first month, because the decisive conventions live only in colleagues' heads. In the demo, one case means six units where the usual standard would be twelve, and a returned product is entered as a negative revenue figure. No model can infer such local habits from general knowledge, however capable it is, because they were never written anywhere.
The remedy is a loop with a strict acceptance test. Whenever the worker fails, a lesson describing the failure goes into the instruction file, and the change survives only if scores improve on work the agent has never seen. Each round therefore ends with a verdict of keep or discard, and the file keeps evolving until the whole set passes. Automation replaces the human who would otherwise hand-edit the instructions after every mistake.
For the demo the presenter rebuilt this loop on his own laptop around a fictitious company's messy month-end sheet, driving everything from a chat window built on Google's agent development kit. Asking the agent to introduce itself reveals the architecture: a worker that reads the messy files, writes a small cleanup script, runs it and saves a clean file; a checker that grades the result against known-correct numbers; a note taker that records each failure; a playbook writer that converts notes into a single instruction change; and a strict manager that tests that change on ten unseen sheets and keeps it only when the score rises.
The test data itself is deliberately ugly, the way real sales exports look. One table mixes several date formats, leaves gaps, spells region names inconsistently, abbreviates units in incompatible ways, and mixes capitalization in product labels. Returns appear as negative revenue lines that a newcomer would easily misread. Anyone who has merged branch-office spreadsheets will recognize the genre instantly.
The finance team's specification adds a second difficulty: it demands sorting by date, region and order id plus standard finance conventions for dates, units, returns and duplicates, yet nobody defined what the local standard actually is. Run without any learned guidance, on day-one behavior, the agent spots missing region values and lists failures but cannot resolve them, which is exactly the point: detection without the local rulebook goes nowhere.
This is where the paper's diagram, three co-workers plus one strict manager, earns its keep, and where the comparison with older techniques lands. Retrieval needs a document to look up, but here the rules exist in no document. Fine-tuning takes days and its lessons cannot be read back. Plain memory records things but never checks whether the memory was correct. WikiSkill learns from errors, writes the lesson in readable prose, and retains it only when it demonstrably helps, which is why the presenter calls the outcome special.
Training ran on twenty files: ten for practice and ten held-out sheets for the manager's verdict. The visible scoreboard starts from a seventy percent baseline, rejects the eleventh and twelfth attempts at forty percent, and finally accepts the seventh iteration with ten out of ten. Attempts three to six collapsed for a mundane reason, exhausted usage quota on the presenter's own account, which is a setup problem rather than a framework failure but still instructive about operating cost.
Then comes the payoff scene: the same file and the same model, this time with the learned guidance loaded. The run passes, every row matches the expected ground truth, and returns are handled with the correct sign. Best of all, the resulting guidance is plain readable prose stating when each finance rule applies and when it does not, including habits the agent inferred on its own rather than receiving them ready-made.
Who is this for? Any team running repetitive work where errors come from house habits rather than world knowledge: month-end closing where each store's export has its own quirks, or invoice routing where only the payables team knows the right cost center. The presenter assembled the whole experiment by describing it in plain language to a coding assistant, uploading the paper, and building on Google's agent kit, and he points to a community implementation plus the paper link for replication. Long-running agents with stable routines gain the most; one-off tasks gain the least.
AI commentary
"What hooked me here is the framing: the hardest rules in a company are the ones nobody wrote down, and WikiSkill turns that problem into a loop the agent runs by itself. I find the demo more convincing than the buzzwords around it."
AI assessment
The strongest objection I can steelman is distribution lock-in: the validation sheets are siblings of the training sheets, so a perfect score may only prove the playbook memorized one family's quirks. Real store exports drift every quarter, and the paper's five controlled evaluation sets do not measure that drift either. A ten-out-of-ten result on twenty related files deserves applause, not trust.
What the video does not test is cost, latency and safety. Every candidate change is re-run over ten unseen sheets, which burns compute and time; the presenter himself ran out of account quota mid-demo, which tells its own story. A worker that writes and executes its own scripts against finance data is also a destructive-command risk with no human approval step in the loop I saw.
On verification: this is a single-person demo where the same hand built the mess, defined the correct answers and judged the result, so the checker's comparison can be circular. The paper's figures need independent replication before I would quote them in a decision, and the presenter's employer disclaimer, while honest, still leaves this as a one-source story. I would want a second implementation's numbers before committing budget.
My practical read, in the first person: for repetitive back-office work with stable formats this is a valuable learning layer, and I would pilot it read-only with human sign-off on every playbook change. For one-off analysis or creative work the setup cost does not pay back. The readable playbook is the real prize for me, because an auditor can inspect a document in a way no weight update allows.
Sources
7 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube Video: AI with Surya — WikiSkill demo
- @arxiv https://arxiv.org/html/2608.27454
- @emergentmind https://www.emergentmind.com/papers/2608.27454
- @developers.google https://codelabs.developers.google.com/codelabs/production-ready-ai-with-gc/3-developing-agents/build-a-multi-agent-system-with-adk
- @adk.dev https://adk.dev/workflows/
Also cited by: The End of the Single Giant Prompt: Getting Started with Graph Engineering on ADK 2.0
- @arxiv https://arxiv.org/html/2607.13104
- @arxiv https://arxiv.org/html/2606.31270v1
wikiskill · ai agents · google research · adk · skill evolution · fine-tuning · enterprise automation