I have lived the same moment in every AI governance meeting: someone worries about a machine touching a high-stakes decision, someone else offers the human in the loop as reassurance, and the tension in the room dissolves. It sounds like a solution. But a person glancing at an output before it ships does not mean the system is governed; most of the time it means responsibility has been laundered through a signature. Phaedra Boinodiris, known for her work inside IBM, calls this picture liability laundering, and the companion essay published under the IBM name spells out the warning.
Boinodiris starts by reminding us how people trust each other. How hard you interrogate a colleague's report depends on whether that person has earned your trust; trust is never automatic, it accumulates over time. It rests on five intuitive behaviors: grounding outputs in real evidence, producing consistent results under the same conditions, choosing methods fit for the decision at hand, working for the good of the team rather than yourself, and remaining the same person when conditions change.
Each of the five maps onto a technical property of AI. Evidence becomes transparency: what data and methods fed the model, and were those choices validated? Consistency becomes explainability: why did it produce this output for this person today? Fitness of method becomes observability: in agentic AI especially, is the system still performing correctly as conditions shift? Self-orientation becomes algorithm optimization: which goals and whose interests the model's tuning truly serves? Continuity becomes robustness against adversaries: has the model been tampered with?
The observability emphasis is no accident: an agent pursuing a goal calls tools, chains decisions, and the risk of drifting out of sight grows with every step. The governance patterns published by the Dataiku teams therefore embed oversight into the flow itself: pause-and-resume checkpoints at critical thresholds, approval gates for risky actions, and escalation paths that hand control back to a person on failure. The picture flagged by Dataiku research is sobering: only five percent of data leaders say AI outputs are traceable every single time. Where there is no trail, there is no oversight.
The law has already moved in this direction. Article 14 of the EU AI Act requires high-risk systems to be designed so that real people can oversee them. Commentaries published by IAPP underline the point: oversight must ensure the system is used as intended and its effects are addressed over its lifecycle. Airia draws the distinction even more sharply: a generic approve button on a screen is weak oversight; strong oversight is policy-driven, recorded, and enforceable. The statute does not say let a human glance at it — it says the human must genuinely see and be able to stop it.
Yet Boinodiris insists the human in the loop contributes something no model replicates: reading circumstances, not just data. A case file can tell one story while a life tells another; the question is not what the output says but what this decision means for this person in this life. Social workers, ethicists, historians, counselors — groups told for decades that their expertise was too non-technical for technology decisions — are exactly the people who ask that question. The human approval at high-risk steps urged by the Synclovis analysis rests on the same intuition: we built systems affecting millions of lives without including anyone who studies lives.
The closing offers a four-question test: can the reviewer see the model's reasoning, does she have the authority to stop the process, are the right measurement instruments and feedback loops in place? If not, what you have is not governance but counting warm bodies in chairs. The ending is not bleak, though: a human who can see the reasoning, stop the process, and bring what no model replicates is the beginning of governance. The question is never whether there is a human in the loop but whether that human actually has a job to do.
AI commentary
"I wrote this piece as a mirror for teams that mistake a checkbox for governance; I want the four-question test hanging on meeting-room walls."
AI assessment
Let me steelman the other side: this frame hands a heavy bill to everyone who takes oversight seriously. Approval gates add latency, reviewers fatigue, and tired reviewers rubber-stamp everything — the effect the literature calls automation bias, which turns the human from safety valve into rubber stamp. In small teams, passing all four tests can be a luxury. Boinodiris would answer that the problem then lies not in the frame but in the appetite for autonomy: do not ship what you cannot oversee.
A second reservation concerns measurability: transparency, explainability, and observability are crisp on paper and expensive in practice. Full disclosure of training data and optimization targets is often a closed box in vendor models, leaving the user to oversee reasoning she cannot see. That is why the frame should become a procurement criterion: a vendor that offers no traceability fails the governance test at the door. Oversight begins with product selection.
Sources
6 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube Watch on YouTube
- @ibm https://www.ibm.com/think/insights/liability-laundering-problem-human-in-the-loop-not-governance-strategy
- @dataiku https://www.dataiku.com/blog/human-in-the-loop-ai-agents
- @iapp https://iapp.org/news/a/eu-ai-act-shines-light-on-human-oversight-needs
- @airia https://airia.com/blog/human-in-the-loop-enterprise-ai-controls
- @synclovis https://www.synclovis.com/blog/the-role-of-human-in-the-loop-in-agentic-ai-governance
artificial intelligence · governance · agent oversight