The Human in the Loop Is Legal Fiction

Let me tell you something that most organizations do not want to hear.

When a company tells a regulator, we have a human reviewing every AI decision, that statement is technically true and practically meaningless. I have seen it firsthand across industries. The human is there. The human is clicking approve. But the human is not reviewing.

There is a difference, and that difference is where real risk lives.

We have invented a term for what is actually happening: liability laundering. The appearance of oversight without the substance of it.

The Three Levels of Human Oversight

When I work with organizations on AI governance, I use a three-level framework to help leadership understand where their oversight actually sits.

Level 1: Human In the Loop

A human reviews every individual AI decision before it is acted upon. This is what most compliance documentation describes. It is also the most resource-intensive model, and in high-volume environments, it is almost impossible to sustain at genuine quality.

Level 2: Human On the Loop

A human monitors system performance and sets parameters, but does not review individual decisions in real time. The human intervenes when patterns suggest something is wrong. This is a legitimate governance model when it is designed intentionally and supported by the right monitoring infrastructure.

Level 3: Human Outside the Loop

A human only gets involved when something breaks. No active monitoring. No meaningful review. Just a name attached to a process for compliance purposes.

Most organizations believe they are operating at Level 1. When I look at the actual workflows, the click rates, the time spent per decision, and the escalation rates, most of them are drifting toward Level 3. Some are already there.

How Human Review Quietly Disappeared

I worked with a healthcare organization that deployed an AI system for clinical decision support. Early results were strong. The system was accurate roughly 98% of the time, which in that context was genuinely impressive.

Months into deployment, the team noticed something. Employees were still reviewing every output manually, exactly as the protocol required. On paper, Level 1 oversight. In practice, something different was happening.

The review times had dropped to a few seconds per decision. Escalation rates had fallen to near zero. The approval rate was above 99.8%. Leadership looked at those numbers and drew what seemed like a logical conclusion: the AI is working, the reviews are confirming it, and we can streamline the process.

So they removed the mandatory review step.

The organization swung from no trust to blind trust without building anything in between. There was no governance model for Level 2 or mechanism to catch the cases where the AI was wrong in ways that did not show up in aggregate accuracy metrics.

The 2% error rate did not disappear. It just lost its last checkpoint.

Why People Stop Challenging AI Decisions

When I describe this pattern to leadership teams, the instinct is to frame it as a people problem. The reviewers got lazy. The team stopped paying attention. If we retrain them, the oversight will improve.

That instinct is wrong, and acting on it will not solve anything.

Complacency in human-AI systems occurs when humans interact with systems that are right almost all the time. We are not designed to sustain vigilant attention to processes that rarely require our intervention. That is not a flaw. It is how human cognition works.

We have seen this pattern before:

  • Commercial aviation introduced cockpit automation that made flying dramatically safer. It also produced a generation of pilots who, in documented incident reports, struggled to maintain manual flying skills and situational awareness during the rare moments automation failed.
  • Consumer vehicles with lane assist, adaptive cruise control, and collision warnings have made driving safer on average. They have also produced drivers who are measurably less attentive, because the car handles the moments that used to require focus.

Knowledge workers with AI tools are now entering the same dynamic. The system handles the routine. The human approves. The human's ability to catch the non-routine quietly degrades.

The question organizations need to ask is not whether this will happen. It will happen. The question is whether the governance model accounts for it.

Presence Does Not Prove Review

The conversation in boardrooms, in regulatory filings, in AI governance frameworks, is almost entirely focused on one question: should humans stay in the loop?

That question matters, but it is the wrong starting point. The more important question, the one almost nobody is asking, is what evidence do we have that the human is actually paying attention?

Meaningful human oversight is not a checkbox. It is a measurable activity. If an organization cannot answer the following questions, the oversight they are reporting is not real:

  • What is the average time a reviewer spends per decision?
  • What percentage of AI outputs does the reviewer escalate or override?
  • How does reviewer behavior change over time as familiarity with the system increases?
  • What is the error rate on the decisions the reviewer approves, compared to a baseline?

If those numbers are not being tracked, the organization does not know whether its oversight is functioning. It only knows that a human is present.

Presence is not oversight.

False Confidence Creates the Biggest Risk

I have a principle I share with every leadership team I work with on this topic.

The most dangerous person in an AI workflow is the one who thinks they are supervising.

The reviewer who knows they are not really checking is a known risk. Leadership can address a known risk. The reviewer who genuinely believes they are providing meaningful oversight, while clicking through decisions at a rate that makes real review impossible, is an invisible risk. That person is the gap between your compliance documentation and your actual exposure.

Build Governance Around Reality, Not Compliance

Fixing this does not require removing humans from AI workflows. It requires being honest about what level of the loop humans are actually operating at and designing governance accordingly.

  • If you are at Level 1, measure it. Track review quality, not just review presence. Build in friction that forces genuine engagement on a meaningful sample of decisions.
  • If you are at Level 2, design it intentionally. Define the monitoring criteria. Build the escalation infrastructure. Make the human role explicit and auditable.
  • If you are at Level 3, acknowledge it. Stop calling it oversight. Build safeguards appropriate to a system operating with minimal human intervention, because that is what you have.

The organizations that will build AI systems worth trusting are the ones that stop counting human presence as evidence of human oversight. The two things are not the same. And in the gap between them is where accountability goes to disappear.

This is only a preview.

The deeper insights, including how AI reshapes education, finance, leadership, cybersecurity, and communication, are inside Neil's Substack, where policymakers, founders, and Fortune 500 leaders get strategies they won't find anywhere else.

Read Disrupting the Box on Substack
← Back to all articles