“Human in the loop” has become one of the most reassuring phrases in artificial intelligence.
Ask an organization how it governs an AI-assisted process and, sooner or later, someone will say that a human reviews the decisions.
That sounds responsible. It can also be almost meaningless.
Consider an insurer that introduces AI into a workflow that previously produced 1,000 decisions or recommendations a week. With AI, the same operation can suddenly produce 5,000. The review team does not become five times larger.
So the organization samples. Perhaps humans review 10% of outputs. As volume increases, maybe that becomes 5%. Eventually, the organization can point to a documented human-review process while the overwhelming majority of AI-assisted decisions pass through without meaningful scrutiny.
The problem is not sampling itself. The problem is confusing a sample of decisions with a system for governing decisions.
That distinction matters as insurers put AI deeper into underwriting, claims, servicing, fraud detection and other consequential workflows. The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers contemplates AI across these activities while emphasizing that existing legal obligations continue to apply regardless of the technology used.
The denominator changed
Traditional quality-assurance programs were designed around human-scale production. A supervisor could review a meaningful portion of an employee's work. Patterns emerged. Coaching followed. Exceptions could be investigated.
AI changes the denominator. It can increase the number of recommendations, drafts, classifications and decisions far faster than an organization can increase the number of people available to inspect them.
More automated decisions + the same review capacity = less meaningful human oversight per decision.
NIST's 2026 work on monitoring deployed AI systems identifies scaling human-driven monitoring alongside rapid rollouts as a barrier to effective AI monitoring. NIST also notes that post-deployment monitoring is necessary because systems operate under changing real-world conditions that controlled pre-deployment testing cannot fully reproduce.
An organization can therefore technically maintain a human audit while steadily reducing the actual strength of the control. That is why I have argued that “human in the loop” is not, by itself, a governance model. The presence of a person somewhere in a workflow tells us very little about whether that person has the authority, information, capacity or responsibility necessary to exercise judgment.
The more important question is not: Did a human review some of the AI's work? It is: What decisions are we allowing AI to influence or make, what could happen if it is wrong, and what evidence tells us the control system is working?
Start with decision rights, not audit percentages
Many AI governance programs begin in the wrong place. They start by choosing an audit percentage: 5%, 10%, 20%. But there is no universally meaningful percentage divorced from the decision being audited.
The NAIC's framework points in a different direction. It says an insurer's controls should be commensurate with the nature of the decision, the potential harm to consumers, the extent of human involvement, the transparency and explainability of the outcome, and reliance on third-party systems or data. Controls for a particular use case should align with the degree of potential consumer harm.
That is fundamentally a risk-based approach.
Before deciding how often humans should review AI decisions, an organization should decide which decisions AI is allowed to make in the first place.
In my work on Decision Debt, I use a three-tier Decision Rights Charter:
- Delegate. The machine may make the decision because the pattern is stable and the consequence of error is sufficiently controlled.
- Augment. AI can analyze, recommend, draft or challenge, but a named human owns the decision.
- Reserve. The decision remains human because it involves values, precedent, people, significant consequences or judgment the organization has deliberately chosen not to delegate.
The audit strategy should follow that decision architecture, not substitute for it.
The NAIC bulletin similarly calls for governance structures that establish scope of authority, chains of command, decisional hierarchies, independence of decision-makers and lines of defense, as well as monitoring, auditing, escalation and reporting protocols.
The question isn't simply whether a human appeared somewhere in the process. It is who had authority to decide.
Audit the control system, not just the outputs
Once decision rights are explicit, sampling becomes much more useful. But an effective audit should test more than whether an individual output was “right.”
An insurer should be able to answer:
- What type of decision was the AI supporting?
- Who owned the decision?
- Which model and version produced the recommendation?
- What information was available to the system and the human reviewer?
- How often did humans override the AI?
- Were errors concentrated around particular products, populations or circumstances?
- Did complaint patterns change?
- Did performance change following a model, data or workflow change?
- What happened when performance moved outside acceptable boundaries?
The NAIC bulletin says regulators examining an insurer's AI use may request inventories and descriptions of AI systems and predictive models, information about data provenance and lineage, measurements and thresholds used in oversight, and documentation of validation, testing and auditing, including evaluation of model drift.
NIST's AI Risk Management Framework Playbook points in a similar direction. Its monitoring guidance recommends documenting the degree of human oversight, maintaining statistics on human overrides, tracking reported errors and complaints, recording adjudication activity, and documenting exceptions and escalation decisions.
That changes auditing from spot-checking answers into testing a control system.
Cadence should follow risk
There is another problem with traditional audit models: cadence is often calendar-driven. Teams review a fixed percentage every week or conduct a larger review every quarter. AI systems do not necessarily change on that schedule.
NIST's current work on deployed AI identifies questions such as what the right monitoring cadence is, whether monitoring should be risk-based, and how automated and human-validated monitoring should interact as important unresolved issues.
That means there isn't a regulatory magic number. There shouldn't be.
A more defensible approach is to establish a baseline review cadence appropriate to the risk and then define conditions that automatically increase scrutiny:
- A material model change.
- A meaningful change in underlying data.
- A spike in human overrides.
- Unexpected differences across customer populations.
- Complaints or adverse outcomes.
- A new product, jurisdiction or use case.
- Evidence of model drift.
- Performance outside established tolerances.
Human oversight should expand when uncertainty or potential harm expands.
The reverse matters, too. If an organization reduces review because an AI-supported process has demonstrated reliable performance, it should be able to show the evidence that justified that decision. Trust should be earned by the task, not granted permanently to the technology.
That is the Calibrate portion of the A.R.C. Protocol I use in Decisive AI: continuously measure where AI performs well, expand delegation where performance earns it, and reclaim decision authority when it does not.
What will an examiner actually ask?
We should be careful about predicting the exact questions of a future examination. Regulatory procedures vary by jurisdiction and circumstance. But we do not have to guess about the kinds of evidence regulators are preparing to examine.
The NAIC Model Bulletin says insurers should expect inquiries into their AI governance framework, risk management and internal controls. It identifies documentation regulators may request concerning AI-program implementation, monitoring and audit activities, model inventories, data practices, measurements and thresholds, testing, validation, auditing and model drift.
As of 2026, the NAIC is also developing an AI Systems Evaluation Tool to help regulators gather information in market-conduct, financial-analysis and financial-examination contexts. The NAIC reports that 12 states were piloting the tool as of March 2026, with adoption anticipated at the 2026 Fall National Meeting. The Market Conduct Examination Guidelines Working Group also has an explicit charge to develop examiner guidance for oversight of regulated entities' use of consumer data and models involving algorithms and AI.
So imagine an examiner asks a deceptively simple question: How do you know this is working?
“We have humans review 10%” is unlikely to tell the whole story.
A stronger answer is: We know which decisions AI may make. We know which decisions require human judgment. We know who owns those decisions. We know what we measure. We know where the system fails. We know when humans override it. We know what conditions increase scrutiny. We know what thresholds require escalation. And we can produce evidence showing what we did when those thresholds were crossed.
That is the difference between having humans somewhere in the loop and having an actual governance system.
Build the evidence before you need it
Don't wait for an examination request to reconstruct this story. For every consequential AI use case, an insurer should be able to produce a coherent evidence package showing:
- Decision authority: what AI can decide, what it can recommend, and what remains reserved for humans.
- Accountability: the named business and technical owners and the governance body responsible for oversight.
- Risk classification: why the use case receives the level of oversight it does.
- Performance: the metrics, thresholds and tolerances used to determine whether the system remains trustworthy.
- Human interaction: overrides, exceptions, escalations and relevant adjudications.
- Consumer signals: complaints, adverse outcomes and other evidence that may reveal problems aggregate performance metrics miss.
- Change history: material changes to models, data, workflows and third-party components.
- Audit history: what was sampled, why it was sampled, what was found and what changed as a result.
- Escalation triggers: the conditions under which additional human review, remediation or suspension becomes necessary.
That list isn't intended as a substitute for a company's legal, compliance or regulatory obligations. It is an operating model for making those obligations demonstrable.
Documentation shouldn't be created because an examiner might eventually ask for it. It should exist because the organization itself should already be asking those questions.
The Hidden Decision Debt
There is a final risk in weak AI auditing that may be harder to see.
The decisions appear finished. The claim moved. The underwriting recommendation was accepted. The transaction completed. The dashboard stayed green.
However, if nobody can explain who really owned the decision, why the organization trusted the system, whether the audit cadence was appropriate, or what would cause that trust to be withdrawn, the organization has not eliminated the decision. It has delegated it by accident.
That is Decision Debt: the accumulated cost of decisions that were deferred, degraded or allowed to migrate away from clear human ownership.
AI can reduce that debt. It can also compound it at machine speed.
The difference will not be whether organizations put humans in the loop. It will be whether they deliberately design which decisions remain human, which can be delegated, what evidence earns that delegation, and what evidence takes it away.
Sources / Regulatory References
[1] NAIC, Model Bulletin on the Use of Artificial Intelligence Systems by Insurers (adopted December 2023). https://content.naic.org/sites/default/files/inline-files/2023-12-4%20Model%20Bulletin_Adopted_0.pdf
[2] NIST, Challenges in Monitoring Deployed AI Systems (2026). https://www.nist.gov/publications/challenges-monitoring-deployed-ai-systems-center-ai-standards-and-innovation
[3] NIST AI Risk Management Framework Playbook — Measure. https://airc.nist.gov/airmf-resources/playbook/measure/
[4] NAIC, Artificial Intelligence — current regulatory work and AI Systems Evaluation Tool. https://content.naic.org/insurance-topics/artificial-intelligence
[5] NAIC, Market Conduct Examination Guidelines Working Group. https://content.naic.org/committees/d/market-conduct-examination-guidelines-wg
