The NAIC's AI Risk Evaluation Supplement is the questionnaire regulators will draw on to examine how insurers govern their AI. It asks whether a model's outputs were tested for accuracy, but only for the models a carrier itself rated high risk. When a customer-facing model gets a coverage question wrong, the policyholder acts on the wrong answer.
That is the exposure a high inherent risk rating exists to capture. Rate one low instead, and the inquiry ends there. The model keeps answering coverage questions, the answers keep reaching policyholders, and nothing requires the carrier to check whether they were right. Few carriers check on their own.
What a low rating leaves unchecked surfaces later, in a claim file rather than an examination. Misrepresenting policy provisions is a listed unfair trade practice in nearly every state, and a dated coverage statement the policy contradicts is what a misrepresentation claim is built from. The customer who heard "you're covered," paid for the repairs, and then had the claim denied is the one who brings the complaint.
When that complaint arrives, a court will ask how often the model's answers were checked and what the checks found. Errors and omissions underwriters, who now ask about AI at renewal, want the same answers. A carrier that rated its model low will have no file to produce. There is no record of what the model said and no comparison against the forms in effect at the time.
What it will have is a self-attested risk assessment concluding that the model was low risk. Plaintiffs' counsel look for exactly that pairing: a carrier that never checked, and a document explaining why it did not have to.
How the rating works
The rating is set early, in an inventory. Each carrier lists its AI models, describes what each one does, and assigns each an inherent risk level. Only the models rated high move on, and there the questions turn specific: were the outputs tested for accuracy, how is performance monitored on a continuing basis, and how is the model reviewed against unfair trade practice and claims settlement laws?
Models rated moderate or low do not reach those questions, and a regulator reading the inventory has discretion to stop at the rating and request nothing further about the model. That is sensible design, since regulators cannot examine everything. It also makes one self-assessed rating the single point of failure for whether a model's accuracy is ever examined, and the carrier is the one holding the pen.
How a customer-facing model goes unexamined
Two judgments decide whether a customer-facing model is ever examined. Does its output reach a consumer, and how much risk does it carry before controls? A carrier can get either one wrong without meaning to.
Destination settles the first question, not who speaks the words. The supplement's background section covers decisions and actions “made or supported by” AI, and a coverage answer that reaches a policyholder is one of them. The pitfall is the supplement's support category, defined as a system that provides information without suggesting a decision or action, which is not counted as having direct consumer impact. A model that drafts a coverage answer is not just providing information. It is supplying the decision the representative delivers, and the support definition excludes exactly that. But because a representative delivers the answer, the model looks on paper like one that only informs an employee, and that resemblance is what puts it in the wrong box. Filing it as support keeps the model out of the count regulators use to decide what to examine, and out of any question about whether its answers were right. Those answers still go out to policyholders, unchecked.
A model that is counted can still be rated low by mistake, and the same representative is usually why. Most carriers put two guardrails around a model's answers. The representative is expected to check it before it goes out, and the platform runs its own internal check against policy language. Both are controls, and because they are built into the process, crediting them when rating the model is an easy mistake to make. The supplement asks for inherent risk, the risk the model carries before any control is applied, and the word "inherent" is easy to read past. Under the supplement's own terms, the representative and the platform check belong in the residual rating, after controls. A carrier that credits them in the inherent rating has rated its model lower than the supplement intends, and taken it out of the testing questions in the process. Both judgments are already recorded, one line per model, in the carrier's own inventory.
Why the cheaper rating costs more
Rating a model low is the cheaper answer. It closes a line on the inventory, raises no question the carrier has to answer, and creates no obligation to find anything out. Rating it high creates recurring work: finding out whether the model's coverage answers were right and continuing to do so while the model runs. The pressure runs one direction, and nothing in the supplement pushes back. The saving lasts a quarter. The gap it leaves is permanent. That work cannot be backfilled later, because the evidence exists only while the conversations are happening, and controls do not close the gap. A representative who catches a wrong answer fixes that conversation and leaves no record of it, so even a carrier with an excellent review process cannot separate an isolated error from a model-wide one. It also cannot see drift. Models do not hold still after launch. They get updated, the policy forms they draw on are amended, and the questions customers ask shift. A test run before deployment describes the model that was deployed, not the one answering calls this quarter. What a low rating costs is a year of conversations nobody measured, and nobody can go back and measure.
I cannot tell you how often AI models give wrong coverage answers, and neither can anyone else, because almost nobody is measuring. The draft is open for comment through Sept. 29, with a vote on a later version expected in November. The questions are being written now, and the evidence they ask for can only be gathered before they are asked. Rating a model low does not change what it told a policyholder. It only changes whether anyone was looking.
