AI's Conundrum for Underwriting

Underwriting AI systems are designed to assist rather than replace human judgment, but this efficiency trade-off may eliminate how juniors become experts.

AI Underwriting

I sell software that encodes expert decisions, so hold that against everything below.

We ran an opinion-drafting model on real renewal files before connecting it to anything. It writes the underwriting opinion the way the senior underwriter does, off claims history and portfolio context. This is health cover outside the US, about 2,500 renewals, where per-condition exclusions still get written into the policy, so the medical exceptions are the job.

Two things came back in the same run. The AI overgeneralized, taking an exception written for one condition and applying it to a whole disease category. It also caught diagnoses the human-written opinions had missed.

Everyone wants to talk about the catch. Both findings landed in the same place, though. A draft, on a human's desk, under somebody else's signature. Nothing we built moved an inch of authority, and that was the design.

Everything ships below the authority line

A great deal of personal lines business quotes and binds without an underwriter looking at it, under filed rates and rules, with authority in an approved rule set instead of a person. That predates all of this by decades. So the real line runs between decisions you can specify in advance and decisions you can't, and what you can't write down stays with the person holding the pen.

The pattern almost everyone has converged on, including us, follows from that. Keep the system decision-negative. Triage, enrich, summarize, score referrals, draft the opinion. The output lands under an existing authority threshold, a human signs, and nobody reopens a governance document. That's why this version gets built, and the ambitious one doesn't. It's an engineering choice made to dodge a governance cost, and we should name it that way instead of calling it a safety feature.

It fails in one specific way. Our reviewers caught the overgeneralization because they could see which prior opinions the draft had reasoned from. Take that away and approvals start to look like review. A reviewer who can't reconstruct why the system said what it said signs because the queue is long, one of the best-replicated results in the human factors work on decision aids.

The line could move

The usual defense of that design is that the alternative is impossible, because the veteran's reasoning lives only in the veteran. I believed that for years. Mostly it isn't true.

Texas requires personal auto, residential property, and workers' compensation insurers to file their underwriting guidelines with the state within 10 days of use, and defines a guideline to include a practice whether written, oral, or electronic. The auto and property filings are public records.

But a filed guideline records the rule, not the reasoning. Nothing in a filing tells you when to depart from it, and the departures are the job. I opened by saying the medical exceptions are where the work sits. So codification retires a smaller claim than it looks. The policy is written down. The judgment isn't.

Noise cuts the other way

There's a finding anyone selling expertise capture should have to answer, and it's a decade old.

In 2015 Daniel Kahneman ran identical cases past 48 underwriters at one large carrier. Executives guessed two underwriters would differ by about 10%. The median difference between any two was 55%. One priced a risk at $9,500 where another said $16,700.

I used to read that as damaging. It runs the other way. Goldberg showed in 1970 that a model built from a judge's own past calls often beats the judge, because the model never drifts. It outperformed 79% of the clinicians it was built from. Noise is the argument for encoding, not against it. What it kills is the premise that there's one veteran in there worth copying whole.

Our own two results are that finding in miniature. The model won on consistency and lost on the rare structured exception.

Reading decisions out of files instead of asking about them is policy capturing, and it predates everyone in this argument. It shows spread that one interview at a time can't. It also flatters the file, which holds the justification rather than the reasoning.

What none of it fixes is the label. The label you have is what the underwriter decided; the label you want is whether the account ran profitably. Seasoning fixes the lag, not the hole, because declined accounts never produce a loss ratio and the only book you can test on is the one he agreed to write. Test on it anyway. Almost nobody in my category does, and I'm not the exception. Everything I've reported here is an agreement measure or a stopwatch, not a loss ratio.

Where the efficiency goes

Suppose you get all of that decision right. The gain still may not land where the business case said.

Carl Van has been in claims since 1980 and runs International Insurance Institute, which sells adjuster training, so read him with that in mind. He was talking about adjusters. The mechanism carries because it doesn't live in the job title. It lives with the executive who signed for the system. By email in August:

"Some companies themselves drive the empathy and humanity out of the adjuster position without realizing it... usually to justify the expense of an improved system, someone invariably promised higher production in some executive meeting.

"Therefore, if you improve the system, the adjuster WILL NOT spend more time being empathetic or human. They will spend the extra time trying to accomplish the higher workload standards imposed by the company."

The capitals are his. Van says the characteristics in his book can be trained, and his firm has trained them for 28 years, but they "won't happen just because the adjusters have more time on their hands." The version that works funds the training next to the system. That line item is the first one cut, because the business case is where somebody already promised the production number.

The part I can't resolve

Our renewal team went from eight people to three. Before anyone else does the arithmetic, I'll do it. Those 2,500 renewals at six minutes each come to roughly 250 hours a year, and taking the review to one minute saves about 208 of those hours. That's a tenth of a person. The rest came from the definition and follow-up work around the review, not from the review. The headline number and the stopwatch number are not the same number, and people in my line of work print both and hope you won't multiply.

What went away was routine renewal review. That's also where a junior underwriter turns into a senior one. You sit with hundreds of ordinary files until an unusual one announces itself.

The obvious fix is to route a share of straight-through-eligible files to a junior anyway and book it as training, the way a teaching hospital protects cases. Teaching hospitals tried it. Matt Beane spent two years inside robotic surgery programs and found residents getting 10 to 20 times less hands-on practice, with the ones who stayed competent managing it by breaking protocol. The fix already failed where I borrowed the metaphor.

So the trade is real and I don't have a way around it. I haven't met anyone who does.

Read More