Download

What an AI Insurance Pilot Can't Tell You

AI pilots prove capability, but production adoption hinges on workflow integration, data relevance and user trust at scale.

Insurance

An AI pilot can answer an important question: Does the capability create value? What it cannot fully answer is whether a broader group of employees will use that capability consistently while doing their everyday work.

That distinction has become increasingly important in our work with AI insurance wholesalers selling to financial professionals such as agents or advisors. A pilot can demonstrate that AI will bring together CRM history, previous product discussions, and outstanding commitments to produce a useful brief before an agent or advisor meeting. Yet production adoption takes place under different conditions. Pilot users are selected, attentive, and supported. Production users are not evaluating the AI; they are preparing for meetings, working with advisors or agents, and moving through a busy day.

For that reason, a favorable response from pilot users is encouraging, but it is not a complete test of adoption. In our experience, three areas deserve particular attention: how users encounter the AI, how the product identifies what matters within production data, and how users verify information without assistance from the pilot team.

1. Pilot Users Seek Out the AI; Production Users Need It Within Their Work

First, pilot users know they are participating in an evaluation. They expect to spend time with the product, initiate the experience, and pay close attention to the result. If generating a meeting brief requires a few deliberate steps, they will generally take them.

Production users operate differently. A wholesaler moving between advisor or agent meetings, email, follow-up, and CRM activity is unlikely to think first about using an AI capability. The immediate objective is preparing for the next conversation—not testing a product.

As a result, adoption depends partly on whether the intelligence appears at the right moment. For meeting preparation, the calendar can provide that moment. An upcoming meeting already establishes when preparation is needed and may help identify the relevant financial professional. The brief can then be delivered in connection with the scheduled work rather than depending entirely on the wholesaler remembering to request it.

Similarly, the location of the output matters. If the wholesaler already works from the calendar and Salesforce, requiring a separate destination introduces another step between the user and the value. Presenting the brief within the existing workflow makes the AI easier to use without asking employees to reorganize how they work.

The larger point is not that users resist new technology. It is that pilot participation creates attention that will not exist at the same level in production. Broader adoption is more likely when the product does not depend on preserving that pilot-level attention.

2. Controlled Data Proves the Capability; Production Data Tests Relevance

Second, a pilot usually begins with a defined set of records. This is useful because it allows the team to determine whether the AI can retrieve, organize, and summarize the intended information.

Production data is less controlled. A financial professional's CRM history may contain years of calls, meetings, emails, marketing activity, and service interactions. Some entries are meaningful to the next conversation. Others are routine, incomplete, or no longer relevant.

Consequently, accuracy alone does not guarantee a useful brief. AI can summarize every available record correctly and still give too much attention to information that does not matter now. The production challenge is not simply to retrieve more data; it is to prioritize the signals most relevant to the work being performed.

For example, an unfinished commitment from the previous meeting might deserve more attention than numerous routine CRM activities. A recent product discussion may be more useful than an older description of the overall relationship. Likewise, recent annuity illustration software activity could indicate that an advisor is actively evaluating a particular product or client strategy.

However, the meeting brief may not need every value and disclosure contained in that illustration. The useful distribution context might be which product was illustrated, when the activity occurred in the annuity illustration software, and whether it led to additional engagement.

Therefore, connecting AI to CRM, product, and customer data is only the beginning. The product also needs enough business context to distinguish between information that is available and information that deserves the wholesaler's attention. A pilot can prove that the AI can process the records; broader use reveals whether it can consistently surface what matters.

3. Pilot Oversight Supports Trust; Production Requires Independent Verification

Third, AI output receives unusual scrutiny during a pilot. Participants know that the capability is being evaluated, and the team running the pilot is usually available to investigate questions. The underlying data may also be familiar to the people reviewing the results.

That level of support does not scale into everyday use. In production, wholesalers need to assess important information without asking the product team to explain how the AI reached a conclusion.

Consider a brief stating that an advisor previously showed interest in a particular annuity strategy. The wholesaler should be able to determine what supports that statement. Was the interest documented in a meeting note? Was an illustration generated? Did the advisor request product information? Or did the AI infer the interest from several activities?

This distinction affects how confidently the wholesaler should use the information in an advisor or agent conversation. It becomes even more consequential when the information moves downstream into a follow-up email, a CRM update or the context prepared for a future meeting.

For this reason, source attribution serves a practical purpose beyond governance. Connecting important statements to their underlying records allows users to verify the context independently. Separating record-based facts from AI-generated suggestions also helps the wholesaler apply professional judgment rather than treating every statement as equally certain.

In a pilot, confidence may be reinforced by the team surrounding the test. In production, the product itself must provide enough transparency for users to decide how the information should be used.

What the Pilot Proves—and What It Does Not

Taken together, these three differences do not diminish the value of a pilot. They clarify what the pilot is designed to establish.

A pilot can show that AI will perform the intended task and that users recognize value in the output. It can validate the central use case, expose data considerations, and provide meaningful feedback before broader deployment.

At the same time, a pilot cannot fully reproduce the conditions under which a larger group of employees will use the capability repeatedly and without special attention. It cannot by itself establish whether the AI will appear naturally within the work, remain relevant across a full production data set, and earn trust without the pilot team nearby.

For us, the central lesson has been that adoption is not simply a favorable reaction to an AI-generated result. It depends on how naturally the capability enters the workflow, how effectively it directs attention, and whether users can act on its information with confidence.

A pilot answers whether AI can create value. Production adoption depends on whether employees can receive that value while doing the work they already came to do.


Jay Singh

Profile picture for user Jaysingh

Jay Singh

Jay Singh is a co-founder and head of client solutions at Hedgeness, which provides AI-powered sales and marketing software for insurance carriers and asset managers. 

He has more than two decades of experience across financial services and technology companies. He is a frequent speaker at industry conferences on AI, financial services technology, and distribution.

AI Apocalypse? Don't Get Distracted

We're suddenly having a debate about whether AI is about to kill us all, and it obscures some pressing issues.

Image
Insurance

While we've suddenly landed in the middle of a debate about whether AI may be about to obliterate the human race, I hark back to a profile I did for the Wall Street Journal about a brilliant AI and robotics researcher from Carnegie Mellon named Hans Moravec.

The focus was his provocative idea that humans would be able to download their brains — their entire consciousness, their full personality, an exact replica of them — into computers, which could then teleport to any spot in the universe or spawn an infinite number of what Moravec called "mind children." 

The memorable headline was:

Good News: You

Can Live Forever;

Bad News: No Sex

I asked Moravec how long it would take for his vision to be realized. "Oh, a long time," he said. "Maybe 25 years."

That was 35 years ago.

So I'm not going to worry much for years about all the talk of impending doom. Timelines on sci-fi-like change tend to be way, way off. But, under the radar, there are plenty of AI issues that should be major concerns right now, including for insurers.

Let's have a look.

I'll start with Bill Gates's recent manifesto, which, among other potential dangers from AI, called out the prospect that AI will supercharge the work of malign actors, perhaps leading to bioterrorism, massive cyberattacks, and more. While we can discuss the potential long-term threats to humanity from AI, these are the kinds of threats I think we need to focus on today. These threats are already being pursued, whether by individuals looking to extort massive amounts of money or by nations looking for weapons in an increasingly belligerent world, and AI clearly provides exponentially more computing capability.

The MIT Technology Review goes into detail about how AI might produce a devastating bioterrorism attack: "A bioweapon might be a highly lethal virus that targets people according to their genes. It could be a fungus that wipes out a crop and causes food insecurity. Perhaps it would be a tasteless, odorless toxin that could be slipped into a region’s water supply, undetected....

"Today, AI bots can answer questions on topics spanning all realms of science. Anyone can use large language models trained on the knowledge and experience of 'almost every scientist who ever lived on this planet,' says Dunja Sabra, a biosecurity researcher at the University of Hamburg in Germany. Those models can provide instructions and video training on how to conduct experiments.

"Combine that with advances in biotech that have made gene editing and synthetic biology tools much more accessible (the “DIY biology” movement has already enabled many people to set up labs at home), and you’ve got a potentially very dangerous situation."

Wired, meanwhile, warns about all the vulnerabilities in software that AI bots are finding, and it's not hard to imagine how those weaknesses could be exploited. In July, hackers, thought to be based in Iran, disrupted 30 municipal water systems in Minnesota, and you can be sure Iran will ramp up attacks as fast as it can. North Korea, China, Russia, and other countries could stage similar, small attacks or could even try to shut down electric grids and stall commerce by using weapons of not-quite war.

Cyber attacks could easily lead to massive business interruptions of the sort insurers routinely cover and could increase geopolitical risks of every flavor.

Businesses and governments understand that the bots are making them vulnerable and are working as fast as they can to plug the holes, but they won't find all the holes, at least not right away. And you be sure that some hacker cartel or foreign government is storing up what are known as "zero day" vulnerabilities that can be unleashed on unsuspecting businesses and societies. 

There is also massive potential for operator error as AI is deployed more broadly. For instance, the plan to use AI in air traffic control, just now going live for DC-area airports, strikes me as an accident waiting to happen. I hope everything goes smoothly, but the potential for trouble is so great that insurers and everyone else should be wary. Air traffic control hasn't exactly acquitted itself well lately, and AI tends to amplify flaws by making everything happen faster. 

That's where I think the focus should be over at least the next couple of years, both for insurers and for society writ large: on the potential bio, crypto and other deliberate, organized attacks that AI makes possible from bad actors who would profit from those attacks, as well as on the potential for catastrophe as AI gets more involved in mission-critical efforts. 

Yes, there is always the possibility that the search for superintelligence could create an AI that will go rogue and do indescribable damage to the whole human race, for no apparent reason, so government officials should be erecting guardrails.

But business is already circumscribing what AI can do. Insurers certainly are. They've realized that justifying a decision with "the AI says so" won't fly, so they're requiring that every decision be explainable and are greatly limiting what actions an AI can take without explicit human permission.

The whole superintelligence debate has so many dimensions even beyond the technical ones — there are political elements, issues related to business models, massive public relations concerns, etc. Here, for instance, is a column in the WSJ that argues the whole apocalypse debate is an attempt by AI's leading developers to duck responsibility. 

The issue reminds me of Winston Churchill's description of the Soviet Union after it allied with Nazi Germany in 1939: "a riddle wrapped in a mystery inside an enigma." And I'm supposed to understand the flow of technology revolutions, having followed them for decades.

Fortunately, I think we can wait to puzzle out all the implications of this AI revolution — as long as we don't take our eye off the ball on the imminent threats it creates.

Cheers,

Paul

 

 

A Water Loss Prevention Program That Truly Works

LeakBot's 65,000 years of US underwriting data proves actuarially relevant loss mitigation: mature programs cut non-weather mains-related water claims ~60%, and the $5/month all-in model delivers positive ROI across market segments.

Insurance

The Problem

Non-flood damage from water costs homeowners billions a year, and insurers pay $15 billion in claims, just in the U.S. One in 60 U.S. homes suffers water damage each year, and the average claim is nearly $14,000—which doesn’t even include the deductible the policyholder pays, or the huge hassle they go through. Typical solutions involve expensive shutoff valves or a host of small sensors placed throughout a home, but providers have struggled to make a compelling case that their programs deliver an ROI.

The Solution

LeakBot takes a different approach, one that begins before a leak occurs and stretches to a solution. The approach starts with a LeakBot sensor that policyholders clip to the main water pipe in the house. The sensor detects any leaks anywhere in the house that exceed a teaspoon per minute. LeakBot’s app walks a policyholder through steps to find leaks, then, if necessary, sends one of its plumbers to visit the house to find any remaining ones, at no charge. The LeakBot plumber documents the repair, giving carriers visual validation of every claim saved.

The approach finds even tiny, hidden leaks weeks or months before they can cause significant damage. In 2025, LeakBot completed 7,000 leak repairs, 1,470 of which revealed water damage that had already begun.

 

More than two dozen carriers provide the LeakBot solution for free, paying LeakBot $5 a month per house, all in. LeakBot reduces the number of water loss claims by 60%, generating a significant ROI for its carrier partners.

LeakBot covers the full loop, detection through documented repair, which makes it the only true end-to-end platform in this space.

The Documentation

LeakBot has produced a white paper that clears the actuarial bar for detail and reliability. The paper draws on hundreds of thousands of device-exposure years. It uses cohort-level (not anecdotal) evidence. It relies on auditable repair records — rather than programs still running on pilot-stage promises.

 

 

Sponsored by Leakbot


Leakbot

Profile picture for user Leakbot

Leakbot

LeakBot is the only end-to-end IoT solution protecting homes from water damage—one that begins with leak detection and can end with a free repair. Backed by more than 10 years in business and 29 patents, a single self-installed device clips onto the home's main water supply line, monitoring water usage and detecting micro-leaks as small as one teaspoon per minute. When a leak is detected, the homeowner can book an appointment for LeakBot's trained employee plumbers to visit, locate, and repair it using specialty equipment — at no additional cost to carrier or homeowner. Homeowners consistently recognize LeakBot's value, reflected in a Net Promoter Score of 82/100 and Customer Satisfaction Rating of 4.9/5. That impact is especially powerful when a hidden micro-leak is found and fixed before it becomes a claim—or a costly plumber bill. That's #PredictAndPrevent in action.

Please connect with us: 

 

Insurers Should Rethink Outsourcing of Data Work

Insurers outsourced data integration, but managed services often trade speed and control for convenience that quietly becomes dependency.

Insurance

When insurers first adopted fully managed data integration, the pitch was simple: hand over the technical complexity, and focus on the business. For years, that trade made sense. Integration work was specialized and hard to staff for, and outsourcing looked like a clean efficiency gain.

But convenience has a way of turning into dependency over time, and dependency tends to stay invisible right up until the moment a business actually needs to move fast.

Take something as ordinary as onboarding a new employer group, or updating a data flow to reflect a new product rule. In a fully managed model, this rarely stays routine for long. It becomes a ticket, then a statement of work, then a wait measured in weeks or months. One major group insurance carrier recently found that executing a single employer group statement of work took an average of 56 days before onboarding of that employer group even began. Not because the technical change itself was complex, but because the business rules governing it lived in a vendor's queue rather than in the carrier's own hands.

That's the cost worth examining closely: not what shows up on the invoice, but what the invoice doesn't show – how much of a company's own operational agility has quietly moved along with the technical work it outsourced.

When Speed Belongs to Someone Else

Every managed integration contract answers a question most companies never ask directly: who gets to move at the speed of the business, and who has to move at the speed of a vendor's backlog?

In a fully managed model, the logic that actually runs the business, the mappings, the transformation rules, the exceptions built up over years of institutional learning, typically lives inside the vendor's systems, in the vendor's tools, understood mainly by the vendor's staff. That works fine when nothing needs to change quickly. It stops working the moment something does: a new regulatory requirement, a partner who needs to be live before a deadline unrelated to anyone's ticket queue.

"Fully managed" often ends up meaning "fully dependent." And dependency isn't really a technology problem. It's a business one.

Knowledge That Lives in Someone Else's Code

There's a second risk that tends to surface only when a relationship changes, whether a vendor gets acquired, a contract comes up for renewal, or a company decides it's time to modernize.

Over years of operation, a great deal of institutional knowledge accumulates inside integration logic: how a specific partner's data should be read, which exceptions matter, or why a rule was written the way it was in the first place. When that knowledge lives entirely within a vendor's proprietary configuration, it isn't really the company's knowledge anymore; it's on loan.

This isn't an abstract concern about vendor lock-in. It's a concrete question: if a company needed to walk away from a managed integration relationship tomorrow, could it actually reconstruct the rules governing its own data? Or would that knowledge have to be rebuilt from scratch, at real cost, because it was never fully owned to begin with?

A Simple Test

Here's a question worth asking honestly, regardless of how an organization currently handles integration: if a business stakeholder needed to onboard a partner or change a rule this week, could the internal team actually do it, or would it have to file a ticket and wait?

For many organizations, the honest answer is the second one. That isn't automatically a crisis. But it's a signal worth sitting with because it says something about whether integration is functioning as a strategic capability or as an outsourced dependency that still runs smoothly enough not to notice.

Before the Next Renewal

A few questions tend to surface the real state of operational control, and they're worth asking before signing the next managed-services renewal or kicking off a modernization effort. Does the internal team have real visibility into the business rules governing its own data, or only a vendor's summary of them? If the relationship ended tomorrow, what would it take to rebuild that logic independently? Is the current turnaround time for a routine change driven by genuine complexity, or by contractual process? And who within the organization actually understands, end-to-end, how its data moves today?

None of these questions have one right answer. Some organizations will look at their scale and resources and conclude that a managed model still makes sense, with trade-offs included. Others will realize that the convenience they signed up for years ago has quietly become a limit on how fast they can move today.

Either way, it's a decision worth making deliberately, as a question of operational control rather than simply cost or convenience. The companies that move fastest over the next few years probably won't be the ones with the most sophisticated integration technology. They'll be the ones who know, with certainty, whether they actually own the rules that run their business, or whether someone else is running the show.

The Next Frontier in Healthcare Risk

Digital musculoskeletal care has proven its market value, but the next frontier is using AI to measure functional decline before it becomes disability.

Insurance

Over the past several years, digital musculoskeletal (MSK) care has evolved from an emerging concept into a validated healthcare category. Companies such as Sword Health and Hinge Health have demonstrated that technology-enabled MSK solutions can attract significant investment, expand access to care, and create new models for managing one of healthcare's largest cost areas. Sword Health has reached multibillion-dollar private valuations, while Hinge Health successfully shifted from a highly valued private company to the public markets, demonstrating strong investor confidence in digital MSK care models.

Their success has proven something important: employers, healthcare systems, and insurers are ready for technology-enabled approaches to musculoskeletal health.

The need for innovation has never been greater. Musculoskeletal disorders represent one of the largest drivers of pain, disability, healthcare usage, and lost productivity. According to research published in the Journal of Medical Internet Research, annual healthcare spending associated with MSK conditions in the United States is estimated to be approximately $300 billion. The impact extends beyond direct treatment costs, as MSK conditions are often associated with other chronic health challenges, including obesity, diabetes, cardiovascular disease, arthritis, frailty, and falls. The true healthcare and societal burden is therefore significantly broader than the cost of treatment alone.

But the next question is even more fundamental:

How do we move from delivering care more efficiently to understanding changes in human function earlier?

The next evolution of healthcare risk management may not simply be better treatment after a condition develops. It may be the ability to objectively measure functional changes before they progress into disability.

This shift is increasingly reflected in healthcare research priorities, including efforts focused on aging, musculoskeletal health, digital health technologies, and artificial intelligence. The National Institutes of Health and the National Institute on Aging are emphasizing the importance of technologies capable of objectively measuring functional changes, understanding age-related decline, and supporting approaches that preserve mobility and independence.

Maintaining mobility and physical function is a fundamental component of healthy aging. Changes in strength, muscle activation, movement quality, balance, and physical performance can contribute to increased risk of falls, frailty, disability, loss of independence, and reduced quality of life.

Yet functional decline is often gradual. A person does not suddenly become frail. Changes may occur over months or years before they become apparent through traditional healthcare encounters.

Today, functional status is primarily evaluated through episodic clinical assessments, performance-based measures, and subjective reporting. These approaches provide valuable information, but they offer only limited snapshots of an individual's health and may not capture subtle changes in physiological and biomechanical function over time.

Healthcare has become increasingly sophisticated at measuring disease. The next opportunity is developing the ability to measure changes in human function. Wearable technology has already transformed personal health monitoring. Individuals can track steps, heart rate, sleep patterns, and activity levels. However, the next generation of digital health technology must move beyond activity tracking toward functional intelligence.

The question is no longer only:

"How much did someone move?"

The more important question is:

"How did someone move?"

Mobility is influenced by complex interactions among muscle activation, coordination, biomechanics, range of motion, and movement quality. Technologies that integrate physiological sensing, biomechanical measurement, and artificial intelligence have the potential to provide deeper insight into individualized functional patterns.

An AI-enabled platform integrating technologies such as electromyography (EMG), inertial measurement units (IMUs), and advanced movement analytics could establish a foundation for objective digital measures of mobility, functional aging, and early changes that may affect independence and quality of life.

The challenge in healthcare is no longer simply collecting data. The challenge is transforming increasingly complex streams of information into meaningful insights that can support better decisions.

Artificial intelligence offers the opportunity to analyze large amounts of physiological and movement information and identify individualized patterns that may not be visible through traditional assessments. The goal is not to replace clinicians, but to provide clinicians, insurers, employers, and individuals with better information to support earlier and more personalized decisions.

AI-enabled healthcare has the potential to move the system away from simply reacting to decline and toward identifying meaningful changes earlier, when there may still be an opportunity to preserve function.

For insurers, this represents a potential transformation in risk management. The traditional healthcare model has largely followed a reactive path: a condition develops, treatment begins, and a claim occurs.

The future opportunity is a more proactive approach in which functional changes are identified earlier, personalized strategies are developed, and risks associated with decline may be reduced before they progress into costly healthcare events.

Objective measures of functional health could support earlier identification of risk, more personalized care pathways, improved management of chronic conditions, support for aging populations, and value-based healthcare models focused on maintaining function rather than simply treating decline.

This shift is especially important as populations age. Aging does not automatically mean disability. Many individuals remain active, engaged, and independent when changes in mobility, strength, and function are recognized early and appropriately addressed.

Preserving function is not only a healthcare goal; it is becoming an essential component of managing healthcare costs, improving outcomes, and supporting a more sustainable healthcare system.

The future of healthcare risk management will require a fundamental shift in how we define prevention. Prevention can no longer focus only on identifying disease after it develops. It must also include understanding changes in human function before they become disability.

The most valuable healthcare technologies of the future may not simply tell us what happened after a problem occurs. They may help us understand what is changing — enabling healthcare systems, insurers, employers, and individuals to act earlier to preserve mobility, independence, and quality of life.

The future of healthcare risk management will not only be defined by how effectively we respond to disease and disability, but by how well we can identify change earlier, preserve function, and help individuals maintain healthier lives.

How to Solve the IT Backlog (and How Not to)

P&C carriers lose millions quarterly when underwriting changes languish in IT backlogs, but uncontrolled self-service creates chaos.

Insurance

In most mid-to-large property & casualty (P&C) carriers, changing a routine underwriting limit or pricing parameter requires riding the exact same deployment pipeline as core application code. The change must navigate an IT ticket, a sprint, a test cycle, a change advisory board, and a release window. By the time a single line of configuration ships, weeks or months have passed.

The result is a business operating at analog speed in a digital market. While executive leadership talks about digital transformation, front-line business teams remain tied to IT backlog ticket systems. Shortening sprint cadences or pushing for faster CI/CD pipelines fails to resolve this friction because the bottleneck is architectural, not process-driven.

The High Financial Cost of "Locked-Configuration Syndrome"

Historically, embedding underwriting guidelines, pricing limits, claims triage, and compliance checks directly into core source code or static configuration files was standard practice. That approach worked when market shifts happened over multi-year cycles, but today it locks basic operational decisions behind full software release schedules.

That friction carries a quantifiable financial penalty. For a carrier writing $2 billion in Gross Written Premium (GWP), delaying a necessary rate or eligibility adjustment on a deteriorating loss segment by a single quarter can easily affect the combined ratio by $15 million to $25 million. That is before accounting for the hidden opportunity costs of delayed product launches and missed competitive moves.

The Pitfalls of Uncontrolled Self-Service

To bypass this bottleneck, carriers often try giving business teams direct self-service capabilities. But granting autonomy without governance just trades slow delivery for production instability. When carriers attempt to bypass IT pipelines using direct database updates or rudimentary rule tools, they consistently run into three major failure modes:

  • Silent Misconfiguration: A business analyst updates a rating factor in a live environment. The change is syntactically valid, but without contextual validation, it unexpectedly alters logic in downstream claims systems. Thousands of policies are mispriced over 48 hours, requiring emergency rollbacks, data patches, and reputational damage control.
  • Audit Vacuum: During a regulatory examination, regulators request the 24-month change history for a specific underwriting rule. Because changes were made informally by multiple users across various tickets, pre-change values, timestamps, business justifications, and formal sign-offs cannot be produced. The result is regulatory penalties and corrective action plans.
  • Environmental Drift: Development teams spend six months testing a major software release in Non-Production Environments (NPE). Upon deployment, critical rules fail because business users made unrecorded configuration changes directly in production over preceding months. Testing in NPE becomes an illusion because test environments no longer reflect reality.
Configuration as a Service: A Governed Control Plane

Fixing this doesn't mean choosing between rigid IT stability and unchecked business speed. The path forward is establishing a dedicated Configuration as a Service (CaaS) layer, a purpose-built "Configuration Center" positioned between business users and core administration systems.

Functioning as a centralized control plane, this layer routes parameter updates through an automated, policy-enforced pipeline instead of relying on IT release backlogs or risky database edits. Core application platforms simply query the engine at runtime to retrieve current values and execute logic safely.

This framework rests on three core pillars:

  1. Risk-Tiered Autonomy

    Granting self-service rights based on the potential financial and operational impact of a change rather than user convenience.

  2. Integrated Governance Controls

    Built-in pre-flight simulation, automated impact analysis, mandatory sign-offs, and immutable audit trails.

  3. Continuous Environment Synchronization

    Ensuring configuration changes made in production automatically propagate back to non-production environments to eliminate environment drift.

Deciding What Deserves Self-Service: The Three Tiers

The most critical architectural decision is not choosing software tools, but classifying configuration assets by risk. Not every parameter should be editable by business teams, and items that are self-serviceable must follow controls proportional to their potential impact. 

HIGH RISK / ~10% OF CHANGES

Tier 1: Tech-Managed

Covers core data models, integration schemas, rating calculation engines, and strict regulatory policy logic. Errors here directly risk financial loss, compliance breaches, or platform instability. These assets stay firmly in IT's domain, following standard SDLC processes and formal CAB review gates.

MEDIUM-HIGH RISK / ~30% OF CHANGES

Tier 2: Governed Self-Service

Covers underwriting eligibility matrices, core rating rules, and workflow decision trees. Mistakes can affect underwriting performance, but risks remain manageable if caught before release. Business teams author and update these rules, but deployment demands pre-flight simulation and dual sign-off from both the business owner and an IT/risk lead.

LOW-MEDIUM RISK / ~60% OF CHANGES

Tier 3: Routine Self-Service

Covers bounded operational toggles, customer-facing advisory messaging, and routine threshold updates. Financial exposure is minimal, with changes mostly driving day-to-day operational efficiency. Business leads manage these end-to-end using single-approver workflows and automated fast-track promotion pipelines.

Accelerating Safety With Artificial Intelligence

Integrating targeted AI capabilities into the configuration pipeline further reduces human error while accelerating throughput:

  • Change Impact Analysis: Large Language Models (LLMs) trained on historical rule changes analyze new proposals, flagging potential downstream system impacts, edge-case scenarios, and historical outcomes before sign-off.
  • Natural Language Authoring: Analysts describe desired business logic in natural language, which AI translates into formal Decision Model and Notation (DMN) tables for validation.
  • Anomaly Detection: Machine learning models continuously evaluate configuration change streams to detect statistical anomalies in timing, parameter ranges, or change volume.
  • Automated Simulation: AI automatically generates and executes targeted test cases against proposed configuration changes in a sandbox environment, attaching results to the approval queue.
Business Outcomes and Key Metrics

Carriers implementing a governed configuration control plane transition from slow, reactive cycles to proactive market responsiveness. Key operational benchmarks demonstrate the impact:

Compressed Cycle Times: Low-risk operational changes move from multi-month IT backlogs to sub-4-hour automated deployment windows.

Incident Reduction: Production incidents stemming from configuration errors drop by 80% to 90%.

Complete Audit Compliance: Full traceability satisfies stringent market conduct standards, recording who changed what, why, and who approved it.

Resource Efficiency: Engineering hours spent managing routine configuration tickets decrease by up to 90%, freeing technical teams to focus on core platform innovation.

The tension between business agility and production stability is real, but it is not irreconcilable. By replacing one-size-fits-all IT release management with a governed, risk-tiered configuration architecture, P&C carriers can deliver the autonomy business teams require without compromising stability or regulatory compliance.

For a deeper architectural breakdown, including failure mode analyses, environment synchronization patterns, and implementation road maps, read the full white paper: Configuration Center: A Blueprint for Core System Agility.

A Way to Tackle Rising P&C Deductibles

Rising deductibles are shifting more storm damage costs to property owners, creating a gap between coverage and affordable recovery.

Insurance

Wind and hail losses are growing more frequent and more expensive, and the shape of that growth is easy to miss if you're only watching total loss dollars. The more significant shift, from where I sit, is in how much of that cost policyholders are now being asked to retain before coverage responds at all.

The pattern in our own data

That shift starts with how the underlying risk itself is behaving. Tornado damage, for example, is showing up in places it historically was not concentrated, a pattern visible in Adaptive's underwriting data over the past several years. Large hail has followed a similar trajectory. Through the first several months of 2026, our data shows events involving hail two inches or greater in diameter, which can damage even newer roofs, have occurred at roughly three times their historical rate.

As those loss patterns change, carriers are adjusting how wind/hail risk is shared with policyholders. One of the clearest results has been the growing use of percentage-based deductibles. A structure that was largely based on a flat dollar amount a decade ago is now commonly 2%, with 5% or higher increasingly standard in higher-risk markets and some coastal and catastrophe-exposed segments exceeding 10%.

The Midwest illustrates this well. Markets that historically carried flat-dollar $1,000 deductibles now commonly see 3% to 5% wind and hail deductibles instead. State Farm's minimum wind and hail deductible in Texas, for example, moved from 1% to 2% deductibles specifically in response to loss frequency in the Dallas-Fort Worth area, a documented, public example of the broader trend.

Run the math on a property insured for $1 million: a 2% deductible means $20,000 out-of-pocket before the primary policy pays a dollar. At 5%, that's $50,000. For larger commercial assets, the absolute numbers scale accordingly. Chicago's Dearborn Station, for example, carries a wind and hail deductible of more than $400,000 under its primary policy.

What retained risk actually does to behavior

The size of the deductible matters. But the more revealing question is what happens to decision-making after a loss, once a deductible is large enough to strain a property owner's finances.

I've watched this play out directly in my own building, a high-rise of more than 30 floors in downtown Chicago. Over the past three years, the property has sustained repeated wind damage to its garage door, with each incident landing just under the policy's deductible threshold. After a significant wind event in March, the HOA changed the building's operating procedure. Instead of filing a claim, it left the garage door open during the day and closed it only at night to reduce further wind exposure.

That may be the most practical financial decision, but it also creates new concerns around issues like security and access. It is a good example of something that broader loss data may not show: a property can be insured, but the deductible may still be high enough that filing a claim does not make sense.

That single building isn't an isolated case. Estimates commonly cited from FEMA put the share of small businesses that don't reopen after a disaster at around 40%. Separately, a 2026 Housecall Pro survey found 77% of homeowners are delaying or scaling back home projects due to rising costs, and 41% report having delayed a repair that ultimately cost more as a result.

Higher deductibles help carriers manage loss volatility, but they also require policyholders to retain more of the financial risk. When a business or homeowner cannot realistically cover that upfront cost, repairs may be postponed. The original damage can worsen, operations can remain disrupted, and the eventual cost of recovery may continue to grow.

For businesses, the effects can extend beyond the property itself to employees, customers, lenders, the surrounding community, and, ultimately, the insurance providers that serve them.

Where deductible buy-back fits into the picture

Wind and hail deductible buy-back coverage is designed to address this specific exposure. It is a supplemental layer that sits alongside the primary policy and reduces what a policyholder ultimately pays after an eligible loss.

Two examples show how the coverage can work across different property sizes:

A small commercial policyholder, such as a children's gymnastics studio carrying a $20,000 wind and hail deductible, could purchase coverage that reduces its ultimate out-of-pocket responsibility to $5,000. The supplemental policy could reimburse up to $15,000 of the deductible, subject to its terms and conditions, for an annual premium of approximately $450.

At the other end of the spectrum, a multiuse commercial building that sustains $1 million in storm damage and carries a $500,000 deductible under its primary policy could see that retained exposure reduced to as little as $10,000 with buy-back coverage in place.

The mechanism is the same in both cases: the coverage reduces the amount the policyholder ultimately retains after an eligible loss.

Depending on the timing and terms of the claim, the property owner may still need to fund some repair costs before reimbursement is received. Once paid, however, the deductible buy-back coverage can help restore working capital and reduce the longer-term financial and operational impact of the loss.

What this means going forward

A new gap is emerging. It is different from the traditional insured-versus-uninsured split.

This one sits between coverage that technically exists and coverage a policyholder can realistically use to recover after a loss.

As deductibles rise, more policyholders may find that the amount they are expected to retain is difficult to absorb. The consequences appear in delayed repairs, deferred maintenance and decisions like leaving a garage door open during business hours because the alternative is not financially manageable.

Closing that gap will take traditional and specialty carriers working from different angles. Deductible buy-back coverage is one of them, a way to give policyholders more control over what they retain and to treat the deductible as a variable worth managing, not a fixed cost of doing business.

It will not be the only solution. As loss patterns, property values, and carrier capacity continue to change, the market will need a range of products that distribute risk more effectively while keeping recovery financially achievable.

The goal is to make sure the risk that remains is something an organization can actually afford to act on when it materializes.

Sources: Adaptive Insurance internal underwriting data; State Farm public deductible notices (Dallas-Fort Worth); Housecall Pro 2026 homeowner survey; FEMA

What Insurers Must Get Right on AI

Converged platforms can unify underwriting, claims and policy data in real time, letting AI act on complete context rather than fragmented systems.

Insurance

There is always going to be another new platform, another integration, and, increasingly, another conversation about AI. But the value of technology isn't in simply having more of it. It's how well it connects and whether insurers get better information when they need it.

At a time when AI is redefining how the insurance industry will function for the next 30 years, a converged environment unifies underwriting, policy administration, billing, claims, and customer engagement on one operating platform, rather than a chain of systems handing data to each other. A claims signal reaches pricing while the claim is still open, informing the next risk assessment in real time. An underwriting decision immediately shapes how service and retention teams treat that customer afterward. Information moves in near real time, following what the business needs rather than the limits of the original system architecture.

The Case for Converged Platforms

Across this platform, AI becomes part of how the workflow runs. An agent reviewing a claim has policy history, the underwriting file, and prior customer interactions available at the same moment, because they live in the same context rather than three different logins. That context lets AI move from flagging an exception to resolving it, and from recommending a next step to taking it, with a human reviewing the outcome afterward. The reasoning behind each action stays visible and auditable, so the people accountable for the outcome can see exactly how the system reached it.

The practical benefits show up quickly. Teams configure new products within a shared environment instead of rebuilding them system by system, so what used to take months of coordinated releases now takes days. Renewal pricing, fraud screening, and service routing draw on the same live data rather than nightly batch exports, so all cross-referenced with each other and reflect the real-time state of the business.

The autonomous AI platform we are describing doesn't mean anyone should abandon the judgment that experienced underwriters and claims handlers bring. A converged environment pairs automation with quality gates. An organization can route a routine renewal straight through and hold a complex commercial risk for human review, using the same data for both. The goal is to make sure the people making the decision are working from the full picture, instead of whatever fragment their siloed, legacy systems offer.

Rethinking Process Design in an AI-First World

It is important that companies should think carefully about the shape of their AI models. There is a natural tendency for people to focus on human-driven, effort-based, bottom-up processes and then apply AI efficiency factors to achieve better productivity results. I believe that approach is fundamentally flawed as it quietly assumes the old process as the correct baseline, when the old process is the very thing that should be questioned.

Companies should work back from the assumption that every function is AI-automated and design that as the starting point, not the improvement point. Then justify, line by line, the human effort that is genuinely needed to be added back: the judgment calls, the client-facing decisions, the regulatory sign-offs, the things an autonomous AI agent demonstrably should not be permitted to do.

The best approach is to question the requirement itself and what value it delivers to the business; wherever possible, delete the step entirely rather than automate it, simplify the real requirements that survive, and only then automate the complete end-to-end process with the minimum human intervention possible. There is no business value in automating a step that should have been removed completely, as this is the most expensive mistake that organizations make.

The clearest sign that this transformation is working comes down to a simple test. Can a service representative see a customer's full history the moment the call connects, without transferring them or putting them on hold while three systems are checked in sequence? Can an underwriter check how similar risks have actually performed against what the rate manual predicts? Can a claims handler settle a straightforward loss the same day it's reported, because the verification, the policy check, and the payment authorization all draw from the same source of truth?

That is the standard worth building toward. For companies where coherence is increasingly a transformational priority, my advice is to build an environment where every part of the business can see what every other part already knows, and where AI has enough context to act on that knowledge instead of merely describing it.


James Hannay

Profile picture for user JamesHannay

James Hannay

James Hannay is chief revenue officer at Sapiens.

Previously, he served as chief growth officer at HCL Software and EVP and general manager at Alight Solutions. Prior to that, he held senior roles at Infor.

NAIC AI Risk Ratings Let Errors Slip

Insurers' self-assessed AI risk ratings determine regulatory scrutiny, but rating customer-facing models too low leaves coverage errors undetected until claims arrive.

Insurance

The NAIC's AI Risk Evaluation Supplement is the questionnaire regulators will draw on to examine how insurers govern their AI. It asks whether a model's outputs were tested for accuracy, but only for the models a carrier itself rated high risk. When a customer-facing model gets a coverage question wrong, the policyholder acts on the wrong answer. 

That is the exposure a high inherent risk rating exists to capture. Rate one low instead, and the inquiry ends there. The model keeps answering coverage questions, the answers keep reaching policyholders, and nothing requires the carrier to check whether they were right. Few carriers check on their own.

What a low rating leaves unchecked surfaces later, in a claim file rather than an examination. Misrepresenting policy provisions is a listed unfair trade practice in nearly every state, and a dated coverage statement the policy contradicts is what a misrepresentation claim is built from. The customer who heard "you're covered," paid for the repairs, and then had the claim denied is the one who brings the complaint.

When that complaint arrives, a court will ask how often the model's answers were checked and what the checks found. Errors and omissions underwriters, who now ask about AI at renewal, want the same answers. A carrier that rated its model low will have no file to produce. There is no record of what the model said and no comparison against the forms in effect at the time. 

What it will have is a self-attested risk assessment concluding that the model was low risk. Plaintiffs' counsel look for exactly that pairing: a carrier that never checked, and a document explaining why it did not have to.

How the rating works

The rating is set early, in an inventory. Each carrier lists its AI models, describes what each one does, and assigns each an inherent risk level. Only the models rated high move on, and there the questions turn specific: were the outputs tested for accuracy, how is performance monitored on a continuing basis, and how is the model reviewed against unfair trade practice and claims settlement laws? 

Models rated moderate or low do not reach those questions, and a regulator reading the inventory has discretion to stop at the rating and request nothing further about the model. That is sensible design, since regulators cannot examine everything. It also makes one self-assessed rating the single point of failure for whether a model's accuracy is ever examined, and the carrier is the one holding the pen.

How a customer-facing model goes unexamined

Two judgments decide whether a customer-facing model is ever examined. Does its output reach a consumer, and how much risk does it carry before controls? A carrier can get either one wrong without meaning to.

Destination settles the first question, not who speaks the words. The supplement's background section covers decisions and actions “made or supported by” AI, and a coverage answer that reaches a policyholder is one of them. The pitfall is the supplement's support category, defined as a system that provides information without suggesting a decision or action, which is not counted as having direct consumer impact. A model that drafts a coverage answer is not just providing information. It is supplying the decision the representative delivers, and the support definition excludes exactly that. But because a representative delivers the answer, the model looks on paper like one that only informs an employee, and that resemblance is what puts it in the wrong box. Filing it as support keeps the model out of the count regulators use to decide what to examine, and out of any question about whether its answers were right. Those answers still go out to policyholders, unchecked.

A model that is counted can still be rated low by mistake, and the same representative is usually why. Most carriers put two guardrails around a model's answers. The representative is expected to check it before it goes out, and the platform runs its own internal check against policy language. Both are controls, and because they are built into the process, crediting them when rating the model is an easy mistake to make. The supplement asks for inherent risk, the risk the model carries before any control is applied, and the word "inherent" is easy to read past. Under the supplement's own terms, the representative and the platform check belong in the residual rating, after controls. A carrier that credits them in the inherent rating has rated its model lower than the supplement intends, and taken it out of the testing questions in the process. Both judgments are already recorded, one line per model, in the carrier's own inventory.

Why the cheaper rating costs more

Rating a model low is the cheaper answer. It closes a line on the inventory, raises no question the carrier has to answer, and creates no obligation to find anything out. Rating it high creates recurring work: finding out whether the model's coverage answers were right and continuing to do so while the model runs. The pressure runs one direction, and nothing in the supplement pushes back. The saving lasts a quarter. The gap it leaves is permanent. That work cannot be backfilled later, because the evidence exists only while the conversations are happening, and controls do not close the gap. A representative who catches a wrong answer fixes that conversation and leaves no record of it, so even a carrier with an excellent review process cannot separate an isolated error from a model-wide one. It also cannot see drift. Models do not hold still after launch. They get updated, the policy forms they draw on are amended, and the questions customers ask shift. A test run before deployment describes the model that was deployed, not the one answering calls this quarter. What a low rating costs is a year of conversations nobody measured, and nobody can go back and measure.

I cannot tell you how often AI models give wrong coverage answers, and neither can anyone else, because almost nobody is measuring. The draft is open for comment through Sept. 29, with a vote on a later version expected in November. The questions are being written now, and the evidence they ask for can only be gathered before they are asked. Rating a model low does not change what it told a policyholder. It only changes whether anyone was looking.

Insurance Agentic AI Needs New Business Models

Asian insurers show that agentic AI success requires reimagining market forces, not just chasing productivity gains through isolated use cases.

Insurance

What if your computer didn't just help you type or access YouTube but acted like a smart assistant instead? That's what we broadly call agentic AI. Some might think AI is here to help us improve productivity and reduce costs, but the change is much deeper (Afanasyev, 2026). It defines what the insurance industry would be in five years and how your organization should adapt.

Asia's First-Mover Advantage and Early Missteps

Asia is the first-mover in AI adoption among financial institutions, compared to developed markets often hindered by tech debt and bureaucracy. The learnings here uncover that, for the industry, agentic AI is a change in the "rules of the game."

Asian financial institutions originally rushed into Agentic AI by treating it like a project focused on multiple use cases, prioritizing 'impact vs. feasibility'. This approach targets quick wins, assuming the organization is the only player implementing AI and hence making the journey risky and expensive. No surprise to see both academia, MIT (Challapally et al, 2025) and consultants, BCG (Apotheker et al, 2025), report that 95% of organizations struggle with Agentic AI. Such organizations miss that their customers and partners are also deploying AI Agents (Citrini and Shah, 2026).

When Productivity Gains Backfire: The Email Case Study

What happens when Agentic AI use cases are assessed without context? Take AI for emails, a productivity use case deployed three years ago, which used AI to help write emails or summarise the received emails. It seemed like a win but lacked context; certain employees used AI Agents to imitate work and secure visibility by generating volumes of emails. AI-generated content is hard (with some researchers claiming impossible) for a human or an algorithm to distinguish from genuinely important messages. The recipients would have to apply AI Agents to search for information. Both senders and receivers myopically reported AI bringing gains, although the insurers likely deployed a use case that resulted in miscommunications and productivity drains.

The New Reality: Everyone Has AI Agents

The right approach for an insurer deploying AI is to accept that others are doing the same, so focusing on internal productivity may not be enough. Think about claims. Your policyholders' and providers' AI Agents are scrutinizing policies and preparing documents to maximize approval. Recall your dental insurance policy, which offers reimbursements only if certain symptoms, such as tooth pain, are present. The AI Agent ensures that the symptom is stated, even when the policyholder visits dentists for routine checks.

In underwriting, similar to how Mythos from Anthropic identifies cybersecurity vulnerabilities (Bloomberg, 2026; Azhar and Williams, 2026), your customers' agents might be identifying gaps. Insurers whose underwriting processes do not handle AI-generated applications may end up underpricing risks. For wealth management, why would a customer having access to LLM models pay advisors who use generic LLMs for generating ideas?

Think about distribution, where an insurance broker deploys AI Agents to identify and message prospects. The competing brokers are doing it too, leading to overwhelmed prospects who might avoid purchases altogether.

The latter is highly illustrative and applies in reverse to a financial institution thinking about its agentic strategy. Previous breakthroughs, such as electricity and the internet, accelerated business but the ability to make intelligent decisions broadly remained intact. AI Agents generate volumes of persuasive, credible content but the overload reduces decision-making ability.

The "Business First" Approach: Three Pillars for Success

Asian financial institutions concluded that to succeed, they need to shift to the "Business First" or "Applied AI" approach, which combines three views—Business domain, Transformation, and Technology—under a single leader (Afanasyev and Milind, 2026). McKinsey used to name leaders with such skills as "Advanced Analytics Translators" and, since the early AI days, has been urging organizations to grow leaders who "ensure that organizations achieve real impact from their analytics initiatives" (Henke, 2018).

Pillar 1: Business Domain — Reimagining Market Forces

The Business domain reimagines how market forces will change, how the industry will provide value, and what an organization's role will be in the long run. This is the least-discussed pillar and precisely what puts organizations at risk.

The majority of Asian organizations are privately controlled, meaning limited dependency on reporting cycles and an ability to prioritize long-term needs. Rapid economic changes over the past two decades have led them to embrace practicality and challenge the status quo. For instance, in developed markets with high labor costs, many insurers seem to be prioritizing Agentic AI technology to improve efficiency and cut costs. Most Asian markets are, however, known to have low labor costs, and local institutions have rightly questioned us about deploying AI for just operational efficiency in the regional context. However, young people, the largest segment of Asia's population, adopt AI Agents for their financial needs the fastest. This has accelerated economic changes not yet noticeable in developed markets.

How Generative AI Disrupts Insurance Fundamentals

We have discussed how Generative AI disrupts signaling theory, which is foundational for modern insurance. In some cases, the policy buyers may know more about the likelihood that they will suffer a loss than the insurance company. The 'market for lemons' theory shows that only the ' worst ' buyers (expecting larger, more likely losses) may end up interested in a policy. Signaling theory explains how parties with superior information (actors, for example, policyholders) convey attributes to parties with less information (decision-makers, for example, policy underwriters) to overcome information asymmetry. For the industry to function, credible signals—think medical records for life insurance—must be costly to imitate. Generative AI has empowered any actor to generate strong signals, disrupting the fundamentals of risk-taking, with empirical studies already demonstrating an erosion of trust.

How can actors credibly signal quality when signals can be cheaply fabricated? The answer, we believe, is grounded in mechanism design, the economics discipline concerned with a decision-maker designing the rules so that rational actors reveal truthful attributes even when their interest might be in using Generative AI. The Applied AI leader has to possess deep insurance domain knowledge and also a practical level of academic perspective. Such a leader analyses disruptions to come up with a view on what the organization's role would be. The outcome could be that the value proposition should be in risk prevention instead of risk pricing; changing target customer segments; vertical integration bringing certain services in-house; or leveling AI to address moral hazards.

Pillar 2: Transformation — Rethinking Processes and Culture

Transformation should focus on agility, changes in processes, decision-making, and human roles. "A growing number of studies reveal that human–AI systems do not necessarily achieve better results than the best of humans or AI alone" (Vaccaro et al, 2024). Many organizations ground their AI performance evaluation on efficiency, accuracy, or productivity. However, "nowadays, productivity gain is no longer the single evaluation criterion. In many instances, computer systems are expected to enhance our creativity, reveal opportunities and open new vistas of uncharted frontiers" (Avital et al, 2009).

Legacy institutions have to unwind certain cultures religiously followed for years. Certain advancements have ensured that employees follow rule books and processes prescribed by management. As a result, many organizations operate as factories, with processes equivalent to production lines of the early 1900s. Employees and mid-management have limited or no discretion outside of what is prescribed. This practice might have been appropriate earlier, but the Agentic AI economy is unpredictable; rule books and processes become outdated in the blink of an eye.

Pillar 3: Technology — Building the Infrastructure for AI Agents

The Technology view is most discussed and depicts how non-human resources are augmented by AI Agents. The learning from Asian progress in Agentic AI is a need for Applied AI leaders to take a broad view on financial industry technology and ground hot topics, such as tech debt, data quality, MCP and A2A, Agentic AI workflow designs, etc.

Asia's Unique Technology Landscape

Asia is home to 5 billion people—about 10x of Western Europe and US populations combined—with interactions among various ethnic groups who have their distinct languages, cultures, and business styles. Asia accounts for about 50% of non-cash transactions globally (Capgemini, 2025), with most countries, including India and China, already cash-free and unimaginable GDP growth in developed markets. Asian institutions deal with unprecedented transaction volumes and diversity, so scale and agility have traditionally been prioritized over politics and offering sophisticated products. Moreover, many overly complicated software products adopted in developed markets for the past two decades found little relevance in Asia.

Before LLMs and Agentic AI, the Asian tech ecosystem had been focused on developing local software solutions to handle volumes. Native software providers have mastered agentic workflow designs using Directed Acyclic Graphs with state management, allowing agents to "Plan, Act, Observe and Reflect". MCP, the "USB-C for AI", has landed in Asia as a standardized interface, allowing agents to pull real-time context from sources like SQL databases, CRM systems, and core banking. For developed markets, MCP would solve the chronic pain point of tech debt by creating an abstraction layer over legacy core systems. Complementing this is the A2A protocol, which facilitates peer-to-peer coordination, and AP2, empowering agents to do financial transactions with or without human presence.

Architectural Challenges: Making AI Agents Work Together

Most legacy systems were architected for workflows that functioned adequately for rule-based operations. Meanwhile, Agentic AI demands open interfaces where workflows are not hardcoded and any given process can be decomposed, assigned to an agent, and reassembled. Organizations are adopting AI capabilities in separate departments and in relative isolation. The containment limits risk, allowing teams to demonstrate value; the actual issue is how these modules function as a coherent whole.

Agent loops, where Agent A triggers Agent B which re-triggers Agent A, require modern architectures to include circuit breakers that terminate chains exceeding defined depth or duration thresholds. Agent disagreements surface when two agents with complementary mandates reach different conclusions. Authority boundary violations occur when agents act beyond their defined scope.

Data Governance and Schema Management

A schema-driven model is recommended. Insurers should adopt a master schema configurator that allows users to govern agent behavior without code. Insurance is challenging in the volume of its unstructured data. Decades of paper-based documentation and inconsistent digitization have left organizations with data states that cannot be queried in real time. For Agentic AI to function, three areas are foundational. Vector databases enable semantic search across document corpora, allowing context retrieval. Structured extraction pipelines convert unstructured documents into queryable records. Master data management maintains reference datasets which can be used for real-time lookup, matching and adjudication.

Regulatory Safeguards: Ensuring Accountability and Control

As regulators, we are of the view that autonomous systems making customer-impacting decisions must be auditable, explainable and interruptible (IAIS, 2024). To ensure agents do not overstep authority, safeguards must be included. Confidence scoring requires that every AI output carry a numerical confidence estimate, with automatic escalation to human staff when results are low. Mandatory review workflows gate high-stakes decisions behind human approval that cannot be bypassed. Continuous feedback loops mean that each correction trains the models. Comprehensive audit trails provide thorough records. 

A Five-Step Modernization Roadmap

A five-step modernization approach is required:

Step 1: Start with structured discovery. Most attempts fail because organizations jump to engineering straightaway. Insurers must produce a clear blueprint of their processes. This discovery phase should identify opportunities to separate capabilities into components — units that can become plug-and-play services with open APIs, allowing future agents and partner systems to come in flexibly.

Step 2: Separate design and validation from execution. Insurers should introduce a solutioning phase where a small team defines the future state by determining which capabilities will be rebuilt or replaced, how workflows should operate, etc. Only once the design has been validated should engineering teams begin executing. AI systems can assist by extracting undocumented patterns from interviews, translating requirements into architectures, and simulating potential models to stress-test assumptions.

Step 3: Build in parallel. The target platform should be developed independently while legacy systems continue running operations. A specific sequence is recommended — load non-member-facing master data first, accepting a short window (as brief as possible) during which updates may need to be reflected in both legacy and new systems. Once it's loaded, validate the setup by running test policies.

Step 4: Migrate data through controlled, incremental processes. A single large cutover concentrates risk and creates extended recovery windows if things go south. Hence, a controlled migration strategy is beneficial. Bring the new system online and then migrate data in staggered batches. Not all data needs to migrate; historical data is often better suited for archival storage or data lakes. Before migration, the new platform should be tested in lower environments. Once validated, the system can go live while legacy systems continue managing the existing portfolio. AI systems support this by inferring undocumented schemas, mapping fields, and extracting structured data from records.

Step 5: Retire legacy systems. The final, frequently overlooked step is decommissioning legacy systems. Leaving them partially operational introduces unnecessary cost and risk. Instead, insurers should archive historical data where necessary, ensure compliance for record retention, and fully decommission said infrastructure.

The Kodak Moment: Adapt or Become Obsolete

Organizations focusing on productivity use cases miss the big picture, but the insurance industry will likely not exist in its current form in five years. Think about Kodak 30 years ago: a leader in optimizing film production costs that missed how digitalization changed preferences. In this democratization, insurers need to shift to business models redesigned to cater to change and structure their AI journeys accordingly.

For the full white paper from which this article is adapted, click here.


Maxim Afanasyev

Profile picture for user MaximAfanasyev

Maxim Afanasyev

Maxim Afanasyev is a senior executive in AI, solutions and product management at Google.

He has more than 20 years of experience in academic research, management consulting, technology and financial services and is an advisor to C-levels of major companies and prominent entrepreneurs.

Afansyev has a PhD from Stanford University in operations, information and technology. He is an adjunct associate professor in AI at the National University of Singapore.


Tomas Holub

Profile picture for user TomasHolub

Tomas Holub

Tomas Holub is the CEO and founder of CoverGo, an AI insurance platform for health, life, and P&C. 

Prior to starting CoverGo, he worked as an insurance and banking consultant at PwC London and also head of operations of an insurance technology company in Singapore. Overall, he has worked in 20 countries across three continents, speaks eight languages and obtained four masters degrees in risk management, international business and public and business administration.

Holub is a frequent speaker at conferences.