Download

How to Move Insurance AI From Pilot to Production

Moving AI from pilot to production requires carriers to master data infrastructure, production architecture, and user experience design.

An artist's illustration of AI

Artificial intelligence has moved from experimentation to active deployment across the insurance value chain. Predictive models are augmenting underwriting, generative AI is accelerating claims and policy servicing, and agentic systems are beginning to coordinate multi-step workflows that previously required human handoffs.

Carriers that move AI into production consistently address three requirements: a data foundation built for AI consumption, an architecture engineered for production conditions, and an experience layer designed for the people who use it.

The data foundation

Connecting data for AI consumption is the first major engineering effort in any serious AI program. Policy systems, claims platforms, loss history, external feeds, and regulatory data have accumulated in separate architectures over decades. Each was built to serve a specific function, and connecting them for AI requires deliberate work. The design choices made at this stage carry forward into every subsequent AI operation.

Where data lives determines the cost and compliance profile of those operations. Running inference inside a governed platform already equipped with access controls, audit logging, and encryption carries lower compliance exposure and lower per-operation cost than routing data to an external model API. That decision is made early and is expensive to reverse.

The highest-value early work in most programs is automation and data engineering. Normalizing loss runs, structuring adjuster notes, and building a reliable integrated view of a risk generate analytical value before any model is involved. These steps build a foundation that extends to subsequent use cases without being rebuilt each time.

Production-ready AI architecture

A production-ready architecture must do three things well: control what runs, make it run reliably at scale, and provide clear visibility into whether it is performing as expected. These elements depend on one another.

First, every agent in production must be pinned to a specific, documented model version. A change to an agent's instructions carries the same functional impact as a change to a model's parameters. Both require the same change management controls. The compute engine decision — in-warehouse versus external — determines data residency, latency, cost, and compliance exposure. That routing decision should be explicit, documented, and revisable.

Second, the system must handle real production conditions. Inference at scale requires prompt caching, token quotas, and cost attribution by agent, use case, and business unit designed in from the start. At peak underwriting volumes or during catastrophe response, uncapped spending quickly becomes a budget event. Long-running tasks need async processing and checkpointing so they can resume cleanly after interruptions. Failure handling — dead letter queues, retry logic with backoff, and idempotency — must be part of the original design so transient outages do not become analyst problems.

Third, the architecture must make performance visible. Override rates serve as the leading indicator of model quality in production. When underwriters or adjusters consistently modify AI outputs, something has shifted in the model, the data, or the business context. Distributed tracing with shared correlation identifiers across every service call turns failure diagnosis from a reconstruction exercise into a lookup. Every AI decision must be recorded with its inputs, agent version, model parameters, and any human override so that when a regulator asks how a specific underwriting decision was made, the answer is already waiting in the log.

The experience layer

Platform selection shapes what provenance is even possible, while the interaction pattern determines how that provenance reaches the user. For work involving policy data, medical information, or PII, the delivery platform must keep that data within the appropriate governed perimeter. The compliance exposure from getting this wrong surfaces at examination time.

The interaction pattern should match how the work actually gets done. Conversational interfaces suit knowledge retrieval. Structured outputs suit decisions feeding downstream systems. Embedded AI integrated into the application an underwriter or adjuster already uses suits workflows where adoption depends on minimizing context-switching.

Underwriters and claims professionals acting on AI-generated outputs need to know what data the output was based on and whether a human reviewed it. Without provenance, usage patterns split between over-reliance and skepticism.

Starting the journey

Build each component for the production environment from the first use case. Size the data foundation so it can extend beyond the initial project. Design the architecture for the actual load it will face, with observability and failure handling included from day one. Shape the experience layer around the real workflows of underwriters and adjusters, using platforms that already satisfy the compliance requirements of the data involved.

Data Standards Key for Insurance M&A

As insurance M&A passes $12 billion just this year, post-merger execution and data standards prove more critical than the deal rationale.

Detailed view of stock market charts and data on a monitor, showcasing market trends.

More than $12 billion worth of insurance M&A deals have been announced so far this year, up from $10 billion last year, demonstrating the importance of M&A as a strategic tool. For carriers, brokers, and technology providers alike, M&A can accelerate growth, expand capabilities, and reposition firms for an increasingly digital and data-driven marketplace. Yet experience has shown that while deals are relatively easy to announce, they are far more difficult to execute successfully.

ACORD's recent Carrier Mergers & Acquisitions study sought to better understand both deal rationale and – more importantly – the implications and imperatives that drive successful outcomes. The study revealed clear differences in both the prevalence and effectiveness of carrier M&A strategies across four distinct deal rationales.

Diversification is an attempt to expand the portfolio by acquiring new revenue and earning sources. It accounted for 41% of carrier deals and delivered strong post‑transaction returns (+14%), making it the most prevalent and one of the most effective drivers of insurance M&A.

Core Expansion increases share across areas in which the insurer already executes, such as products, geographies, channels, and customer segments. These transactions represented 29% of deals and generally produced positive outcomes when valuation discipline, relationship preservation, and integration execution were well managed.

Scale & Scope goals include amortizing fixed costs and improving resource access by increasing absolute size, or expanding scope across strategic and tactical dimensions. These deals ranked third in frequency but were the only category to generate negative returns (‑14%), reflecting the risk of overstated synergies, underestimated integration risk, and diseconomies from added complexity.

Capability Acquisitions intend to optimize the risk, cost, and time associated with developing new or enhanced internal capabilities. While only 6% of transactions, these generated the highest returns (+28%) by closing strategic gaps and improving competitiveness rather than relying on scale alone.

Post-Merger Execution: Risks and Imperatives

Analysis of post‑transaction performance shows that outcomes are shaped far more by integration execution than by deal rationale, creating a clear divide between successful and underperforming transactions. Across the more successful transactions, acquirers focused on converting the deal thesis into a small number of high‑impact initiatives with clear ownership, targets, and governance under a single integration authority.

Early leadership and cultural alignment, protection of core operations, a defined target operating model and technology roadmap, and proactive talent retention were critical. The best performers track synergies through a single source of truth, communicate consistently, and sequence integration to deliver near‑term, no‑regret value before more complex change.

For transactions that fell short, value destruction was driven primarily by execution, not deal logic. Cultural friction, leadership misalignment, and technology integration complexity routinely slow decision‑making, erode talent, and dilute expected scale and efficiency benefits – risks amplified by regulatory burden and integration fatigue. Without disciplined value‑capture management, synergies identified in diligence often dissipate during integration.

The Role of Standards in Post‑Merger Integration

Standards are a critical execution lever in post‑merger integration, particularly in data‑intensive and highly regulated insurance environments. Industry data standards provide a foundation for reducing integration risk and accelerating value capture by addressing common sources of post‑deal friction. In many transactions, value leakage stems not from flawed strategy but from inconsistent data definitions, fragmented messaging, and limited interoperability across legacy and acquired systems. Common standards mitigate these challenges by establishing a shared insurance data language across underwriting, claims, policy administration, billing, reinsurance, and finance.

In the near term, industry standards enable effective "bridge integration," allowing disparate systems to exchange standardized data and messages without immediate core system replacement. This accelerates the consolidation of operational and financial reporting, enabling leadership to establish a credible single source of truth for synergy tracking, performance management, and regulatory oversight. Standardized data also reduces reconciliation effort, manual workarounds, and control gaps that frequently obscure results and slow integration progress.

Common data standards support several high‑value integration use cases across the insurance value chain. In underwriting and product management, common data models improve portfolio visibility and enable faster rationalization across products, lines, and geographies. In claims and servicing, messaging standards support more consistent customer and distributor experiences during integration, reducing disruption and service degradation. In reinsurance, finance, and risk management, standardized data structures enhance exposure aggregation, capital reporting, and regulatory compliance across the combined enterprise.

Over the longer term, standards-based architectures provide a stable foundation for phased modernization and system rationalization. By reducing reliance on bespoke point‑to‑point integrations, standards lower cost, improve flexibility, and shorten time‑to‑value for future acquisitions.

As insurance M&A accelerates and transactions grow larger and more complex, post‑merger execution – not deal ambition – will continue to drive shareholder value. Leading organizations will distinguish themselves by treating integration as a strategic capability, embedding discipline, governance, and data alignment from the outset. Industry data standards enable speed, control, and transparency while preserving future optionality. Positioned correctly, standards help protect the franchise, manage execution risk, and sustain value creation in a data‑driven insurance industry.


Dave Sterner

Profile picture for user DaveSterner

Dave Sterner

Dave Sterner is the senior vice president of research & development at ACORD. 

He has over 20 years of experience in insurance. 

Sterner is a graduate of Drexel University's LeBow School of Business Administration, where he earned both a bachelor of science degree in finance and marketing and an M.B.A.

Governance Infrastructure Is Key for Agentic AI

Agentic AI's rapid deployment in underwriting and claims is outpacing the governance infrastructure insurers need.

Close-up of a professional in a suit reviewing a document with focused attention

TLDR: Deploying agentic AI without governance infrastructure accumulates regulatory and operational exposure faster than most carriers recognize. Carriers scaling with confidence are those that made governance foundational and not remedial.

The Governance Infrastructure Insurers Need

Agentic AI is taking on consequential decisions across underwriting and claims at unprecedented speed and scale. The governance infrastructure at most insurers has not kept pace.

A simple prompt change - a few lines of text updated in a configuration file - can alter how an AI agent reasons about risk across every submission it processes. In a traditional predictive model, the equivalent change requires a full retraining cycle - weeks of documented work, validation runs, and a formal change record. In an agentic system without governance infrastructure, the same functional impact happens with no audit trail. Regulatory and operational exposure accumulates as a result.

Organizations scaling agentic AI with confidence have recognized this gap early and built the infrastructure to close it - because the ability to answer basic questions about their AI systems is a precondition for operating them responsibly at scale.

Why Agentic AI Strains Traditional Model Governance

Insurance organizations have spent years building model risk management capability for predictive AI - pricing models, fraud scores, and reserve estimates. These systems are well understood in governance terms. They take defined inputs, apply learned parameters, and produce a single output that a human then acts on.

Agentic systems change this structure in fundamental ways. Reasoning chains span multiple steps - querying data, evaluating evidence, calling external tools, forming intermediate conclusions - before producing a final result. A simple input-output log is no longer adequate.

Prompts are the governing parameters for AI agents. Changing a system prompt is functionally equivalent to changing a model's weights. Tool calls are data transactions that can potentially send confidential information to third-party systems. Human oversight placed only at the end of the workflow misses the dozens of consequential micro-decisions made along the way.

What Regulators Are Expecting

SR 11-7 is the de facto governance baseline. State insurance regulators are applying its principles - model inventory, independent validation, change management documentation, continuing monitoring - in examination practice. Any carrier deploying AI that influences underwriting or claims decisions should treat SR 11-7 as the minimum standard.

The NAIC AI Model Bulletin's eight principles are now shaping market conduct exams. Explainability carries the sharpest operational bite: if an AI system influenced a coverage decline, the insurer must be able to explain why in specific, contemporaneous terms. Colorado SB21-169 adds testing for proxy discrimination, documentation of external data sources, and annual certification. Similar requirements are advancing in California, New York, and elsewhere.

One pressure that did not exist two years ago is now concrete - carriers without documented AI governance frameworks are facing coverage exclusions and premium increases on their own AI liability policies. The industry that applies governance scrutiny to its insureds is now applying it to itself.

Governance Capabilities

A governance framework for agentic AI requires six distinct capabilities. Each addresses a specific gap in how these systems are built, operated and governed.

  1. Asset registry - Answers the first question regulators ask: what AI is in production, who owns it, and what version is live. Every agent, task, prompt, and tool is stored as a versioned database record - not hidden in code. SR 11-7's model inventory requirement and the NAIC transparency principle both resolve to this capability.
     
  2. Lifecycle framework - Enforces the change management discipline that agentic systems otherwise lack. Every asset version moves through a defined sequence of states - from draft through shadow deployment to champion - with human approval gates at the points that matter. A prompt change cannot reach production without the same controls applied to a code change.
     
  3. Testing & validation - Replaces ad-hoc demonstration with structured evidence. Before any agent version reaches production, it is tested against ground-truth datasets labeled by subject-matter experts, producing precision, recall, and fairness metrics stored against the specific version. Colorado SB21-169's bias testing requirement and SR 11-7's independent validation requirement both have a direct answer here.
     
  4. Execution control - Ensures that what runs in production is exactly what governance approved. At runtime the framework pulls the current champion configuration from the governed registry. Agents can access only authorized tools and approved parameters. Governance is enforced at execution, not assumed after the fact.
     
  5. Decision logging - Produces the contemporaneous record that the NAIC explainability requirement and Colorado's adverse-action provisions demand. Every AI invocation is logged with the exact prompt version, inference parameters, tool calls, inputs, and output. When a market conduct examiner asks how a specific decision was made, the answer is a query, not a reconstruction.
     
  6. Compliance & reporting - Makes the governance data useful to the people who need it. Model inventory reports, model cards, override-rate analysis, and regulatory evidence packages are generated on demand from the records the other five components accumulate, not assembled manually when the examination notice arrives.
The Infrastructure Argument

Policy documents alone cannot operationalize these capabilities. A policy requiring documented approval for all agent changes creates the obligation but not the mechanism. Under operational pressure, informal processes prevail.

The insurance industry has already solved an analogous problem in actuarial pricing systems. Algorithms are versioned, changes require documented approval, prior versions are retained for audit, and the system generates its own compliance record. No one would consider deploying a new rating algorithm by editing a configuration file with no version control. That standard of infrastructure is exactly what is needed for AI agents.

When agent behavior is embedded in code, compliance teams cannot access it without engineering support. Storing agent definitions as versioned configuration records changes this - any authorized reviewer can see exactly what instructions any agent version was operating under at any moment in time. Built on that foundation, the framework can enforce the lifecycle mechanically, log every decision with the version that produced it, surface override patterns automatically, and generate regulatory evidence on demand.

The carriers that have built this treat governance as an engineering problem, not a policy exercise.

Where to Start

Start with the agent registry. For each AI system in production, name it, document what it does, record the live version, and identify its owner. Add prompt version control before the next agent change, and decision logging before the next production deployment. Governance built incrementally as infrastructure - before the examination, before the finding, before the failure - is faster to production and more durable than governance assembled in response to one.

The regulatory direction is clear, and the examination questions are already being asked. The difference between carriers that answer them confidently and those that cannot is infrastructure that existed before the question arrived.

How to Reframe Operational Challenges

Operational challenges often become rationalized clutter; reframing them through expertise rather than experience unlocks breakthrough solutions.

Long External Stairs in the Facade of the Building in greyscale

Has a family member ever given you a gift you can't bear, yet can't refuse, and it simply becomes part of the decor? It might be a decanter so impractical that it's ornamental, but it has to be brought out every time they come over; or a portrait that asks fundamental questions about the nature of your relationship, yet over time you no longer register that it's there. 

Operational challenges can be like this; unwanted gifts that become clutter, obstacles that are easy to rationalize. As they accumulate, they require incremental effort to navigate and leach efficiency. Yet when we approach a familiar operational challenge from inside the organization, we risk framing the challenge so narrowly that we're boxed in with too few options available. We refer to this as approaching challenges through a lens of our experience - and it can become part of the problem, rather than a means of solving for it.

Seeing a problem through the lens of our experience describes a way of seeing that includes all our knowledge of the history of the problem. All the attempts to resolve it, the failures, the frustrations; it's the voice that says, "We've tried that before and it didn't work." Returning to our furniture metaphor, it's not dissimilar to saying, "We can't move that painting. We took it down once, and my brother got upset." The lens of experience is effective at keeping you on the same track but it's less likely to help change direction. Evaluating a persistent operational challenge through a lens of expertise is vastly more effective.

Approaching a familiar challenge through a lens of expertise means stepping outside of the challenge, viewing it more objectively, and applying our knowledge to that problem. This is the secret sauce of consulting, the classic "outside-in perspective," yet it's possible to strengthen this capability within your own organization. The key is understanding how changing the structure of a problem helps to create new ways of seeing it. By carefully evaluating a problem and adjusting its constraints, experienced operators can see a familiar challenge with a broader perspective, and then bring their hard-won expertise to bear.

I worked with an insurance property repair firm whose leaders shifted their focus from a lens of experience to a lens of expertise with spectacular results. They were part of an insurer's repair vendor panel and found themselves competing across a broad range of repair categories, tackling jobs that ranged from minor fence repairs and garage doors through to major reconstructive work for insureds. Smaller jobs only required general handymen - low cost, low risk, and the pool of available contractors was broad - whereas the larger jobs required more skilled trades and more oversight - higher cost, higher risk, and a narrower pool of trades. Larger firms on the panel could absorb the occasional job that went off the rails, but this firm was small enough that even one or two jobs that went over budget hit profits hard. That was the model. Until this firm opted to re-imagine and renegotiate their panel membership.

The repair firm reimagined their business in two stages: first, they negotiated with the carrier to remain on the panel as a "small repairer." They would only accept smaller repair work but take higher volumes. This was feasible because the pool of trades was large and - given the nature of largely weather-related property damage - jobs were often geographically co-located. One trade could attend multiple sites in a day, which allowed for bundling and improved efficiency. In exchange, the repairer would offer a reduced rate because they weren't subsidizing larger jobs.

Second, they re-designed their operations from within by re-structuring their project management approach. They turned the entire model upside-down, from how they hired trades and retained them to how they would project manage each job. Each repair was broken into its discrete segments (plastering, painting, electrical, and so on) and were arranged such that the right trade attended at the right time - a virtual production line. Trades tapped in and out on their cellphone app, which gave the business visibility of their activity, plus allowed for them to estimate the time required for each job - a feedback loop that informed project, pricing, and contract-hiring forecasts.

The results were significant. The carrier ultimately integrated the model directly into its property claims flow, allowing customers to move from first notice of loss to completed repairs with a speed that hadn't previously been possible. Customer satisfaction ratings exceeded 90%. The firm had transformed itself not by responding to competitive pressure, but by isolating the fundamental conditions of their business and restructuring them to reveal entirely new ways of operating.

Your operations function may be more or less complicated than this example, but there's likely at least a handful of persistent challenges you'd love to unpick. Start by examining the assumptions and constraints that shape how you interpret the problem. Change those, and new solutions will follow.


Chris Bassett

Profile picture for user ChrisBassett

Chris Bassett

Chris Bassett is a management consultant with over 10 years of experience in operations strategy. 

He is the founder of Green Bean Consulting Group, which helps leadership teams step outside familiar thinking to tackle complex operational challenges more effectively.

Claims AI Requires Strong Operational Guardrails

The most important question in claims AI is not whether a model performs well on average. It is what happens when it does not.

Winding Forest Road in Early Spring

Artificial intelligence is already changing insurance claims operations. It can shorten cycle times, improve fraud detection, reduce administrative costs, and help carriers handle routine claims with greater speed and consistency. Those benefits are real. But the difference between a useful AI system and a risky one is rarely the model itself. It is the control environment around it. 

After 15 years in financial operations across telecommunications, banking, and healthcare, I have learned that systems do not usually fail because they produce outputs. They fail because organizations do not build the right controls for what happens when those outputs are wrong. That lesson is especially relevant in insurance claims, where AI can recommend payments, trigger denials, or escalate fraud investigations at speed and scale.

This is why the most important question in claims AI is not whether a model performs well on average. It is what happens when it does not. Who reviews the outlier decision? What happens when source data is incomplete or inconsistent? Which claims are allowed to move straight through, and which require human judgment? Without clear answers to those questions, automation creates exposure faster than it creates value. 

The insurance industry has made real progress. Many carriers now use AI in some part of claims handling, especially for low-complexity workflows. But mature deployment remains limited. The gap is not just technical. It is operational. Insurers often struggle with fragmented data, inconsistent workflows, weak escalation paths, and governance models that are more aspirational than enforceable. In practice, that means claims AI often performs inside silos rather than inside a coherent control framework. 

In financial operations, this kind of weakness is familiar. I have seen organizations lose significant revenue not because the systems were incapable, but because no one had defined what should happen when an exception appeared. In one credit control role, I identified more than $10 million in revenue leakages. Those leakages persisted not because no system existed, but because process gaps allowed errors to go unchallenged. Claims AI creates the same risk, except with higher speed, broader scale, and greater regulatory sensitivity.

So what guardrails actually work?

Human review for non-routine claims. Straight-through processing can be appropriate for low-value, low-complexity claims where the decision logic is narrow and well tested. But once a claim involves material exposure, medical complexity, ambiguity in coverage, or fraud indicators, human judgment must re-enter the process. This is not resistance to AI. It is sound risk design.

Explainability for adverse decisions. If an AI system recommends denial, escalation, or fraud review, the rationale must be understandable to the people accountable for that outcome. An adjuster cannot meaningfully supervise a recommendation that cannot be explained in plain terms. Explainability is not just a technical preference. It is the basis for accountability, defensibility, and fair review.

Continuous data-quality control. AI systems do not fail only because of bad models. They also fail because of incomplete, stale, fragmented, or poorly governed data. In claims operations, a data issue is not a minor defect. At scale, it becomes a multiplier of bad decisions. Regular review of upstream data sources, transfer points, and exception patterns is essential.

Defined exception and escalation pathways. Every model has edge cases. Effective governance assumes this from the start. Claims that fall outside confidence thresholds, conflict with policy logic, or present unusual fact patterns should move automatically into a structured review queue with identified owners and documented next steps. In strong operating environments, exceptions are not left hanging. They are routed.

Active regulatory monitoring. AI governance in insurance is no longer an internal policy matter alone. Carriers now operate in an environment of increasing scrutiny around disclosure, fairness, bias, consumer protection, and human oversight. Any organization deploying AI in claims must treat compliance monitoring as part of the operating model, not as an afterthought.

It is equally important to be clear about what does not work.

Principles without enforcement do not work. A statement about responsible AI is not a control unless it is backed by auditability, accountability, and operating discipline.

Black-box decision making in high-stakes contexts does not work. A model that cannot be explained may still produce accurate outputs in aggregate, but it creates real risk when applied to adverse decisions that affect claimants and attract scrutiny.

Deployment on unvalidated source data does not work. AI does not fix weak data foundations. It accelerates the consequences of them.

Minimal staff training does not work. Claims professionals do not need to become data scientists, but they do need enough AI literacy to interpret outputs, question recommendations, recognize limitations, and escalate when needed.

The operational stakes are high. Carriers that deploy AI well can improve speed, consistency, and cost performance. Carriers that deploy it poorly can create regulatory exposure, claimant harm, and reputational damage that overwhelms any efficiency gain.

In the end, the real issue is not whether AI belongs in claims. It does. The issue is whether insurers will build the operational discipline required to make AI trustworthy. The winning organizations will not be the ones with the most impressive demos. They will be the ones with the clearest controls, the strongest escalation design, the cleanest data discipline, and the most accountable governance.

AI can make claims operations faster. Only guardrails make them reliable.

Why Most Insurance AI Strategies Will Fail

Every major insurer has an AI strategy, but most will fail without the operating model to support it.

AI

Every major insurer has an AI strategy. Most of them will fail. Not because the technology isn't ready — it is. Not because the use cases don't exist — they do. The strategies will fail because organizations treat AI as a point solution rather than a platform, and they underestimate how fundamentally it demands a different operating model.

I spoke about this at ONUG (Open Networking User Group) last fall under the title "Beating the 4%: Why AI Fails." The thesis is straightforward: the vast majority of enterprise AI initiatives stall at the pilot stage, not because of technical limitations, but because of misalignment between technology investments and business operating models. Insurance, with its complex distribution relationships and legacy infrastructure, is particularly exposed to this failure mode. The question for carriers isn't whether to adopt AI. It's whether they have the architecture — organizational and technical — to make it stick.

Building the Right Foundation

The organizations that succeed with AI don't start with models; they start with alignment. Before any model goes into production, the business objective has to be clear, the data must be trustworthy, and the teams have to understand what the tool is solving and why. That sequencing matters more than the technology itself.

A consistent pattern across successful AI programs is centralized governance, a single framework through which development and deployment are coordinated. This isn't bureaucracy. It's how you prevent fragmentation. Without it, you get dozens of disconnected pilots, inconsistent data practices, and no coherent path to scale. With it, you build institutional muscle: teams that know how to evaluate, deploy, and improve AI solutions within a common framework.

Early applications often show up in operational improvements – streamlining workflows or improving access to knowledge — but those are table stakes. They're the foundation, not the destination.

Owning the Experience Layer

The real competitive battle in insurance isn't being fought in the back office. It's being fought in the experience layer — the digital surface where financial professionals and clients actually interact with your products and your brand. Carriers that own that layer will win. Those that cede it to distributors, aggregators, or fintechs will spend the next decade competing on price alone.

The real opportunity lies here: embedding AI directly into the workflows of financial professionals, not as a separate tool they have to context-switch into, but as intelligence woven into the systems they already use. That can take many forms – from meeting preparation and product insights to tools that help advisors refine how they engage with clients and improve over time.

The unifying thread is data. Personalized, trustworthy, AI-powered experiences are only possible when data is unified enterprise-wide. Without that foundation, you are not personalizing — you are guessing.

Democratizing AI Across the Organization

I've seen this pattern play out across industries – insurance, real estate, professional sports, and media & entertainment. The organizations that win with transformational technology are never the ones that centralize it in an IT function and call it done. They are the ones that democratize it: making AI capability accessible, legible, and useful to colleagues across every function.

The organizations making real progress are not just building AI tools – they are building AI fluency: helping teams understand how to interpret outputs, where to trust automation, and where human judgement remains essential.

The Carriers That Will Lead

With $483 trillion in projected retirement savings shortfalls by 2050, the demand for trusted financial guidance is only going to intensify. The carriers that are positioned to meet it will not be the ones with the most AI tools. They will be the ones that have built the right operating model, owned the experience layer, and treated AI not as a pilot project but as organizational infrastructure.

That is a harder problem than it looks. But it is the right one to solve.

Agentic AI Transforms E&S Policy Binding

As E&S market surges, agentic AI cuts policy binding from 21 days to three, transforming specialty insurance operations.

Side profile of an artificial intelligence face with data on the screen

Against a challenging commercial insurance landscape, the excess & surplus (E&S) market continues to demonstrate strong momentum. For the sixth consecutive year, E&S premiums have grown at double-digit rates, with U.S. domestic direct premiums written increasing 13% year-on-year to $98.2 billion in 2024, according to S&P Global Market Intelligence. This sustained growth reflects rising demand for flexible, non-standard risk coverage as traditional markets tighten underwriting appetite.

As the E&S market grows, inefficiencies in policy processing have become more pronounced. Agentic AI, by enabling autonomous, intelligent execution, directly addresses these gaps and delivers speed, consistency, and accuracy, redefining policy workflows for specialty and E&S insurers.

Market Dislocations and Operational Challenges

Despite strong growth, the E&S market continues to face structural dislocations. Segments such as umbrella and excess liability, catastrophe-exposed property, construction, commercial auto, and healthcare continue to face profitability pressure. Contractor liability in construction defect states is particularly impacted by long-tail exposures, complex legal environments, and inflationary cost dynamics. At the same time, increased competition is gradually softening the market, even as emerging risks such as supply chain disruptions and generative AI create new opportunities.

Within this environment, operational inefficiencies remain a significant constraint. Policy binding is still heavily manual and fragmented, with 73% of underwriters citing clause review as their number one-time drain. Each policy requires manually reviewing thousands of clause variants, often without intelligent recommendation support. This is compounded by the need to analyze 50–100-page risk engineering reports, where critical insights can be missed.

The result is a slow, error-prone process. The average time to bind a complex commercial policy remains around 21 days, while manual handling increases errors by 45% and annual rework costs by approximately $2.3 million per carrier. In multi-party environments such as the London Market or U.S. E&S segments, these inefficiencies are further amplified, delaying decision-making and affecting broker relationships.

Why Agentic AI and Why Now?

Traditional rule engines have long provided structure and compliance in underwriting workflows, mostly for admitted lines. However, they are not designed to handle unstructured data such as broker emails, PDFs, and bespoke clause language. They lack the ability to interpret context, detect nuanced conflicts, or adapt to evolving risk scenarios.

Agentic AI addresses this gap by combining large language models with multi-agent orchestration. These systems can parse unstructured data, recommend clauses from libraries exceeding 10,000 variants, detect conflicts in real time, and generate plain-language explanations for decisions. Importantly, agentic AI complements rule engines rather than replacing them, handling contextual reasoning while rules enforce deterministic compliance.

This shift enables insurers to move toward adaptive, intelligence-driven workflows, resulting in 60–99% faster quote-to-bind cycles and 3–5% improvements in loss ratios.

Reimagining Policy Binding Across Specialty and E&S Lines

Agentic AI is transforming workflows across both specialty and E&S insurance. In specialty insurance, an AI-powered policy binding solution enables rapid, compliant, and highly customized workflows to meet the market's complex needs. Submissions are seamlessly captured from multiple channels, with AI-driven extraction and validation of unstructured data, including bespoke clauses and risk details. Binding agentic AI automates clause selection, real-time conflict checks, and scenario-based underwriting, ensuring regulatory compliance and accuracy.

According to recent research on U.S. insurance sector growth in 2025, digital-first binding solutions have reduced cycle times by up to 50% and improved pricing precision. Such a solution also streamlines customer communication, automates documentation, and integrates with downstream systems, empowering underwriters to focus on risk assessment and strategic decision-making, while ensuring faster and error-free policy binding.

In E&S markets, where flexibility and customization are essential, agentic AI enables more contextual and dynamic decision-making. It builds multi-dimensional risk profiles using unstructured and external data, supports scenario-based underwriting, and facilitates faster negotiations through real-time analysis of broker inputs. In certain specialty segments, these capabilities have reduced binding times by up to 50% while improving pricing accuracy.

Across both markets, risk assessment becomes more comprehensive, placement decisions more precise, and negotiation cycles significantly shorter. At the binding stage, agentic AI ensures that all compliance and authority checks are completed before execution, while automating documentation and downstream processes.

Transforming Roles With Agentic AI

The impact of agentic AI is not limited to process efficiency; it is fundamentally reshaping roles across the insurance value chain. For commercial underwriters, AI-driven clause recommendations reduce what was once a four-hour manual search to under eight minutes. With pre-built risk briefs, underwriters can shift their focus from data gathering to strategic judgment and decision-making.

For wordings and compliance analysts, the benefits are equally significant. Agentic AI can detect conflicts across more than 200 clauses simultaneously while automatically validating jurisdictional requirements for every endorsement. This reduces manual review effort while improving consistency and regulatory adherence.

Insurance brokers experience faster turnaround times, with many policies moving to same-day binding. AI-generated counter-clause responses in plain language improve negotiation efficiency, while automated coverage summaries enhance client communication and transparency.

Operations and binding teams also see substantial gains. Pre-bind checklists are validated automatically, ensuring no conditions are missed. Policy documents are generated and distributed at the point of binding, and downstream systems, such as CRM, billing, and reinsurance platforms, are all updated seamlessly without manual intervention.

A Real-World Shift in Specialty Insurance

A leading public specialty U.S. insurer's transformation illustrates how these capabilities translate into practice. Facing fragmented workflows and manual processes, the organization modernized its operations by digitizing and streamlining end-to-end policy-binding workflows using a customer communication management platform and an enterprise content management platform. This improved turnaround times, enhanced compliance tracking, and provided a unified view of policy and customer data. As a result, the insurer reduced manual effort while strengthening its ability to manage complex risks and respond more effectively to market demands.

Delivering Measurable Business Impact

The adoption of agentic AI is delivering tangible results across the board. Policy binding times are reduced by 86%, from 21 days to just three days, while clause selection effort drops by 93%, from hours to minutes. Compliance breaches are reduced by 93%, significantly lowering regulatory risk.

Underwriter productivity increases by 175%, enabling them to handle 18–22 policies per week, while rework costs decline by 83%, from $2.3 million to approximately $380,000 annually. Brokers benefit from faster responses and improved service levels, and operations teams gain efficiency through automation. Together, these improvements allow insurers to scale operations without proportional increases in cost or headcount.

The Road Ahead

Agentic AI represents a turning point for the commercial insurance industry. By enabling faster, more accurate, and scalable policy binding, it allows insurers to move from reactive processes to proactive, intelligence-driven operations. As competition intensifies and risks evolve, the ability to process unstructured data and act with speed will define success. The future of policy binding is not just faster, it is smarter, more adaptive, and built for complexity.

AI Penetration Testing Transforms Cyber Security

AI penetration testing transforms annual compliance snapshots into continuous security assurance without sacrificing the depth of manual expert testing.

Cyber Locks

Penetration testing (pentesting) is a simulated cyberattack conducted by security professionals to identify and prioritize vulnerabilities in your systems, applications, or networks that can be exploited -- before a real attacker finds them first. Unlike automated scanners that generate lists of potential issues, penetration testing validates exploitability with evidence and proof of exactly how an attacker would get in, what they would access, and what it would take to stop them. Penetration testing follows a structured process governed by internationally recognized frameworks, including the Penetration Testing Execution Standard (PTES) and OWASP Testing Guide.

AI is fundamentally changing what "continuous security assurance" looks like through AI pentesting in 2026. 

Before any testing, the pentester and client must define the rules of engagement, including which systems are in scope, what testing methods are permitted, and what constitutes a "safe" level of disruption. This phase also covers legal documentation (authorization letters, NDAs) and defines what success looks like.

Here are the recommended steps:

A Pentester Initial Check Box

Clients choose among three testing postures:

  • Black Box: Tester has no prior knowledge of the environment (simulates an external attacker with no insider information)
  • White Box: Tester has full access to source code, architecture diagrams, and credentials (deepest coverage, fastest to execute)
  • Gray Box: Tester has partial knowledge - typically a standard user account (simulates an insider threat or compromised credential scenario)
Mapping the Attack Surface

The tester maps the attack surface using passive and active techniques:

  • Passive reconnaissance: This includes OSINT (Open Source Intelligence), DNS enumeration, WHOIS lookups, LinkedIn scraping for employee names and technology stack clues - all without touching the target system directly.
  • Active reconnaissance: Common methods are port scanning (Nmap), service enumeration, web crawling, banner grabbing. The output is an inventory of exposed systems, services, technologies, and potential entry points.
Threat Modeling

Not all vulnerabilities are equally dangerous. Threat modeling is where the tester (or in AI-powered pentesting, the reasoning engine) evaluates which discovered entry points represent the highest risk given the specific business context. This is where context matters. An SQL injection vulnerability in a payment processing endpoint is materially more dangerous than the same vulnerability in a public-facing blog comment form. Traditional scanners assign the same CVSS score to all vulnerabilities. A skilled pentester (or a context-aware AI agent) weighs them correctly.

Vulnerability Analysis

With reconnaissance complete and attack paths prioritized, the tester performs systematic vulnerability analysis. This includes:

  • Automated scanning (Nmap, Nikto, OpenVAS) to baseline known CVEs
  • Manual analysis to identify business logic flaws that scanners miss - authentication bypasses, insecure direct object references, race conditions
  • OWASP Top 10 coverage for web applications - injection attacks, broken authentication, sensitive data exposure, security misconfigurations, and more

The key distinction between vulnerability analysis and exploitation is that analysis identifies potential weaknesses. The next step is to determine whether those weaknesses can actually be leveraged.

Exploitation
  • In this step, the tester actively attempts to exploit identified vulnerabilities to prove their impact. This includes:
  • SQL injection to extract database contents or bypass authentication
  • Cross-Site Scripting (XSS) to hijack user sessions
  • Privilege escalation to move from a standard user account to an administrator account
  • Chaining vulnerabilities by combining multiple low-severity issues into a critical attack path that neither issue would represent individually
Post-Exploitation and Lateral Movement

Once initial access is achieved, the tester assesses how far an attacker could realistically go. Questions to be addressed include:

  • Can they move laterally to other systems on the same network?
  • Can they escalate to domain administrator or cloud root access?
  • What sensitive data (PII, credentials, financial records) could they exfiltrate?
  • How long could they maintain persistence without triggering detection?

This phase answers the question your C-suite will ask after a breach: "How bad could it have been?"

Reporting, Remediation Guidance, and Retesting

The final deliverable is what separates a useful penetration test from an expensive PDF. This last point matters more than most teams realize. Paying for a pentest and a separate retest engagement is the standard model. It is also where AI-powered penetration testing changes the economics since retest runs become instant, not billed separately.

Expected results from a solid penetration test report include:

  • Executive summary: Business-language explanation of risk severity and top findings for the CISO and board
  • Technical findings: Vulnerability details with CVSS scores, evidence screenshots, and attack chain diagrams
  • Reproducible proof-of-concept steps: Exact steps your team can follow to confirm the vulnerability before fixing it
  • Remediation guidance: Specific, actionable fix recommendations - not "update your software" but "apply patch CVE-2025-XXXX to Apache 2.4.x and rotate the following credentials."
  • Retest confirmation: A follow-up assessment to verify that remediations actually closed the vulnerability
AI Penetration Testing

Traditional penetration testing forces a choice: you can have depth (manual testing by skilled humans) or frequency (automated scanning run continuously). You cannot have both - not at a cost that scales. AI-powered penetration testing changes the underlying economics. An autonomous AI agent can:

  • Map an attack surface and enumerate vulnerabilities without human supervision.
  • Adapt its attack logic in real time based on how the application responds - mimicking the reasoning of a human ethical hacker rather than following a static script.
  • Validate exploitability with safe proof-of-concept execution.
  • Deliver remediation guidance in a developer-ready format immediately after the test completes.

The result is the equivalent of a week or more of manual penetration testing, delivered in hours and available on demand.

What Makes an AI Pentest Agent Different from a Scanner

A vulnerability scanner applies pattern matching. It looks for known CVE signatures, compares version numbers against databases, and flags anything that matches a rule. It is deterministic and static.

An AI penetration testing agent applies adaptive reasoning. It observes how the application responds to an input, infers what that response suggests about the underlying architecture, and adjusts its next action accordingly. It can:

  • Notice that a 500 error on a specific input suggests a backend database query is being passed as user input, and pivot to SQL injection testing.
  • Recognize that a redirect loop suggests a flawed authentication state machine, and attempt to exploit the race condition.
  • Chain a low-severity information disclosure finding with a medium-severity IDOR vulnerability to demonstrate a critical data exfiltration path.

This is the difference between automation (doing the same thing faster) and autonomy (reasoning and adapting independently).

AI Pentesting for Continuous Security Assurance

With an AI agent that can run a full assessment in hours, security teams can:

  • Test every significant release before it reaches production
  • Re-validate remediations immediately after they are deployed (instead of waiting for the next engagement to confirm a fix actually worked)
  • Run targeted retests after CVE disclosures that may affect your tech stack
  • Build a longitudinal trend view of your security posture over time, not just a point-in-time snapshot

AI-powered penetration testing replaces annual compliance with continuous security. The most transformative application of AI penetration testing is not replacing the annual manual engagement - it is enabling continuous assurance between those engagements.


Sumedh Barde

Profile picture for user SumedhBarde

Sumedh Barde

Sumedh Barde is chief product officer at Simbian, a provider of autonomous AI agents. 

Prior to Simbian, he was head of product for Microsoft's cloud data security products.  He also previously held a position as director of security programs at Meta. 

Barde obtained his B.Tech in computer science and engineering from IIT Bombay.

The Onset of 'Death by AI' Claims

Gartner projects that there will be at least 2,000 legal claims of "death by AI" this year, as the complexities of AI adoption move to a new phase. 

Image
AI Robot Hand with Legal Image

The insurance industry can take pride in the fact that innovation can't happen without it. Until innovators and their insurers figure out how to defray the risk from driverless cars, commercial space flight, etc., they can't go to market. But innovation also can't happen without lawyers. While we non-lawyers complain about how they slow things down, innovations can't scale until the legal system develops a framework for adjudicating the inevitable problems. 

Generative AI is moving into its early legal phase, according to a report from Gartner Group. The report predicts that by the end of the year there will be more than 2,000 legal claims worldwide related to "death by AI," as mistakes by the software or by those implementing it may be the root cause of fatalities. 

The implications will be most immediate for health insurers but will be felt soon enough in just about every corner of the insurance industry, especially where AI is being used to try to anticipate and prevent losses.

Let's have a look.

Gartner frames the "death by AI" issue as a broad one for companies in all industries, suggesting that general counsels need to be aware of the risks and need to work with insurers to purchase coverage. Gartner predicts that by 2030 there will be a 60% increased in corporate spending on security and governance related to AI. From that standpoint, AI looks like a big, new opportunity for insurers.

I'm more concerned about the potential surprises that may be waiting for insurers. 

Those insuring medical practices, for instance, may be caught by surprise if the caretakers turn tasks over to AI that then go awry. Human doctors are still very much in the loop at the moment, but there's a real push toward instituting a combination of telemedicine and automated AI advice, especially to reach people who live in remote areas or other "healthcare deserts." So decisions will real consequences may start moving quickly into the AI. 

The theory is great. You outfit people with wearables that monitor their health, alerting doctors of any warning signs. You coach people on eating, sleeping, exercise and so on. Doctors are reachable by Zoom for consultation and diagnosis. 

But what happens when the AI misses the signs of an impending stroke? What happens when it misdiagnoses a diabetic? 

A columnist in the Washington Post recently wrote about an experiment in Utah that raises all of these questions. It's a very responsible test, limited to having AI refill prescriptions, and could have major benefits. The columnist, an MD and former health commissioner in Baltimore, writes: 

"Right now, getting a prescription refilled can be challenging. Many patients call a doctor’s office and struggle to reach the right person or are told it’s not possible without an in-person visit, which requires time and travel. Some end up putting off that visit and go without medications, which can be dangerous for those with chronic diseases such as hypertension, diabetes and cardiovascular issues."

But she also quotes a professor at Harvard Medical School who says that, "while some drugs might appear to be low-risk on paper, prescribing them is often complicated and patient-specific. He noted that many drugs require ongoing monitoring, including regular lab tests, attention to side effects and careful and nuanced discussions with patients. 'It’s not clear that AI is fully able to replicate that,' he said."

And I believe that people -- including those on juries -- hold machines to higher standards than they do humans. Humans can make errors in the heat of the moment. We know we aren't perfect. But software is written by very smart people who aren't under instant time pressure and are vetted by large, responsible organizations (with deep pockets). So AI can't just be good. It has to be perfect.

The potential for legal surprises won't just relate to "death by AI," either. There will also be "injury by AI," at a far greater rate. (While more than 40,000 people die in car accidents in the U.S. each year, for instance, some 2.5 million are injured.) 

And the claims won't just hit healthcare providers that may have misdiagnosed or mistreated someone. I worry about the companies that use AI to detect dangerous situations in workplaces. What happens when they miss one and someone is hurt or killed? What happens when sensors don't detect the electrical problem in a home that leads to a fire, or the leak that's about to become a flood? When the forward-looking dashcam doesn't spot the deer that has jumped into the road? 

As I've written, consumer advocates are already blaming the big, bad algorithm for any decisions they don't like on underwriting and claims. Those legal issues are about to broaden, especially for those promising prevention via AI.

We'll get through this. The legal framework will gradually develop, and we'll learn what the rules are going to be. But we need to brace ourselves for complications like the coming wave of "death by AI" claims.

Cheers,

Paul

 

The Critical Flaw in Insurance AI

Agentic AI exposes insurance's critical flaw: Insurers cannot consistently deliver decision-ready data when and where it matters.

Techy Image

AI in insurance is advancing, but it is not yet transforming the industry. We are moving beyond systems that analyze and recommend and toward ones that can act, by initiating claims workflows, flagging fraud in real time, adjusting underwriting decisions, and orchestrating next-best actions. This shift toward agentic AI is often described as a turning point, and it is, but not for the reasons most narratives suggest.

While the technology is evolving rapidly, most insurers remain constrained by a more fundamental issue: they cannot consistently deliver the right data, at the right time, in the right context to support real-world decisions. Until that changes, autonomy will remain limited, no matter how advanced the models become.

Most insurers are not lacking data or platforms. Over the past decade, they have invested heavily in data lakes and lake houses, advanced analytics and AI tools, and integration and data engineering pipelines, yet progress beyond pilots remains slow.

The problem is not access to data. It is making that data usable, trusted, and actionable at the moment a decision is made.

In insurance, this challenge is amplified by fragmented policy, claims, and customer systems, dependence on third-party data such as telematics, weather, credit, and health data, regulatory and compliance constraints, and the need for real-time decision-making in customer-facing processes.

Agentic AI does not solve this problem. It exposes it.

Why a Shared Data Layer is Not Enough

Many organizations respond by building a shared data foundation — a unified layer where humans and AI agents can access the same information. While this is directionally right, it is incomplete. The challenge is not that organizations lack a shared data layer; it is that they struggle to deliver the right version of data for each decision, at the moment it matters.

Insurance operates on multiple, decision-specific views of data, each with distinct requirements:

  • Claims decisions depend on real-time, enriched incident data
  • Underwriting relies on forward-looking risk models and external signals
  • Fraud detection requires cross-entity patterns and behavioral analysis
  • Customer servicing depends on a simplified, current policyholder context

These are not variations of the same dataset, they are purpose-built representations of data, shaped by different latency, governance, and semantic needs, which becomes even more critical with agentic AI. Different agents operate at different points in the decision lifecycle, and require different data, in different forms, at different times.

A shared layer can provide access, but effective decisions depend on context.

From Data Access to Decision Activation

This is where many AI strategies stall. Most architectures are designed to store, process, and analyze data, but not to activate it at the point of decision. There is a fundamental gap between data being available and data being usable within real-time workflows.

Agentic AI operates directly in this gap. Without access to live, governed, and contextually aligned data, agents operate with partial understanding, and their outputs become unreliable. This is why many AI initiatives remain stuck in experimentation.

To move forward, insurers need to rethink how data is delivered. Not as raw datasets or reports but as data products — a reusable, governed, and outcome-aligned data asset designed to support a specific decision or workflow. Instead of exposing raw data, insurers should deliver contextualized, decision-ready views, with embedded governance and policy controls, consistent business semantics, and real-time access to internal and external sources.

For example:

  • A claims data product unifying FNOL, policy data, repair estimates, and external signals
  • A fraud data product combining claims history, network relationships, and behavioral indicators
  • An underwriting data product integrating internal risk data with third-party enrichment

These are not static datasets. They are dynamic, purpose-built representations of data, aligned to the decisions they support.

Why Real-Time, Governed Access Matters

For agentic AI to deliver value, data must be live, governed at access, semantically consistent, and traceable. This is where a logical data layer becomes critical, not just as an integration approach, but as a way to connect distributed data in real time, apply governance dynamically, and deliver consistent, business-ready views across systems. This enables both humans and AI agents to act with confidence, without introducing further fragmentation.

The insurers that lead in 2026 will not be those with the most advanced models. They will be the ones that connect AI directly to business outcomes. That means starting with the outcome, such as reducing claims cycle time, improving fraud detection, increasing underwriting precision, or enhancing customer experience, and working backwards to define the decisions, data and systems required to support them.

This is how AI moves from experimentation to operational impact.

Where AI Success is Won or Lost

The next turning point for AI in insurance will not come from smarter models. It will come when organizations accept a deeper truth; AI is only as effective as the data it can access, interpret, and act on, in real time.

Agentic AI accelerates this realization. It makes clear that data must be trusted, contextual, available at the moment of decision, and aligned to outcomes. Those who solve this will scale AI successfully, and those who do not will continue to pilot without transformation.

The future of insurance will not be defined by whether humans and AI agents share the same data. It will be defined by whether they have the right data, in the right form, to make the right decisions. That requires a shift from shared data to decision-ready data, from access to activation, and from experimentation to measurable outcomes. That is the real inflection point for AI in insurance.


Errol Rodericks

Profile picture for user ErrolRodericks

Errol Rodericks

Errol Rodericks is director of product marketing for EMEA and LATAM and global solutions director for vertical industries at Denodo.

He previously held leadership roles at Boomi, ServiceNow, HP, CA Technologies, and IBM. He founded Technology Concepts.

Rodericks holds an MSc in digital systems from the University of Wales, Cardiff, and a BSc (hons) in electronics and communications engineering from the University of North London.