The Unknowns of Enterprise AI Deployment

Property & casualty insurers face systemic unknowns when scaling AI beyond pilots into regulated workflows like underwriting, claims and pricing.

Deployment

Property & casualty insurers are moving fast from narrow machine learning pilots to enterprise-scale deployments that blend predictive models, generative AI, and agentic workflow automation. The hardest barriers to this transition are not primarily technical. They are the unknowns: the uncertain, interdependent, and often non-obvious failure modes that surface when AI systems get embedded in regulated, long-tailed, and economically sensitive insurance processes like underwriting, pricing, claims, reserving, and reinsurance.

Insurance executives must shift from model-centric thinking to system-centric thinking. Strong data and model controls are necessary but not sufficient. Without secure integration patterns, operational monitoring, model risk discipline, and clear accountability, even a high-performing model will struggle to become a safe, compliant, and profitable production system. Enterprise AI risk is not merely model risk. It is systemic risk arising from the coupling of data, models, workflows, humans, vendors, and core platforms.

Why P&C is a special environment

P&C is uniquely difficult territory for enterprise AI. The product is a promise made under uncertainty, and the balance sheet carries long-tail obligations, so decisions made or supported by AI can influence loss emergence years later through selection effects, reserving assumptions, and litigation pathways. This drives an unusually high cost of model error and governance failure. Several structural features amplify the unknowns: exposure to catastrophe clustering and tail events, rapid changes in external cost drivers like repair and medical inflation, the potential for proxy discrimination through correlated variables, complex multi-party ecosystems spanning brokers, MGAs, TPAs, and repair networks, and the fragmented reality of U.S. state-based regulation.

A taxonomy of unknowns

The key is a structured taxonomy that classifies the unknowns into eight categories, each with concrete P&C examples and matching guardrails. These span data unknowns (coverage gaps, inconsistent cause-of-loss codes, third-party data drift), model behavior unknowns (overfitting to recent inflation, LLM hallucination, proxy discrimination), system integration unknowns (automation triggering payments without adequate checks, silent integration failures), operational unknowns (drift during catastrophe season, retraining backlogs), security unknowns (prompt injection, data exfiltration, model theft), regulatory unknowns (varied state DOI expectations, market conduct exam demands), economic unknowns (unclear ROI, behavioral feedback loops, non-linear computing costs), and human and organizational unknowns (overreliance on models, adjuster workarounds, incentive misalignment). This taxonomy works both as an executive checklist for risk identification and as a technical planning artifact for control design.

Two unknowns receive special emphasis. The first is the feedback loop problem: when AI is used to price, select, investigate, or settle, it changes the composition of the book and the behavior of insureds and internal teams, which in turn alters the future data the AI is trained on. A model may appear to improve loss ratio in the short term while quietly increasing adverse selection, litigation frequency, or churn over longer horizons. The second is the tail and regime shift problem: since P&C risk is dominated by tails, models trained in routine years can fail under catastrophe clustering or new social inflation regimes, and validation that optimizes average error will systematically miss tail risk.

Mapping unknowns to governance frameworks

Rather than inventing a new compliance regime, the unknowns should be mapped onto established frameworks to create a shared language across technology, business, and regulators. The NIST AI Risk Management Framework serves as the backbone, with its four functions of Govern, Map, Measure, and Manage. This is complemented by the NAIC's 2020 AI Principles and its December 2023 Model Bulletin on the Use of AI Systems by Insurers, which set regulatory expectations for governance, documentation, and oversight, including for vendor-acquired systems, and stress compliance with existing unfair trade practice and unfair discrimination laws. Also important are NIST CSF 2.0 and the NIST Privacy Framework for cyber and privacy integration, the NAIC Insurance Data Security Model Law for data security standards, and SR 11-7 / SR 26-2 model risk management discipline adapted from banking. The Colorado AI Act also offers guidance on what the future of AI regulation in the industry looks like.

A control mapping table connects specific unknowns to control objectives, framework hooks, and evidence artifacts, and the highest-leverage leadership move is treating evidence as a product: every AI system should ship with documentation, test results, monitoring plans, and audit-ready logs.

Reference architecture and generative AI patterns

A six-layer reference architecture is needed for the heterogeneous, hybrid environments that carriers actually run: business process and orchestration, AI application, model, data and feature, platform and operations, and a cross-cutting security, privacy, and governance layer. MLOps must be treated as first-class production engineering rather than project-based delivery, because the true cost of AI is dominated by post-deployment work like monitoring, incident response, recalibration, and security patching. The main generative AI patterns in production insurance settings are: retrieval-augmented generation grounded in policy forms and claims manuals, constrained tool use, targeted fine-tuning, and agentic workflows with supervisory layers and circuit breakers for financial actions. Security by design extends existing controls while adding AI-specific safeguards drawn from OWASP's LLM vulnerability taxonomy and MITRE ATLAS.

Tiered guardrails and continuous assurance

A key insight is that not all use cases warrant the same governance intensity. A three-tier model based on decision impact is needed. Tier 1 (informational: search, summarization, document classification) allows advisory-only outputs with basic guardrails. Tier 2 (decision support: underwriting triage, pricing indications, claims severity and fraud scores) requires segmented validation, explainability, fairness tests, and human-in-the-loop review. Tier 3 (acting and automation: auto-routing claims, automated payments within limits) demands dual control for payments, transaction limits, rollback, and continuing monitoring. This tiering supports proportional governance without over-controlling low-risk productivity use cases. 

On measurement, there should be a shift from one-time validation to continuous assurance, combining pre-deployment testing (data readiness, model validation, fairness and security tests, documentation), post-deployment monitoring across technical, business, compliance, and security signals, AI-specific incident management, and exam readiness.

Operating model, roadmap, and research agenda

Deployments fail when accountability is unclear, and it challenges a common myth: that accountability for all AI initiatives should sit with the CTO, CIO, CDAO, or CAIO. This contradicts basic operational risk principles, and accountability should instead be defined by roles, responsibilities, and use case within functional areas. What is needed is a governance structure (an AI Steering Committee, an AI Risk and Controls Council, product owners with a value realization office, and an independent validation function) and a detailed table of CXO responsibilities and the specific unknowns each leader must own.

Avoid the trap of scaling models before scaling controls, and instead sequence the work: baselines in the first 90 days, enterprise repeatability at three to six months, and institutionalization at six to 18 months, including alignment to ISO/IEC 42001 and 27001. Set a research agenda covering causality and feedback loops, fairness under distribution shift, tail-risk stress testing, generative AI assurance, AI security metrics, and standardized evidence for regulators. Practical appendices provide an executive checklist mapped to NIST AI RMF, a model card outline, AI risk register fields, and an incident response playbook.

To read the full paper this article is drawn from, click HERE.


Kushal Shah

Profile picture for user KushalShah

Kushal Shah

Kushal M. Shah has 25 years of experience in insurance industry.

He is the author of "The Unknowns of Enterprise AI in Regulated Sectors" and innovator of patent-pending aiV-Cube framework for AI risk assessment and underwriting. 

He holds active producer and adjuster licenses across multiple states and has completed the Associate in Claims from The Institutes. He is currently a candidate for Associate in Insurance AI (AIAI).

Read More