The industry no longer doubts whether AI benefits submission intake in commercial property and casualty (P&C) lines. Recent AI rollouts by Zurich, AIG, Markel, and other market leaders show impressive gains come very quickly: intake timelines for complex submissions compress from hours to minutes, straight-through processing (STP) rates increase from 10% to 95%, underwriter productivity grows by 100%+, and submit-to-bind ratios improve by 35%.
Yet many P&C carriers are still figuring out how to make AI work for their business. Which parts of their submission intake workflow can realistically be automated? Where to start so that we can move out of pilots fast? How can we explain, audit, monitor, and defend what the AI did? And what AI solution design should we actually pursue to scale well?
These are the questions I hear most often in ScienceSoft's engagements with commercial P&C clients. In this article, I'll share my perspective on where AI works best in submission intake today, why many initiatives never move beyond pilots, and what insurers should focus on if they want to launch submission intake AI over the next six to 12 months.
Where AI Delivers the Fastest Value in Submission Intake
From my experience, the quickest wins come from automating the parts of insurance submission intake that are high-volume, repetitive, document-intensive, and largely governed by business rules. Strong starting points in commercial P&C include:
- Capturing and classifying submissions.
- Extracting data from ACORD forms, loss runs, schedules of values, and other submission documents.
- Checking submissions for completeness and identifying missing information.
- Summarizing and triaging risks for underwriters.
- Prefilling policy administration and underwriting systems.
- Drafting underwriting files and follow-up questions for brokers.
These activities create a large operational workload but require little underwriting judgment, which makes them ideal candidates for AI automation. Modern AI systems can handle these connected intake operations almost entirely, only involving humans in complex cases. In this first half of the intake pipeline, STP rates of 90–95% are a realistic target.
They are also the safest processes to automate from both a business and regulatory perspective. In each case, AI prepares data and doesn't touch high-impact decision-making.

Anything involving decision-making is a different story. Commercial P&C underwriting is rarely standardized. It often requires negotiation, exceptions, and collaborative expert judgment, with every submission carrying its own nuances. Most of the cases are just too complex for reliable straight-through AI decision-making.
That doesn't mean AI has no role here. It can absolutely assist in eligibility assessment, pricing, and coverage decisions — for example, by highlighting relevant risk factors, surfacing similar cases, or recommending next steps — but I'd keep it away from making those decisions autonomously. Human experts should continue to own the final judgment.
Moreover, even when AI acts only as an assistant, the governance bar remains high. You need mature controls, explainability, and human oversight to defend AI suggestions to brokers, policyholders, auditors, and regulators. That's why I usually recommend that clients avoid complex assistive use cases during early AI deployments. A safer path is to first confirm AI accuracy, auditability, and workflow impact across lower-risk data intake tasks, then gradually scale into underwriting decision-support areas.
Overall, I don't think anything close to fully autonomous underwriting should be the near-term objective for commercial P&C insurers. The "AI assists, humans decide" approach, where AI prepares files and experienced underwriters make the final decisions, has proven the most practical operating model. It lets insurers preserve human accountability where it matters most while still delivering huge business returns. I recently came across a case study where deploying AI for submission intake routines alone brought a 646% ROI for a large property carrier through faster, more accurate, and more efficient data processing.
Data Foundations for AI-Powered Submission Intake
First of all, you can effectively start with whatever data you have. That's an important point because I've seen many insurers unnecessarily delay AI initiatives out of fear that their existing data is too scarce or too low-quality for AI processing.
The first necessary data foundation you need in commercial submission intake is a clear document taxonomy. However strong AI may be at data tasks, it still struggles with unstructured, consequential data containing insurance-specific terminology and context. One classic example is loss runs: they vary by provider. Claim descriptions often contain abbreviations. Severity indicators may be buried in narrative text. The AI may read the document correctly but fail to determine what's material from an underwriting perspective.
A clear taxonomy creates enough structure around the data so AI can distinguish between document types and apply the right extraction logic, validation rules, and review requirements to each. It also helps determine which documents AI should treat as authoritative when data conflicts across sources.
What Might Keep AI Stuck in Pilot Mode, And How You Avoid It
By far the biggest blocker is AI integration into existing submission intake workflows.
Most commercial P&C insurers operate heavily fragmented environments. Documents sit in one place, policy data in another, rating logic somewhere else, and underwriting notes in spreadsheets or workbenches. People have built manual workarounds over the years because their legacy core systems do not support the connected workflow.
Making AI work with these core systems doesn't mean you need to modernize or replace all of them. But you do need a practical integration pattern to avoid building numerous point connections. Off-the-shelf AI architectures often hit a wall at this point: most mass-market tools, by design, introduce AI as another isolated interface sitting alongside the existing process. AI product vendors can't possibly account for every system or data format their clients may still be using to build a product that integrates well with legacy stacks.
Another blocker is what I call "pilot thinking." During pilots, people often focus too much on proving that the AI model can accurately classify, extract, summarize, and answer questions. That's necessary, but it's only one part of the equation. Moving to production requires proving that the AI can operate reliably inside a real insurance workflow: integrate with existing systems, handle exceptions, support human review, produce audit evidence, and scale across thousands of submissions. That's where many otherwise successful pilots stall.
The goal of a pilot should be validating not just AI accuracy but rather the operating model around the AI. Can the workflow route low-confidence cases for review? Can underwriters easily verify and override AI outputs? Can the solution recover from missing or conflicting information? Can it integrate with core systems without creating manual workarounds? Those capabilities determine whether the AI can become part of daily operations.
The third blocker is poor ownership models. Someone has to review AI performance, define business rules, validate compliance, and measure outcomes. When ownership is unclear, issues randomly fall between teams. The technology may work, but if nobody takes responsibility for adoption, governance, and continuing improvement, your promising AI initiative may quickly lose momentum after the pilot stage.
You'll need a board of AI governance owners from underwriting, IT, data or AI engineering, and compliance or risk before moving beyond the pilot. The board should regularly review AI performance, override rates, exceptions, user feedback, and regulatory risks, and decide when the solution is mature enough to expand into new workflows or business lines.
Addressing AI Risks That Insurers Often Underestimate
A lot of insurers immediately think about AI hallucinations, and that's a real risk.
To address that risk, the AI reasoning should be grounded in source documents, and the system must show where each important field or summary statement came from. Adding deterministic validation of AI outputs after each processing iteration also helps prevent error creep. These are rule-based checks that verify required fields are present, figures are consistent, references match the source documents, and the output complies with predefined business rules before the workflow continues.
Yet, in commercial P&C submission intake, I think one of the most overlooked risks is silent workflow bias.
By that, I mean AI may not be making the final underwriting decision, but it may still influence which submissions move faster, which are routed to senior underwriters, which are treated as gapped, and which receive follow-up. Those workflow decisions can create different outcomes over time, even if the model never explicitly uses protected-class data.
Another risk is silent portfolio drift. The process gets faster, productivity looks better, and everyone is happy. But over time, the mix of business may change. Maybe more borderline risks get through because the process feels smoother. Maybe underwriters stop asking certain follow-up questions because the AI-produced summary looks complete. None of this looks like a major failure on day one, but it shows up later as leakage, adverse selection, and portfolio quality issues.
The way to manage this is through continuing outcome monitoring. During pre-launch tests, compare AI-assisted cases versus manually processed ones to establish execution benchmarks. After rollout, monitor trends in quote-to-bind rates, referral rates, missing-data rates, post-bind corrections, override rates, and downstream loss performance. Also regularly sample fast-tracked cases for expert review and ask underwriters: would we have handled the submission the same way without AI? The goal is to detect when AI begins influencing portfolio quality in unintended ways.
Regulatory compliance is a known source of risk, and it requires governance from the start. In the US, regulators increasingly look beyond underwriting outcomes and examine whether AI affected submission routing, triaging, eligibility assessment, and broker interactions. The NAIC Model Bulletin on AI, adopted by 24 states and the District of Columbia, calls for controls against AI discrimination. New York, California, and Connecticut have issued their own AI guidance for insurers. You need explainable AI logic, audit trails, mandatory human reviews for high-impact decisions, and evidence that you regularly monitor for bias and drift to withstand regulatory scrutiny.
An Actual AI Solution for Commercial P&C Submission Intake
In an end-to-end commercial submission intake scenario, we would be looking at a multi-agent system coordinated by an orchestrator. We need specialized AI agents to perform different tasks: one classifies incoming documents, another extracts and validates data, the third drafts broker follow-up requests, and so on.
ScienceSoft prefers this pattern because submission processing involves many distinct activities with different accuracy requirements, permissions, and risk profiles. An input classification agent doesn't need access to the same data as a risk file preparation agent. An agent responsible for drafting emails shouldn't be reasoning on eligibility. Separating responsibilities reduces the blast radius of errors and makes governance much simpler.
And then, the orchestrator acts as a control layer of the agentic workflow. Unlike narrow agents, it doesn't perform submission-intake tasks itself. Its only job is to coordinate how work moves between AI agents, business rules, systems, and people. It decides which agent should act next, applies predefined routing rules, manages confidence thresholds, sends exceptions to human reviewers, and maintains an audit trail across the entire process. Think of orchestration as an AI traffic controller. Without it, you have just a set of AI capabilities.
Must-Haves of a Production-Ready AI Architecture
Three engineering principles separate a production-ready agentic AI system from a one-off pilot: decoupling AI from core insurance systems, implementing a multi-level AI authorization model, and building agents in a modular way.
Avoiding tight coupling between the AI workflow and the insurer's existing systems is essential for interoperability. Otherwise, adding AI would require rework across existing systems, and even small changes to those systems would trigger changes to AI-supported operations. In an ideal setup, the orchestrator and AI agents act as a back-end operational layer between your current systems without replacing them or requiring expansion. A layered architecture with a dedicated agentic layer and an integration layer sitting between the AI workflow and other insurance systems works well for that.

Consider applying event-driven integration patterns, where business events (think submission arrival or document upload) automatically trigger the next AI task. This keeps workflows synchronized across multiple systems without creating tightly coupled point-to-point integrations and allows AI to react immediately as an event occurs, making submission intake faster. Another major advantage is easier integration with legacy systems. Older apps that don't support APIs can often participate in the workflow by sending or receiving event messages, removing the need for custom integrations.
One more practical move is to expose the integration layer through a single gateway. This way, you don't need to integrate every AI agent separately with your existing systems and get a single place to capture audit logs. This approach simplifies integration across fragmented environments and supports a consistent audit trail required for compliance. It also minimizes integration maintenance overhead: if you later change your core platforms, the AI integration contracts remain stable.
The second principle is multi-level AI authorization. This approach aims to limit AI autonomy where business and regulatory risks are high while maximizing overall automation. We typically use three authorization levels. The first level allows fully automated actions for low-risk tasks like document classification or completeness checks, provided the AI outputs pass predefined deterministic checks. The second allows AI recommendations but requires human approval before execution, for example, when drafting broker follow-ups or preparing underwriting summaries. The third covers decisions that may affect eligibility, pricing, or other material underwriting outcomes. Here, AI prepares the supporting analysis and data, but humans remain the decision-makers.
You also need the system to log every action, including inputs, outputs, agent actions, source references, validations, and underwriter overrides. With this complete log, you can explain how decisions were reached, trace outputs back to evidence, review AI behavior long after a submission pack is processed, and prove regulatory-aligned controls during audits.
The third principle is modular agent design. With a modular architecture, each agent is built as an independent component. Such a design lets you add, improve, and replace task-specific agents without redesigning the whole system. This means you can first deploy agents for only a few processing tasks, one product, or one submission channel, prove accuracy and adoption, and then expand. That ability to scale gradually lets you start your AI journey with moderate upfront investment and with minimal implementation risk. If you're ready for a broader rollout from the outset, the modular architecture still makes the solution easier to maintain and evolve as business needs change.
The same modularity also allows you to use the best-performing and cheapest technology for each agent. For example, agents that assess eligibility and summarize risks benefit from the reasoning capabilities of large language models (LLMs). Implementing them with tailored retrieval-augmented generation (RAG) pipelines ensures outputs are grounded in actual submission documents. For document classification and entity extraction agents, fine-tuned encoder models trained on a fixed set of documents typically deliver very high accuracy at a fraction of the cost of LLMs. Plus, these models do not generate anything, so there's no hallucination risk at early intake steps.
If you're wondering what a workable automation scope, AI architecture, and data foundation would look like for your submission volume and current intake process, don't hesitate to reach out to discuss your case.
Contributing to this article were: Vadim Belski, head of AI, principal architect, ScienceSoft, andStacy Dubovik, financial technology & AI researcher, ScienceSoft.