The insurance industry has been arguing all summer over whether AI liability can be insured. Most carriers have decided - at least for now - that it can't. They are narrowing policy language on the GL, cyber, and product liability forms, adding AI exclusions, trying not to repeat the "silent cyber" problem from last decade. A smaller group of MGAs, Lloyd's coverholders, and reinsurer-backed programs has started writing affirmative AI liability coverage. Their approaches have little in common: performance warranties, adversarial testing, governance reviews, litigation-based models or just AI endorsements bolted onto existing cyber policies.
AI liability claims are still sparse, but the risk exposures are concrete: an autonomous agent that approves $50,000 in payments it shouldn't have, or an AI hiring tool that screens out protected classes. The potential buyers: any company deploying AI where errors hit third parties. Healthcare, finance, legal, HR screening.
These are different perils, different policy forms, different loss dynamics, and that's part of what makes measurement so hard. We're in the early innings of AI insurance. Some skeptics say you can't price it at all. I wouldn't go that far. But what are we actually trying to insure, and on what basis
Who's Writing AI Liability Today
The market is small enough that you can name the entire first cohort.
Munich Re launched the first dedicated AI insurance products, aiSure, as a performance warranty: if the model drifts below defined accuracy thresholds or produces discriminatory outputs, the policy pays. They've been at this since 2018 and remain one of the few reinsurers with a dedicated AI liability product in market.
At Lloyd's, three coverholders have launched AI liability products. Armilla AI offers standalone coverage, with underwriting informed by over 500 AI system evaluations. AIUC takes a different path: it certifies an AI model first, then adds insurance. Their AIUC-1 framework puts systems through thousands of adversarial simulations. The insurance is written on Beazley paper, with ElevenLabs, an AI voice generation company, as their first public customer. Testudo builds its underwriting off AI litigation data, though the dataset is still thin.
Corgi started offering an AI and algorithmic liability endorsement on top of D&O, E&O, and cyber. Cowbell added AI-specific underwriting factors to its cyber risk-rating framework in July. That's probably the near-term path for most cyber MGAs: AI bolted onto cyber, not a standalone line.
And then there's the rest of the market. Technology companies can buy standard Tech E&O from Vouch, Hartford, or Hiscox, with an AI endorsement added. No AI-specific risk assessment, no evaluation of model behavior. Premium is based on revenue, headcount, vertical, prior claims, the same way you'd price a SaaS platform or payroll tool.
A handful of players are experimenting with AI-specific underwriting. A much larger market isn't measuring AI risks at all. Total dedicated premium for AI liability: immaterial.
From Signal to Pricing
AI liability is one of the hardest emerging risks to insure.
Parametric insurance works because the trigger is a pre-agreed, observable, independently verifiable number: a NOAA weather station, or a cat model from Moody's RMS.
AI evaluation scores don't work that way. A hallucination rate changes with the test set, the evaluator, the model version, and the business context. Armilla and AIUC evaluate AI models, but when the evaluator is also the insurer, the data isn't independent anymore, by definition.
What underwriters need is a validated chain: an observable AI signal that correlates with loss frequency, that maps to expected severity, that can be modeled across a portfolio. I haven't seen a public demonstration of this chain from signal to pricing.
Gallagher Re laid this out in their June 2026 report "Anthropic's Fourth Way": current AI evaluation methods "were not designed for underwriting and are not fit for that purpose." Benchmarks measure how models perform on controlled tests. But losses happen in deployment, not in testing. Their conclusion:
"If a model cannot be tested, insurers end up pricing uncertainty rather than risk."
Without better evaluation, that leads to two failure paths: AI losses absorbed silently into existing policies until carriers exclude them, or standalone products launched without foundations that collapse after early losses.
Without independent model evaluation, the market defaults to underwriting governance: deployment approval processes, human oversight. Most AI underwriting is done that way today, but I'm not convinced that is enough. Cyber insurance tried governance-based underwriting for years, remember? Applicants checked "yes" on the MFA question, and half the time they didn't have it deployed properly. It took a catalyst, the ransomware wave of 2020/21, to force the industry to verify independently what applicants were telling them.
Where the Cyber Comparison Breaks Down
BitSight and SecurityScorecard built outside-in security scores for cyber, based on a client's open ports, SSL certificates, and malware infections. The same approach doesn't work for AI risk. A publicly visible chatbot might reveal something about prompt-injection resilience. It reveals nothing about the agent's authority, data flows, approval thresholds, or what happens when it makes a wrong decision. No external attack surface to scan.
There is a trust problem underneath all of this. Policyholders resist giving insurers access to internal data. They worry it might get used against them in a coverage dispute. And insurers have their own reasons not to look. I learned this at a cyber MGA: if you discover a vulnerability in a client's network through an internal vulnerability scan and don't act on it, you're exposed to E&O claims. Better not to know. For AI, it gets worse, as the vulnerabilities are harder to define, and there's no patch to deploy.
The Missing Layer
What the nascent AI insurance market needs is an independent measurement layer, something that takes technical AI evidence and turns it into data an underwriter can price from.
AI governance platforms like Holistic AI and Credo AI score AI systems on bias and robustness. They were built for compliance teams, not underwriters, and tell you whether a system meets a regulatory standard. They don't tell you the likelihood and cost of a liability loss.
Neither side can build this layer alone. The insured won't share data they fear could be used against them. The insurer faces the E&O problem I described. This layer needs to sit between them, the way a credit rating agency sits between borrower and lender.
AI insurance will need at least one credible, independent source of underwriting-grade evidence. Whoever builds that becomes infrastructure.
What would you trust as underwriting evidence for an AI agent: pre-deployment testing, runtime telemetry, an independent rating, a contractual warranty, or some combination?
