7 Takeaways from the Big Data and Artificial Intelligence Working Group at the NAIC 2026 Summer National Meeting

Every seat was taken when the Big Data and Artificial Intelligence Working Group convened at the NAIC’s 2026 Summer National Meeting in Columbus. Attendees stood along the walls and filled the back of the room, reflecting how quickly AI has become an operating question for insurers, regulators, and technology companies.

We attended the session, which combined an update on the NAIC’s AI Risk Evaluation Supplement pilot with a presentation from Edin Imsirovic, Director at AM Best, on governance across predictive, generative, and agentic systems. The discussion was wide-ranging, but seven takeaways stood out.

1. The NAIC’s AI Risk Evaluation Supplement is heading toward public exposure

The working group opened with an update on the AI Risk Evaluation Supplement pilot, which has involved 12 states since March. Participating regulators have met nearly every week to compare their experiences.

The pilot has not followed one uniform process. Some states incorporated the supplement into planned market conduct or financial examinations. Others used it as an ad hoc questionnaire. States could modify or add questions based on their needs.

The NAIC plans to expose a revised version for comment in early September. The published timeline anticipates two exposure periods, additional company feedback, and possible adoption at the Fall National Meeting. The process remains a pilot, and the supplement is not a product certification, rating system, or new insurance law.

Its new name matters. Earlier versions were called the AI Systems Evaluation Tool. Calling it a risk evaluation supplement more accurately describes its intended role: giving regulators a structured way to ask how an insurer uses and oversees AI.

2. AM Best placed AI inside existing innovation and ERM questions

Imsirovic was explicit that his presentation did not change AM Best criteria, methodology, or rating guidance, or create a formal AI governance assessment framework.

AI already fits within two areas AM Best uses to understand insurers: innovation and enterprise risk management. The innovation lens considers leadership, culture, resources, processes, and measurable operating results. The ERM lens asks whether risk-management capability remains commensurate with the risk profile being created.

That framing avoids treating AI as an isolated technology category. A deployment that improves claims efficiency or underwriting quality may also create operational, regulatory, cyber, data, or third-party risk. Both outcomes can develop at the same time.

The relevant question is not whether an insurer can describe an impressive AI capability. It is whether the capability produces durable operating results while governance keeps pace.

3. Predictive, generative, and agentic AI create different control problems

The AM Best presentation separated AI into three capability classes.

Predictive AI scores, ranks, classifies, or flags information. Examples include pricing models, fraud scoring, and claims triage. Its governance questions concern validation, fairness, documentation, drift, and performance.

Generative AI creates content from documents, narratives, images, or other information. It can summarize claim notes, extract details, or draft language. Its control questions include hallucinations, prompt changes, sensitive information, and vendor updates.

Agentic AI can use tools and work through multiple steps toward a goal. It may retrieve information, update systems, or route work. Governance therefore has to address permissions, task boundaries, action logs, escalation, rollback, and the ability to stop the system.

The categories are not cleanly separated. One insurance workflow may combine all three, and later capabilities retain many of the risks associated with earlier ones.

4. The use case matters more than the AI label

One of the clearest messages from the session was that novelty alone does not determine risk. A predictive model is not automatically low-risk, and an agentic system is not automatically unacceptable.

Imsirovic offered five factors for assessing a particular deployment: the importance of the decision, the system’s autonomy, the insurer’s ability to explain an outcome, the speed at which behavior can change, and dependence on third parties.

The claims examples made the distinction concrete. A system preparing a claim summary for adjuster review may affect an important consumer outcome, but it remains assistive. A system routing and paying claims has both high impact and delegated action, placing it in a much more demanding governance category.

This is a better way to evaluate insurance AI than starting with a model name. Insurers need to document which decisions a system touches, how much it can do on its own, and what happens when it produces an unexpected result.

5. Written policies are not the same as operational evidence

The presentation drew a useful distinction: documentary governance answers with policies, while operational governance answers with evidence.

Monitoring should lead to observable action. If a predictive model drifts, the insurer might change a threshold, limit its use, or add review. If a generative AI vendor updates a model, the insurer should know whether the application was retested. If an agentic system crosses a boundary, the insurer should be able to reconstruct the action, stop it, and return the process to a safe state.

Human review also needs to be substantive. Reviewers need enough information, time, and authority to question the output. A company should be able to show where review changed an outcome, triggered escalation, or stopped a process. If every review ends in approval, it may be little more than a procedural click.

Audience members pressed this point during the discussion. They asked whether insurers are producing written procedures, governance schematics, and standardized documentation, or primarily describing their controls in meetings. Imsirovic cautioned that his view was limited, but said that much of what he currently hears comes through conversation rather than a formal collection of AI governance documentation.

6. The operating model may be the industry’s biggest weakness

When asked where insurers struggle most, Imsirovic identified the operating model around AI. He sees companies using multiple tools without necessarily redesigning the workflow, accountability, and controls surrounding them.

The presentation contrasted controlled integration with fragile tool layering. In controlled integration, the workflow, governance, and accountability are redesigned together. The organization defines what decision the tool touches, where a person intervenes, who owns the outcome, and what evidence the process retains.

Fragile tool layering occurs when AI is placed on top of an unchanged process. A pilot works, usage spreads, and the organization realizes later that the workflow and governance never caught up. The tool may still save time, but the value remains isolated and the risks become harder to see.

This is not only a compliance problem. Imsirovic said the same weakness can reduce the benefit an insurer receives from the technology. Governance that sits inside the operating process helps value and controls mature together.

7. Agentic AI expands governance beyond the model itself

Agentic insurance deployments remain early. Imsirovic said the implementations he currently encounters tend to involve lower-stakes work such as research, FAQs, or chatbots, and he expects the conversation to look different within the next year.

Audience members compared an agentic system to a nonhuman employee that requires supervision, defined authority, contingency planning, and a tested shutdown process. They also raised system-level failures that could affect the enterprise rather than one consumer interaction.

For these systems, a model inventory is no longer enough. Insurers may also need to inventory the tools a system can call, the information it can access, the permissions it holds, the actions it can take, and any memory it carries between tasks. Vendor contracts, audit rights, update notices, portability, and shared infrastructure become part of the governance perimeter.

The control boundary has expanded from what the model says to what the full deployment can do.

What these takeaways mean for claims correspondence

At Voltaire, the session reinforced the value of a bounded, observable claims use case. Claims correspondence can be evaluated through drafting and completion time, policy-language assembly, citation and formatting, reviewer corrections, rework, and escalations. Voltaire reduces repeatable drafting work within the carrier’s claims workflow, while the coverage position and approval remain within that process. The operating outcomes are faster letter completion, more consistent correspondence, and less manual copy, paste, formatting, and review work.

The public Claims Correspondence Compendium and its Columbus, Ohio claims correspondence reference show why this work is not uniform across jurisdictions. Claims leaders should map one live correspondence workflow, define the system’s role, and measure whether it improves completion and review without adding operational ambiguity. Teams ready to examine that bounded approach can explore Voltaire’s AI claims letter solution.

The broader message from Columbus was straightforward. Insurers will not get durable value from AI by buying a model and surrounding it with policy documents. They will get it by redesigning the work, defining authority, preserving evidence, and making governance part of the operating process itself.

More from this Author