Who Governs the AI Agents?
Zywave CTO Doug Marquis on the Emerging Need for AI Oversight and Evaluation
Last month an AI agent built by OpenAI was handed a routine cybersecurity test. Instead of solving it, the agent found a flaw in its sandbox, escaped, and hacked into Hugging Face, a repository millions of people rely on for AI models and tools. This “unprecedented cyber incident,” as OpenAI called it, made a theoretical question all too real: When an AI agent acts on its own and something breaks, the answer to “who is accountable” isn’t complicated: It’s whoever deployed it.
As CTO, I expect vendors to take testing, evaluation, and guardrails seriously. But their diligence doesn’t transfer accountability. I select the vendor, the models, and the constraints. Yet my company owns decisions about where and how those agents operate. For insurers, meeting that responsibility means building a new kind of oversight that addresses how agents are supervised, what they’re allowed to do, and how to ensure they remain within bounds.
Defining AI governance
Oversight has to start with observability, which has become one of the most difficult issues in running agents. Traditional software was predictable. It followed a set sequence — A, B, C, D. If it stopped at C, you knew exactly where it broke, and a post-mortem was enough because nothing happened that you couldn’t trace and undo.
An agent removes that safety net. Because an agent decides its own steps, you can’t know in advance what path it will take. By the time the mistake shows up in the logs, the damage is often already done. Oversight has to happen differently: organizations need visibility into how an agent operates while it’s working: what information it uses, what systems it accesses, what actions it takes, and when a human needs to step in.
Observability answers the question of what an agent is doing. Governance addresses the question of what it should be allowed to do. In insurance, that distinction will become increasingly important as agents move from answering questions to taking action. An agent that helps a customer understand a policy has a very different role than one that recommends coverage changes, adjusts a claim, or influences an underwriting decision. The more an agent can affect a customer’s outcome, the more its boundaries and oversight matter.
Treat agents like employees, not software tools
That is why organizations should govern agents less like software and more like employees with defined responsibilities and limits. In our IT department, a developer assigned to one application has no standing access to the others. Agents should work the same way: access rights and restrictions tied to the task at hand with specific consequences if it gets it wrong.
In our R&D group, we trust agents to write and test software. The speed advantage is tremendous. But a firewall separates those agents from the production environment. A dropped database or exposed records stay inside our network; in production, those errors could be catastrophic. Trust should be proportional to the consequences of failure.
Trust is earned, but never permanent
An agent should earn more responsibility only after both the vendor and the company deploying it have vetted it. The vendor’s testing and the company’s own monitoring must show that the agent operates consistently within its limits. And that vetting never ends. Governing agents isn’t a one-time policy; it’s a continuous evaluation loop. Models change as they are updated and retrained. A new version, new data source, or new tool capability can change how an agent behaves, and each one deserves the same scrutiny as the original deployment.
This isn’t IT’s problem alone
Once agents move from experiments into daily operations, governance means understanding their business impact in dollars and legal terms, not just technical ones. On the cost side, calls, compute, and downstream services add up fast, so spending must be part of what gets monitored. On the contracts side, traditional software agreements assume the customer is the one operating the application, but agent agreements need to account for software that can make decisions, interact with other systems, and take action on a customer’s behalf.
Because agents touch spending, contracts, security and customer relationships, governance can’t sit with one team alone. HR, finance, and legal are obvious examples, and as agents gain more influence across organizations, more functions will need a seat at the table. Regulators have a role here too, and will likely need to set up expectations around transparency, accountability, and the use of agents in high-impact decisions. However, those rules should establish clear guardrails without freezing innovation or putting companies in any one region at a competitive disadvantage.
The real test for insurers
The promise of agents is that they can handle more work with less human intervention. The challenge for insurers is knowing which decisions can be delegated, which require oversight, and how to prove that the system is operating within the boundaries they set.
Doug Marquis is chief technology officer of Zywave, a Milwaukee-based insurance technology (insurtech) company that delivers unmatched insurance intelligence using the Zywave Apex™ AI platform so insurance brokers, agencies, MGAs (managing general agents), and insurance carriers can grow their business.
