Skip to Content
OPINION

Notes from the AI Governance Center: AI auditing is getting amplified

AI incidents, industry proposals and emerging regulations are increasing focus on AI audits, independent evaluation and assurance programs.

Published

Contributors:

Ashley Casovan

Managing Director, AI Governance Center

IAPP

Editor's note

The IAPP is policy neutral. We publish opinion pieces to enable our members to hear a broad spectrum of views in our domains.

This month has seen another whirlwind of generative AI incidents. Or at least the reporting of them. From more AI agents getting up to no good, to the first known intentional agentic AI attack reported to a regulator, followed by the first, and then second government data breaches by an AI agent, Axios has reported that companies are currently probing tens of thousands of security incidents.

While these AI incidents are starting to sound commonplace, each event should be taken very seriously. As we can see from this month's news, while many of these incidents can be traced back to a frontier lab experimentation gone awry, we are also starting to see how criminals are leveraging these tools.

Perhaps what captured the public's attention the most this month were the first public warnings from current and former employees of these AI companies who anticipate that AI will cause a catastrophic end of the world within 10 years. Whether your P(doom) has been altered as a result of this news or not, it is important to note that the CEOs from most U.S. AI companies have also joined the chorus to say some action needs to be taken.  

Action equals AI audits

Not going as far as calling for regulations, the composite of these incidents has spurred more calls for AI governance controls from these companies and governments than I have ever seen. 

In Anthropic CEO Dario Amodei's recent essay Pacing the Frontier, he recommends establishing "embedded evaluators" within AI companies as a way to align and safeguard their models. Others have referred to AI auditors as a function of an AI assurance program. 

Putting aside the current immature state of standards for various types of AI systems needed to perform evaluation, we agree that the idea of having people within organizations responsible for performing assessments, evaluation, and testing and ongoing monitoring is good AI governance practice.  

Amodei expands to share that all frontier labs should have "embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes." 

This call was immediately supported by other CEOs of frontier AI companies including Sam Altman and Elon Musk. On 29 Sept., the White House announced a new Accord recognizing the importance of controls and audits for AI aligned with the ideas put forth by Amodei. The Accord points to both internal and external audit functions, including executive oversight from an independent committee of the board of directors. 

While all these efforts are currently focused on frontier AI labs, it is important to note that this will have downstream effects on their customers and will help to establish best practices for all types of AI uses. 

These concepts are not new

While the recent announcements of incidents and calls for slowing down the development of AI have captured the public's attention, the concept of embedded evaluators or AI auditing is not new. Several organizations, including the U.K. government, Ada Lovelace Institute, Center for Democracy and Technology, TechUK and the Partnership on AI all identified that AI audits were a useful tool to mitigate risk from existing AI systems. 

As discussed in much of this research, the idea of independent evaluation, verification and assessment is common in many industries — in particular industries whose systems and services could lead to safety incidents. 

So, while the suggestion of embedded evaluators might be novel for AI companies to have, and beyond what any regulation currently requires, it is a concept that has worked well in other industries. 

A note on independence: Amodei argues that with employee-like access these evaluators will be able to operate as independent safety checks for the organization while remaining neutral. I think that this will be important to watch. 

In an op-ed for The New York Times, former U.S. Federal Trade Commission Chair Lina Khan notes there are real concerns with AI companies self-governing. She points to existing laws that AI companies already need to adhere to which could include external review requirements from government as well as a growing number of lawsuits against these companies. 

In her article, "In AI evaluation, access is not evidence," Brookings Artificial Intelligence and Emerging Technology Initiative Director Elham Tabassi points to important questions about the nature of evaluators and how to ensure they are meaningful, asking, "who controls the methodology? Who controls access? Who funds the work? Who controls publication?" She warns, "Those dependencies should be disclosed and, where possible, constrained. An evaluation is not made credible simply because someone outside the company ran it — or, for that matter, someone inside it." 

Additionally, what is also common in other industries is that the people performing these evaluations are doing so against a set of standards and have the knowledge, skills and capabilities to do so. The research also points to the need for increased training, standardization and certification to formalize and increase the maturity of the people tasked with performing these evaluations. If not, we are opening ourselves up to companies writing — or funding the writing of — their own exam and hiring evaluators who do not need to demonstrate capability to perform these evaluations. 

Governments agree

While the EU is the only major jurisdiction with a binding, horizontal conformity-assessment regime through the AI Act, with core obligations coming into effect in December 2027, other regions in the world have identified the need for audits in AI regulation, including China, which requires compliance audits every two years "for companies processing personal information for more than 10 million individuals." Additionally, regulators can order external audits as needed, and there are separate algorithm filing and security assessments for public-facing AI. 

In the U.S., some states are thinking about the role of assessments and evaluations. In Illinois, SB315 requires audits and third-party assessments for catastrophic risk, and the Great American AI Act 2026 is a similar bill at the federal level. While not required, but recommended, California and New York also identify the value of audits to mitigate catastrophic risk. Bills with a more narrow scope are also including audit requirements. California's SB1119 is the first legislation passed with a focus on companion chatbots. Many more bills with audit requirements have been introduced across U.S. state legislatures. 

To support the increasing evaluation and audit demand, California's AB1405 and SB813 and Virginia's SB384 establish regulatory frameworks for independent verified organizations. California's legislation goes further than Virginia's by creating a central registry for approved AI auditors with the requirement for all AI auditors to have a relevant certification by January 2029. The Great American AI Act also proposes a licensing regime for IVOs, with required audits for catastrophic risk models every six months.

More common are voluntary programs established by governments, most popularly in Singapore and the U.K., with incentives and infrastructure to support companies that would like to leverage these tools. With high adoption rates of these voluntary tools in both regions, the need for training and certification against voluntary tools has a similar market demand to required laws at present.

The new White House Accord notes at the end that the control and audit actions "will give each company, its customers and the public confidence that the technology is operating as intended" and "over time, it may make sense to codify these steps into laws or regulations."

Where do we go from here?

While the future of embedded evaluators within AI companies remains to be seen, it does seem clear that independent evaluators and auditors for AI are becoming a more common concept. 

As these tools become even more embedded into our professional and personal lives with September also seeing the launch of the first personal assistants from Meta and OpenAI, ensuring that they operate in a safe and responsible manner is getting personal. 

AI auditing and assurance is something that I have been interested in and researching for many years. Alongside our upcoming IAPP P.S.R. + AIGG 2026 conference, I have the opportunity to partner with the CDT on a session to better understand the current AI assurance and AI audit landscape. If this is something that is also of interest to you, please reach out to me directly to be included in this conversation which will take place 7 Oct.  in Seattle. If you can't make it, expect more information on the AI assurance space from us soon.

This monthly column originally appeared in the AI Governance Dashboard, a free weekly IAPP newsletter. Subscriptions to this and other IAPP newsletters can be found here.

CPE credit badge

This content is eligible for Continuing Professional Education credits. Please self-submit according to CPE policy guidelines.

Submit for CPEs

Contributors:

Ashley Casovan

Managing Director, AI Governance Center

IAPP