Skip to Content
OPINION

Notes from the AI Governance Center: Security breaches, safety debates and transparency rules

Recent AI security incidents, open-model debates and new EU guidance highlight the growing importance of practical AI governance.

Published
Subscribe to IAPP Newsletters

Contributors:

Ashley Casovan

Managing Director, AI Governance Center

IAPP

Editor's note

The IAPP is policy neutral. We publish opinion pieces to enable our members to hear a broad spectrum of views in our domains.

The goal of this "Notes from the AI Governance Center" series was intended to be a way to take a step back and reflect on one AI governance issue that might not be getting the most attention. While this remains the goal, I can’t help but reflect on some more newsworthy items this month. 

From AI security breaches to new European Union guidance that might change the way we interact with AI, July has proven to be another monumental month for the AI ecosystem.  

AI escapes its sandbox

In case you missed it, mid-July saw the first major fully autonomous cyberattack. The incident involved a set of two OpenAI models tasked to perform a cybersecurity evaluation in a secure testing environment. Instead of working within their sandbox to solve the assigned task, they devised that an answer key for their exam was available on the open-source hub Hugging Face. The models found a way out of their sandbox and proceeded to breach Hugging Face's production environment. 

Prior to determining these models were from OpenAI, Hugging Face reported this intrusion was agentic and was capable of not only accessing their production environment, but accessed a "limited set of internal datasets and several credentials used by (their) services" and alerted the authorities.  

A precursor to this event occurred earlier this year when an early version of Anthropic's model, Mythos, was tasked with escaping its sandbox. Not only was it able to do this, it went beyond its assignment and found a way to alert its developer, identify zero-day vulnerabilities and post to GitHub. 

Whether these incidents were due to AI models exceeding imagined capabilities, poor prompting, inadequate alignment or limited regulations for foundational models, these incidents have raised a lot of questions and "told you so's" from AI safety researchers. 

In a joint statement Open AI and Hugging Face underscored that this attack was an "unprecedented cyber incident," from which both organizations indicate they have learned a significant amount. However, they also predict these types of autonomous attacks will "become more commonplace."

Safety researchers including, the AI Futures Project, have been predicting an agentic attack such as this in their AI 2027 project. Additionally, organizations including Guidelight AI Standards suggested that if their standards for AI foundation models would have been followed this incident could have been avoided. 

Whether you are in the camp that thinks this is a routine cybersecurity incident or a major alignment concern that requires government intervention with up-to-date guardrails, this incident demonstrates that downstream deployers of these foundational models need to have strong incident and cybersecurity defenses in place to match the capabilities of these new models.

Additionally, cyberattacks are not the only incidents that AI safety researchers worry about. While this was an example of an AI causing a security breach, other harms that are tracked through projects like the AI Incident Database demonstrate that a robust and comprehensive AI governance program that anticipates these risks as part of a mitigation and procurement strategy is necessary for appropriate adoption of these systems. 

Hugging Face concluded their incident report by sharing that "Autonomous, AI-driven offensive tooling is no longer theoretical ... Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defense to keep pace."

Open vs. closed

In response to these growing threats and incidents, this week we saw industry leaders including Nvidia, Adobe, RedHat, Cloudflare, IBM and several others launch an Open Secure AI Alliance for AI Safety and Security.  

They indicated their goal is to build and share open tools to promote responsible use and trust in AI and argue that both open and closed models are required to adequately defend and protect. This coalition calls on governments to invest in open infrastructure. 

This sentiment is also at the heart of what appears to be a growing geopolitical debate between tech rivals, China and the U.S., between open- and closed-sourced models. It was a significant discussion at the recent AI for Good conference and G7 meetings. 

Taking the other side of the debate, in response to the launch of this alliance, Anthropic CEO Dario Amodei explained why Anthropic didn't join this coalition. In an open letter, Amodei pushed back on accusations that Anthropic opposes open-weight models to protect its own proprietary business, stating the company "has never advocated for a ban on open-weight models" and calling capable-but-safe open-weight models a public good that benefits businesses, developers and researchers at little cost beyond compute.

Where Amodei diverges from the signatories is in how the debate should be framed relative to China. He argued the real risk isn't open-weight models themselves but authoritarian governments, particularly the Chinese Communist Party building AI systems powerful enough to secure permanent military superiority or enable deep domestic repression. He also flagged broader misuse risks, including AI-enabled cyberattacks, biological weapons development and unresolved alignment problems in powerful models.

To address these concerns, Amodei proposed a three-point plan. The first element being a ban on selling advanced AI chips and chipmaking equipment to Chinese firms, paired with a crackdown on hardware smuggling. His reasoning is that China's limited domestic chip production means that, given scaling laws, it cannot build more powerful models than the U.S. without access to U.S. chips making export controls the most direct way to mitigate this threat.

This appears to be the start of a potentially pivotal moment for AI development with the U.S. contemplating a ban on the use of AI models from China, which has prompted significant domestic debate. News from today indicates that the U.S. administration is now seeking to ban technology like AI-enabled humanoid robots from China. 

Will AI labels become as ubiquitous as cookie banners?

Staying out of the fray and continuing on with their regulatory agenda, the European Commission may not have garnered the same attention, but significant progress was made with the EU AI Act Digital Omnibus coming into force. 

The proposed simplification amends the AI Act's harmonized rules, easing burdens for subject matter experts and small mid-caps more than for large organizations. Beyond tweaks to data governance, technical documentation and value chain rules, three changes stand out. Article 4 on AI literacy reduces developer and deployer burdens and places more effort on regulators. Article 4a identifies that processing of special category data to detect and correct bias in high-risk systems is permitted under strict conditions, and Article 113 addressing high-risk compliance deadlines are pushed to December 2027 for Annex III systems and August 2028 for Annex I systems, due to a lack of technical standards. I shared more details on the changes via the AI Act Omnibus in an earlier article

However, finalized guidance on transparency obligations for AI-generated content were released in support of Article 50 received the most attention. It still seems like these guidelines are only being reviewed by a niche group despite having significant implications.

Requirements for Article 50 enter into force on 2 Aug. The finalized Code of Practice on Transparency of AI Generated Content is intended to help providers and deployers with compliance efforts. It indicates providers must mark and make AI-generated audio, image, video and text detectable via machine-readable, effective and interoperable technical solutions. Deployers must label deepfakes and AI-generated text on matters of public interest — unless the content underwent human review with editorial responsibility. There is a standardized set of icons deployers are to use for labeling.

Adherence is voluntary, but the underlying Article 50 obligations are binding. The Commission and AI Board have confirmed the code as an adequate tool for demonstrating compliance, giving signatories predictability and legal certainty across member states and access to Signatory Taskforces for shared implementation practices. Organizations complying by other means must independently prove those measures are adequate to national market surveillance authorities. 

I anticipate that as implementation progresses, there will be significant lessons learned on evolving best practices for how to comply with these new labeling requirements. We will follow this and share as they emerge. 

Additionally, a new action plan for cybersecurity and AI was shared. Given the above news, this seems like a timely and important effort. The action plan aims to promote the safe and responsible use of advanced AI, reinforce the EU's cybersecurity and resilience, and scale up Europe's AI capabilities for cybersecurity. It complements existing legislation, including the AI Act, Cyber Resilience Act, NIS2, DORA and the Cyber Solidarity Act, aiming to help Europe capture AI's benefits while staying resilient to emerging threats.

All supporting guidance and action plans connected to the AI Act can be found on the Commission's AI Act webpage. Here you will also find a list of forthcoming guidelines

What does this all mean?

For me, these three developments are substantial. As we see more real-life incidents of AI going rogue, it is clear to see the potential impact these systems will have. For better or worse. 

One important takeaway from Hugging Face is how essential AI was to identify and thwart the cyberattack. Despite the EU’s best efforts to anticipate best practices for the management of AI systems via guidelines, real life incidents will continue to be a driving force in shaping AI governance in practice. 

However, it does appear that the rules, standards and regulatory infrastructure that has been developed by Europe and other nations are useful frameworks. As we saw with the rapid response from industry players to the cybersecurity attack, they are calling for a lot of actions that governments, international organizations and safety researchers have already been publicly working toward. Either directly or indirectly informing the evolution of industry best practices. 

Looking at all of this through the current geopolitical lens is useful for all AI governance professionals to be aware of when determining what best practices your AI governance and digital governance teams will be establishing. Guided by requirements in legislation will provide a solid foundation, but it will also be important to understand both process and technical infrastructure needed to adequately govern the latest models.

This article originally appeared in the AI Governance Dashboard, a free weekly IAPP newsletter. Subscriptions to this and other IAPP newsletters can be found here.

CPE credit badge

This content is eligible for Continuing Professional Education credits. Please self-submit according to CPE policy guidelines.

Submit for CPEs

Contributors:

Ashley Casovan

Managing Director, AI Governance Center

IAPP

Tags:

AI and machine learningEU AI ActAI governance

Related Stories