Skip to Content

EDPB discusses draft anonymization, AI web scraping guidelines

The European Data Protection Board's Gwendal Le Grand joined an IAPP LinkedIn Live to unpack the board's draft anonymization and AI web scraping guidelines while they each sit under public consultation through October.

Published
Subscribe to IAPP Newsletters

Contributors:

Lexie White

Staff Writer

IAPP

The European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology.

The guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus. The proposal raises the potential for fundamental alterations to the landmark data protection regulation, with notable changes to the definition of personal data under consideration.

During a LinkedIn Live with IAPP Research and Insights Director Joe Jones, EDPB Secretariat Deputy Head Gwendal Le Grand noted that while the EDPB supported the package, "We advised quite strongly against any change to the definition of personal data." He stressed the necessity for the guidelines while touting the EDPB's ability to provide timely resources to help companies navigate the compliance landscape before and after the finalization of the omnibus, which is still being negotiated between EU institutions.

Anonymization

Anonymization was last addressed by the EDPB through guidance issued in 2014 when the board was still functioning as the Article 29 Working Party. The guidance predated the adoption of the GDPR, leaving more than a decade of developments around data use cases and privacy-enhancing technologies to consider.

Reidentification within a specific context was a key focus of the new draft guide, according to Le Grand.

"Rather than asking if an individual is identified or identifiable in an absolute sense, the question is rather about the likelihood that the individual will be identified or identifiable by some entity," Le Grand said. "This may vary from one entity to another, and therefore anonymity has to be assessed from each relevant entity's perspective, meaning that any party for whom the data is intended to be anonymous."

To help organizations conduct assessments for anonymization, the EDPB outlined both contextual and simplified approaches for reviewing the legal standard for anonymized data.

Le Grand indicated the simplified approach offers a means to "voluntarily shift the risk from false positives to false negatives, where the anonymization controller treats anonymous data as personal data because they have, in effect, overestimated the likelihood that certain means will be used."

Though simplified efforts may lead organizations to unnecessarily consider information to be personal data, Le Grand said, "it can provide greater confidence, and it can be complemented with the contextualized approach to refine the finding."

As companies expand their abilities to process data, Le Grand said organizations should continue to reassess their compliance with the guidelines and the GDPR's obligations for processing sensitive data during their anonymization processes.

AI web scraping

The web scraping guidance runs complementary to the EDPB's parallel work with the European AI Office on joint guidelines detailing compliance with the GDPR and the EU AI Act.

The draft guidelines state organizations should only collect data they consider necessary for AI training while stressing data controllers must comply with the GDPR's data minimization and purpose limitation obligations when collecting and storing consumers' personal data. Organizations must also provide transparency to individuals when required under the GDPR, including when personal data is collected from publicly available sources.

"Regardless of the safeguards you put in place, it's quite difficult to have 100% assurance that you're not collecting special categories of data," Le Grand said.

The CJEU's Case C-136/17 judgment ruled search engines may unintentionally process sensitive data. Le Grand said when using web scraping for AI, the same reasoning applies.

"The incidental and residual collection of sensitive data for AI training is not unlawful if the controller implements certain measures to prevent the dissemination of data," Le Grand said.

To avoid questionable personal data processing altogether, the EDPB recommended companies consider using synthetic data to train AI models. According to Le Grand, organizations could also "apply syntax-based filtering, replace some or all of the real data with synthetic data if it's possible, or anonymization to anonymize."

"We've made a lot of efforts to engage more proactively and be more transparent about our stakeholder engagement and about how we use the input that we receive through the stakeholder consultations," he said. "I can assure you that all the input you send us is considered, everything is analyzed and taken into account."

CPE credit badge

This content is eligible for Continuing Professional Education credits. Please self-submit according to CPE policy guidelines.

Submit for CPEs

Contributors:

Lexie White

Staff Writer

IAPP

Tags:

AI and machine learningLaw and regulationPrivacy-enhancing technologyRegulatory guidanceGDPRPrivacy

Related Stories