The Privacy Partnership Podcast with Robert Bateman

Web scraping under the GDPR: The EDPB's uncharacteristically pragmatic solution

treborjnametab1

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 7:27

Can you scrape the internet for AI training data without completely running afoul of the GDPR? The European Data Protection Board (EDPB) has finally offered an answer: Yes, but get ready to implement a massive amount of filtering.

In this episode of the Privacy Partnership Podcast, Robert Bateman breaks down the EDPB’s newly adopted Draft Guidelines 03/2026 on web scraping for generative AI. Robert begins by exploring the political context behind this unexpectedly pragmatic guidance, discussing how the EDPB is effectively front-running the European Commission’s upcoming "Digital Omnibus" proposal to cement its authority over how privacy law applies to AI development.

Then, Robert walks listeners through a practical, 10-point checklist for developers and privacy teams trying to navigate this regulatory minefield, from mapping out complex controllership arrangements to leveraging a fascinating loophole for the "incidental and residual" scraping of sensitive, special category data.

Key Topics Discussed:


The Digital Omnibus Context: Why the EDPB’s new guidance is "deceptively permissive" and how it serves as a strategic maneuver to preempt upcoming EU legislation.

Controllership in the AI Supply Chain: How to define your role—whether you are dictating instructions to a scraper, co-determining collection criteria, or buying a pre-scraped dataset.

Establishing a Lawful Basis: Why consent is a non-starter at this scale, how to lean on Legitimate Interests, and why a missing "robots.txt" file does not equal a green light.

Designing the Collection: The end of indiscriminate web hoovering, the importance of data minimisation, and respecting technical barriers (like CAPTCHAs and ai.txt).

Transparency at Scale: How to utilize the Article 14 "disproportionate effort" exception while maintaining a highly detailed, searchable public scraping notice.

Cleaning and Accuracy: Applying syntax-based filters to weed out format-identifiable data on the fly, and utilizing synthetic data where feasible.

The Article 9 Workaround: How the EDPB is applying the 2019 GC & Others CJEU search engine ruling to allow the incidental scraping of special category data—and the rigorous output filters required to justify it.

Accountability: The massive documentation burden required to prove your technical measures and filters remain effective against the evolving state of the art.