The European Data Protection Board (EDPB) adopted Guidelines 03/2026 on web scraping in the context of generative AI (Version 1.0) on 7 July 2026 and opened them for public consultation through 30 October 2026. The guidelines address large-scale automated extraction of personal data from publicly accessible web sources for the purpose of training generative AI models, applicable to any entity engaged in that activity where GDPR jurisdiction is triggered.
The EDPB confirms in the guidelines that public accessibility does not remove personal data from GDPR protection. Consent under Article 6(1)(a) is identified as generally unavailable as a legal basis for web scraping at scale because the requirements of Article 7, including free, specific, and informed consent with a genuine right to withdraw, cannot realistically be met in the context of large-scale automated collection from public websites. Legitimate interest under Article 6(1)(f) is the primary available legal basis but requires satisfaction of a three-part test: identification of the specific legitimate interest pursued by the controller, necessity of the processing for that interest, and a balancing of that interest against data subjects' rights and reasonable expectations. The EDPB states the balancing outcome is not automatic and depends on the specific context.
The guidelines directly affect AI developers, large language model providers, academic institutions and research bodies building training corpora, and cloud platforms supplying training data pipelines. Controllers relying on web scraping as a data source must publish detailed public privacy notices accessible at the point of collection where feasible, and must implement mechanisms for data subjects to exercise rights, including pre-collection opt-out. Robots.txt files, ai.txt files, CAPTCHAs, and login walls are characterised as relevant indicators of data subjects' reasonable expectations; controllers that technically override such signals face heightened scrutiny in the legitimate interest balancing assessment.
The guidelines are in draft form and subject to revision following the consultation period closing 30 October 2026. Member state supervisory authorities may take divergent interim positions pending final adoption. The guidelines do not resolve the application of Article 28 to data supply chain arrangements between scraping contractors and AI model developers, nor do they address Chapter V transfer obligations where scraped data is processed outside the EU by non-EEA entities.
Licentium advises AI developers, data processors, and technology companies on GDPR compliance for AI training pipelines and data acquisition strategies. Contact us at www.licentium.io if this development is relevant to your operations. Work we undertake includes legal basis assessments for training data collection, web scraping compliance reviews, data subject rights framework design, privacy notice drafting, and regulatory consultation responses.
Source: EDPB, Guidelines 03/2026 on web scraping in the context of generative AI (v1.0), 7 July 2026