At its July 2026 plenary, the European Data Protection Board adopted Guidelines 03/2026 on web scraping in the context of generative AI (Version 1.0). The guidelines address when and how GDPR requirements apply to large-scale automated extraction of publicly available personal data used to train generative AI systems. A public consultation runs until 30 October 2026, after which the EDPB may revise the text before issuing a final version.
The EDPB concluded that GDPR applies whenever a web scraping operation collects or processes personal data, regardless of whether the source was publicly accessible online. Article 5(1)(b) of GDPR governs purpose limitation: data originally published for a different purpose may not be re-used to train AI models without an independently valid legal basis. Processing of special categories of data under Article 9, including health, racial origin, and political opinion data incidentally scraped, requires explicit consent or another enumerated Article 9(2) ground.
Companies building or operating generative AI systems in the EU face the most direct impact. Foundation model providers, fine-tuning operators, and data aggregators supplying training sets must assess each scraping operation against the legal basis, purpose compatibility, and transparency requirements set out in the guidelines. Organisations relying on legitimate interest under Article 6(1)(f) must conduct a balancing test that accounts for the scale of the scraping and the absence of individual notice to data subjects.
The guidelines are in consultation phase and the EDPB may revise specific positions before finalisation. Separately, the EDPB adopted guidelines on anonymisation at the same plenary. Web scraping that produces genuinely anonymised outputs before those outputs enter training pipelines would fall outside GDPR scope, but the anonymisation guidelines set a high threshold for what qualifies as sufficiently anonymised.
We advise on GDPR compliance for AI training data programmes and maintain a partner network for multi-jurisdictional data protection engagements. Contact us to discuss your specific training data practices. Work we undertake includes AI training data audits, web scraping compliance reviews, data protection impact assessments for generative AI systems, and legitimate interest balancing documentation.