What Do the European Data Protection Board’s Web Scraping Guidelines Mean for AI Training Datasets?

On July 7, 2026, the European Data Protection Board (EDPB) published draft guidelines on web scraping for generative AI (Guidelines). The Guidelines are intended to provide practical GDPR guidance in one of the more complex areas of AI development and will be of direct relevance to any organization building or procuring generative AI systems trained on internet-sourced data.

To Whom Will the Guidelines Apply?

The Guidelines apply to both organizations that scrape data from external internet sources (directly or via a third party) to train generative AI systems and those that acquire and reuse pre-scraped datasets from third parties (e.g., a data broker). The EDPB’s focus is firmly on controller obligations, though the Guidelines map out how scraper and AI developer relationships fall across the controller/processor spectrum (see below).

Six Key Takeaways

In addition to providing context on the technical elements of the web scraping process, the Guidelines focus on six key areas of GDPR compliance:

1. Data Protection Roles: The EDPB confirms that role allocation is fact-specific and should be analyzed on a case-by-case basis. Practically speaking:

Role When It Arises
Processor Scraper acts on documented instructions from the AI developer (e.g., specified data sources and categories) with no independent decision-making. AI developer would be the controller.
Joint controller AI developer and scraper jointly agree on collection criteria, even if each performs distinct processing activities.
Separate controllers Where an AI developer sources a pre-scraped dataset, each entity is responsible for its own processing (i.e., the scraper is not responsible for the reuse of the data by the AI developer).

Practical takeaway: Organizations should consider assessing the GDPR role they play in the scraping ecosystem to confirm what obligations may apply. Organizations acquiring third-party datasets may want to conduct upstream due diligence, as the AI developer bears independent controller responsibility for its own use of that data.

2. Legal Basis: Personal data shall be processed lawfully. The Guidelines focus on two GDPR legal bases:

    1. Consent: The EDPB considers that consent will likely be unavailable. Where personal data are collected indirectly and at scale, there is typically no direct relationship with data subjects, and they are rarely identifiable in order to be able to obtain consent. Freely given, informed consent is therefore likely unworkable in practice, and the EDPB is clear that simply making data publicly available does not amount to giving consent to its scraping.
    2. Legitimate Interest: The EDPB considers this to be the more viable route. Legitimate interests would require genuine engagement with all three limbs of the legitimate interest test – and substantive detail is provided by the EDPB in this regard. Controllers must: (1) identify a legitimate interest, (2) demonstrate that the processing is necessary to achieve it, and (3) show (through a documented balancing exercise) that their interests are not overridden by those of the data subjects. The EDPB acknowledges that growing public awareness of online data use may weigh in a controller’s favor but is careful to note that this does not give controllers carte blanche for all situations and all purposes.

Practical takeaway: Where looking to rely on legitimate interests, the balancing test must be documented and tailored to the specific processing activity. Further, organizations may want to consider implementing technical and organizational measures that reduce privacy impact (e.g., consider filtering out sensitive data categories and pseudonymization).

3. Data Minimization: Personal data shall be adequate, relevant, and limited to what is necessary in relation to the purposes for which they are processed. The EDPB is clear that, despite this being a “major challenge” for large-scale scraping, compliance with the data minimization principle applies from pre-collection through output.

Practical takeaway: Build data minimization principles into the scraping architecture itself. For example, collection criteria can be defined in advance; consider filters to exclude sensitive data categories, high-risk websites (e.g., those directed at minors), and those which clearly prohibit scraping; and post-collection cleansing of the data can be implemented. The use of synthetic data as a substitute for real personal data should also be considered.

4. Transparency: Personal data shall be processed in a transparent manner in relation to the data subject. The EDPB acknowledges that it is often “difficult, impracticable or, even, objectively impossible” to identify and notify data subjects individually when scraping at scale. However, the disproportionate effort exemption can be viewed as other than a default fallback. Reliance on this exemption requires case-by-case justification with consideration given to the number of data subjects, the age of the data, and any appropriate safeguards adopted. The EDPB provides examples as to what may or may not satisfy the test:

    • Likely disproportionate effort: Large-scale collection of data from a variety of sources spanning 20 years, with no direct identifiers, covering thousands or millions of individuals, where the controller has published its privacy notice and excluded directly identifiable data.
    • Unlikely disproportionate effort: Targeted collection from a closed group of 5,000 identifiable individuals spanning two years.

Practical takeaway: Where relying on the exemption, a publicly accessible privacy notice is required. This may include: (a) a list of the sources “to the greatest extent possible,” (b) whether the sources are publicly accessible (or not), and (c) where applicable, the crawler’s characteristics. However, good practice may also require inclusion of domain names and URLs of scraped pages (in searchable format), the date range of collection, and, where data was purchased, contact details of the originating controller. Other measures to consider may include undertaking a data protection impact assessment (and making it publicly available).

5. Accuracy: Personal data shall be accurate and, where necessary, kept up to date. Scraped data can be unreliable (i.e., it may be outdated, user-generated, or sourced from sites with no editorial standards). In turn, compliance with the GDPR’s accuracy obligation — which applies both to the data collected and to the outputs the trained model produces (with a view to minimizing the risk of incorrect output) — can be challenging. However, the consequences for getting this wrong can be far-reaching, as inaccurate training data increases the likelihood of outputs that are factually wrong or harmful.

Practical takeaway: Consider timestamping data at collection, restricting scraping to credible sources where possible, and validating the data before it enters training pipelines.

6. Special Category Personal Data (SPD): The EDPB generally recommends that a controller should implement measures to prevent the collection of SPD. Where that is not technically feasible, the EDPB confirms that an Article 9(2) GDPR condition will be required and that this must be assessed on a case-by-case basis. However, the Guidelines also introduce a nuanced framework drawing on GC & Others (C-136/17) — a search-engine case — to address the incidental and residual collection of SPD inherent to large-scale scraping. In short, (a) where the processing is analogous to a search engine, (b) SPD collection is incidental rather than deliberate, (c) it is genuinely difficult to assess whether the SPD is present, and (d) the controller has implemented robust technical and organizational measures to prevent collection and dissemination, the EDPB suggests that the processing is not automatically unlawful.

Practical takeaway: To rely on the GC & Others framework, each of the criteria (a) through (d) identified above should be met and documented. Technical and organizational measures implemented should span the full lifecycle (e.g., filtering at collection, deletion post-collection, resistance to privacy attacks during development, and output filtering post-deployment with ongoing monitoring).

Conclusion

The Guidelines are a practical step forward. They do not resolve the fundamental tension between large-scale AI training and the individual-centric protections GDPR was designed to provide, but they do provide a clearer map of the compliance landscape and a set of concrete measures that, taken together, represent the foundations of a defensible position.

The core message is clear — publicly available data is not freely useable data, and it remains subject to the GDPR. The compliance burden falls on the controller, and it arises before scraping begins.

The EDPB’s consultation for feedback on the Guidelines remains open until October 30, 2026.

This post is as of the posting date stated above. Sidley Austin LLP assumes no duty to update this post or post about any subsequent developments having a bearing on this post.