EDSA who? What can guidelines do?
The EDPB (English EDPB) brings together the data protection authorities of the EU and EEA member states and the European Data Protection Supervisor (EDPS). Austria sits at the table via the DSB. The goal of the EDPB is to promote the uniform application and enforcement of the GDPR. To this end, the EDPB publishes various materials intended to facilitate the application of data protection law. EDPB guidelines are certainly not laws and do not have a comparable effect. They are not binding on data controllers, courts, or authorities; the CJEU has the final say anyway, and it has already contradicted the committee in the past. And yet: the guidelines are a benchmark. The authorities work with them; anyone who argues against them in proceedings carries the burden of argumentation.
Personal reference: Depends on who is asking!
The anonymization guidelines 02/2026 supersede the 2014 testing approach of the Article 29 Data Protection Working Party. The reason for their drafting and publication is essentially a 2025 judgment by the ECJ in Case C-413/23 P (EDPS/SRB).
The case can be roughly explained in one sentence: An authority sends pseudonymized statements to an external consultant and keeps the key for identifying the individual persons. The question: What does this mean for the application of the GDPR? The Court of Justice's answer: For the authority, they are personal data, a crystal-clear case. For the consultant, this may be viewed differently. What is decisive is the respective controller's ability to re-identify individual persons in the dataset. Against this background, the ECJ concludes: The view that pseudonymized data is always personal is incorrect.
This (long-discussed) topic can be well summarized by the term „relative personal reference.“ When assessing whether personal data is present, one must—to put it even more simply—say: „It depends on who is asking!“ Anonymity in this sense is no longer a uniform label that you stick (or don't stick) onto a dataset. Rather, the term also conveys a statement about a specific entity and what other information it has available in relation to a dataset (e.g., data keys).
The EDPB takes up precisely this topic in its Guidelines 02/2026. The decisive factor for the re-identification of individuals (and thus for the delimitation of the scope of the GDPR) is „the means [of the respective entity] that can reasonably be likely to be used.“ One must ask: „Who could, if they wanted to, establish a reference to a person and what would it cost them?“. This always concerns the perspective of a specific entity. The testing framework consists of three criteria that must be met cumulatively (Guidelines 02/2026 marginal nos. 52 ff):
- No isolation – the data does not represent a unique combination of characteristics of a single person.
- No merging – they cannot be combined with other sources for the same person.
- No inference – they cannot be used to trace back to a specific person.
All three criteria met: no personal data. If one fails, it does not automatically mean data protection law is applicable. The EDPB then requires further analysis (guidelines marginal nos. 95 et seq.). So between yes and no, there is some work. There are not only clear answers in this matter.
Web scraping for AI training: public does not mean ownerless
The 22 pages of the EDPB Guideline 03/2026 state whether European AI models can be trained legally (measured against data protection law). The defining principle here is initially unspectacular: as soon as personal data is collected, stored, organized, or retrieved, the GDPR applies. The fact that someone has put their data online themselves (whether on social media, in a forum, or on a corporate website) does not change this.
In practice, the legal basis is the controller's legitimate interest (Art. 6(1)(f) GDPR). The EDPB specifies the balancing criteria from its already published Opinion 28/2024 with examples for the scraping context. There is no automatic rule in either direction: neither does the legitimate interest justify web scraping per se, nor is it ruled out in every case. Necessity and a documented balancing test are required – as everyone knows and Art. 5(2) GDPR supports: if you do not write down this balancing test, in case of doubt, you haven't done it either.
New and practically relevant about the new guideline is the link to the principle of accuracy: The EDPB recommends scraping only from reliable sources, recording the timestamp of collection, and validating the data prior to training. In addition, there are requirements for data minimization as well as a special focus on purpose limitation and transparency – whereby informing data subjects may be omitted if it proves impossible or requires disproportionate effort.
The EDPB takes the strictest stance on sensitive data. Their processing is fundamentally prohibited under Article 9(1) GDPR; anyone who scrapes them as well needs a legal basis under Article 6 and generally must also fulfill an exception under Article 9(2) GDPR. For data that is accidentally or residually captured, the EDPB considers the ECJ jurisprudence in Case C-136/17 (GC et al. v. CNIL) to be potentially applicable, provided that the controller acts within the scope of its responsibilities, powers, and capabilities, and implements appropriate TOMs against collection and redistribution. However, the EDPB explicitly emphasizes that there is no general exemption from Article 9 GDPR. Case-by-case assessment remains required.
What to do by October 30 – two points
- Safeguarding your interests: Check your processing of (supposedly) „anonymous“ data against the three criteria of Guideline 02/2026 and document the results. For AI training data, provably record the origin, timestamp, and legal basis, and have the balancing test for legitimate interests to justify the processing in writing ready at hand—preferably before anyone asks for it.
- Get involved and take a stand: The public consultation is open to companies and associations, and practical input is especially welcome regarding legitimate interest and anonymization requirements. It is the only phase in which practitioners can help write the standards instead of simply being measured against them. Anyone who considers the requirements unrealistic later on would do better to say so now rather than during the audit procedure.
Conclusion
The EDPB is trying to bring order to difficult topics. For data-driven business models, this is mostly good news—provided that the respective responsible parties have the resources and discipline to carry out the necessary analyses and act decisively based on them. Things could become uncomfortable in the future if people settle for unsubstantiated claims: „We removed the names, so it's anonymous“ generally does not pass the aforementioned three criteria for testing personal data. „The data was freely available online, so we were allowed to process it“—that also misses the mark somewhat. Both statements harbor the potential for trouble—but that doesn't have to be the case. Against this background, it is clear that action is needed.
Frequently asked questions about the new EDPB guidelines (FAQ)
1. Are the new EDSA guidelines legally binding?
No, EDSA guidelines are not laws and do not directly bind companies or courts. However, they act as an important „benchmark“ for supervisory authorities. In the event of data protection proceedings, companies that deviate from the guidelines bear the full burden of argument.
2. What does „relative personal reference“ mean in the context of anonymization?
The term implies that anonymity is not absolute. Whether a dataset is personally identifiable depends on the ability of the respective entity to re-identify the data. It depends on who is processing the data and what additional information or technical means are reasonably available to that entity.
3. Am I allowed to simply scrape data from the internet for AI training?
Only to a limited extent. Publicly available data is also subject to the GDPR as soon as it is collected or processed. Legitimate interest (Art. 6 (1) (f) GDPR) can be a legal basis, but it strictly requires a documented balancing test as well as compliance with principles such as data minimization and purpose limitation.
4. What criteria must be met to classify data as „anonymous“?
The EDPB establishes three cumulative criteria: (1) no isolation (no unique combination of characteristics), (2) no linkage (no merging with other sources), and (3) no inference (no drawing of conclusions about a person). If all three criteria are met, there is no personal data reference.
5. What is scraping?
Web scraping refers to the automated extraction of data from websites using special software („crawlers“ or „bots“). This data is subsequently extracted, stored, and structured—for example, to generate datasets for training AI models. Since personal data is frequently collected in the process, the entire procedure is subject to the strict requirements of the GDPR.
Are you planning to use AI or do you need a legally compliant data concept?
As experts in Austrian data protection law, the lawyers at ATB.LAW advise companies on the implementation of these new EDPB guidelines. We review your data processing procedures and AI pipelines, draft data processing agreements, and protect your business model against regulatory bans. Contact Stefan Knotzer and Roman Taudes at any time under office@atb.law or by phone at 01 39 12345 for a non-binding initial consultation.