The EDPB’s New Anonymisation Test Starts With What Someone Can Still Infer
The EDPB adopted draft anonymisation guidelines in July 2026, offering a practical framework for deciding when data is truly anonymous under the GDPR.
“Anonymous” is often used to mean that a name was removed. Privacy regulators are asking a harder question: can someone still be singled out, linked to other information, or inferred from what remains? On July 8, 2026, the European Data Protection Board adopted draft guidelines on anonymisation. The EDPB says the guidance is intended to clarify when data is anonymous, taking account of European court decisions and the perspective of the entity handling the information. The guidance is still open for public consultation through October 30, 2026. It should be read as a framework under development, not as a final safe harbor. Removing names is not the same as anonymising data A dataset can omit names and still describe people in ways that make them distinguishable. A timestamp, location sequence, rare diagnosis, device pattern, or combination of attributes may provide enough structure to identify or track a person when joined with another dataset. The EDPB frames the analysis around two questions: does the information relate to an individual, and is that person identified or identifiable? Whether a person is identifiable depends on the context and on means reasonably likely to be used by the relevant entity. That perspective is important. A small research team and a large platform may have different datasets, tools, and capabilities. Data that appears harmless to one organization may be linkable by another. Three ways a dataset can reveal too much The guidance describes three criteria that help test whether data is truly anonymous: Singling out: Can one person or device be separated from the rest, even without a name?
Linkage: Can records be connected to one another or to an outside dataset?
Inference: Can someone derive a meaningful fact about the person from the remaining information? If any of these remain possible, an organization needs more analysis. Aggregation can reduce risk, but it does not automatically remove it. A group of one, a very small group, or a group described with unusually precise attributes can still expose a person. Why this matters for AI training data Generative AI systems increase the scale of collection and reuse. A company may scrape public pages, combine records, and train a model without having a simple list of names. That does not answer whether personal data was processed. An anonymisation analysis must consider what was collected, how it was transformed, what the model or downstream system can retain, and what additional information can be used to identify or infer something about a person. “The data was publicly available” and “the names were removed” are not complete answers. The same issue appears in evaluation and telemetry datasets. A developer may believe a log is anonymous because it contains an identifier instead of an email address, while the sequence of requests or device attributes makes the user easy to distinguish. What teams should document Organizations working with sensitive or large-scale datasets can prepare by documenting: The data source, purpose, and legal basis for collection.
The transformations applied and what information they preserve.
The realistic means available to the organization and likely recipients.
Whether records can be singled out, linked, or used for inference.
How the risk changes when data is combined with public or commercial sources.
What happens when a model, vendor, researcher, or partner receives the result. The EDPB’s draft is valuable because it moves the conversation away from labels and toward capabilities. Privacy is not restored merely because a column called “name” was deleted. The real test is what the remaining data still allows someone to see.