SecurityTechInsider AI security & governance
EN/ NL
Data

Why pseudonymised AI data stays under the GDPR: the distinction EDPB 02/2026 sharpens

EDPB Guidelines 02/2026 clarify when AI data is truly anonymous and when pseudonymised data stays under the GDPR, with consequences for training and inference.

29 August 2026 4 min
Illustration for this article: Why pseudonymised AI data stays under the GDPR. Layers of translucent film lifted apart, light scattering between the sheets.
Pseudonymised data in AI workflows remains subject to GDPR obligations and cannot be treated as anonymous without passing a three-part cumulative test. Image: SecurityTechInsider — original editorial illustration

You must now treat pseudonymised data in your AI workflows as personal data subject to the full GDPR, not as a privacy-preserving exit from regulation. Hashing, tokenisation and coding are safeguards, not anonymisation. Your governance must distinguish which datasets are genuinely anonymous under a three-part test and which remain pseudonymous and therefore require data protection compliance.

The prompt is an analysis of 29 August 2026 of the distinction between pseudonymised and anonymised data under EDPB Guidelines 02/2026, which argues that pseudonymisation does not remove data from the scope of the GDPR and that much data labelled as anonymous in practice is legally pseudonymous. The analysis applies this framework to clinical trials and AI model training, showing how re-identification risk must be assessed contextually and over time. In our assessment, this clarification closes a gap in how many organisations have treated encoded or hashed datasets: the legal consequence is that your AI workflows cannot treat pseudonymisation as a route outside data protection law, and you must now document which datasets have passed a rigorous anonymisation test and which remain under GDPR obligations.

What makes data truly anonymous under the new framework?

The EDPB Guidelines 02/2026 establish a cumulative three-part test. Data qualifies as anonymous only if it meets all three criteria: No Record Isolation (the dataset cannot be broken down into individual records), No Linkage (no means exist to link records to an identifiable person), and No Inference (no information can be derived that re-identifies an individual). A dataset does not become anonymous simply because a name has been replaced with a code or hash. The Board stresses that anonymity must be tested per relevant entity and reassessed over time as re-identification techniques evolve. If any of the three criteria fails, the data remains personal data under the GDPR.

Where does pseudonymisation fit in your data governance?

Pseudonymisation is a security safeguard that reduces linkability but does not sever the link to an individual. It is a processing technique, not a privacy status that exempts you from data protection law. Common techniques—hashing, tokenisation, subject ID recoding—all amount to pseudonymisation as long as re-identification remains technically possible, which it usually does. Pseudonymised data therefore remains personal data and falls under Article 4(5) of the GDPR. The practical distinction is sharp: anonymised data can be shared or published under different rules; pseudonymised data cannot.

What controls must you document for pseudonymised AI workflows?

  1. Identify which datasets are pseudonymous and which are anonymous — record the anonymisation test results and the three-criteria assessment for each dataset in your AI pipeline.
  2. Map re-identification paths and key management — document where decryption keys, linkage tables or re-coding schemes are held and who may access them.
  3. Conduct a Data Protection Impact Assessment for pseudonymised processing — treat pseudonymised data as personal data and complete a DPIA before training models or running inference.
  4. Record the lawful basis for each pseudonymised dataset — document which legal ground (consent, contract, legal obligation, vital interest, public task or legitimate interest) justifies processing.
  5. Establish a verification layer in your workflow — implement a check that distinguishes pseudonymous from anonymous datasets before sending data to external AI models or services.
  6. Review re-identification risk contextually and over time — reassess whether new techniques or data combinations could re-identify individuals in your pseudonymised datasets.

What are the failure modes if you treat pseudonymisation as anonymisation?

  • Regulatory breach — processing pseudonymised data without GDPR compliance exposes you to supervisory action and enforcement.
  • Uncontrolled re-identification — sharing or publishing data you believe is anonymous when it is pseudonymous enables third parties to link records back to individuals.
  • Loss of audit trail — failing to document which datasets are pseudonymous means you cannot demonstrate compliance with data minimisation or purpose limitation.
  • Inadequate key management — treating pseudonymised data as anonymous removes the incentive to control access to decryption keys or linkage tables, increasing breach risk.
  • Model training on uncontrolled personal data — feeding pseudonymised data to external AI services without contractual safeguards may violate processor obligations or data transfer rules.

How should you structure AI workflows around this distinction?

Your workflow design must make the boundary explicit. Pseudonymised data should be processed on infrastructure you control, with documented key management and access controls. Before sending any dataset to an external AI model or service, you must verify whether it has passed the three-criteria anonymisation test. If it has not, it remains personal data and requires a data processing contract, a lawful basis, and technical safeguards. If it has, you may treat it under different sharing and publication rules. This is not a binary choice made once; it is a design decision you must revisit as your datasets, techniques and threat landscape evolve.

Tooling can help you track which datasets are pseudonymous and which are anonymous, and can flag when a workflow is about to send pseudonymised data to a service without the required safeguards. But the final judgement on whether your anonymisation test is rigorous enough, whether your re-identification risk assessment is complete, and whether your governance structure meets the GDPR remains yours. No system can replace that professional accountability.

Sources: This article draws on reporting and guidance from EDPB, Mdp-data, Secureprivacy and Iliomadhealthdata.

Noor El Amrani

Written by

Noor El Amrani

Data protection, anonymisation practice, and what regulators actually accept as evidence.