SecurityTechInsider AI security & governance
EN/ NL
Policy

AI and the GDPR in 2026: why scraping and anonymisation must now demonstrably hold up

The EDPB clarifies with Guidelines 03/2026 and 02/2026 how the GDPR applies to AI training, web scraping and anonymisation. What does that mean for your AI setup?

18 August 2026 4 min
Illustration for this article: AI and the GDPR in 2026. A ventilation grille in deep shadow, cold air distorting the light passing through it.
Organisations must now document the legal basis for each AI workflow and demonstrate that models and training data meet GDPR requirements. Image: SecurityTechInsider — original editorial illustration

You must now document, per AI workflow, which model processes which personal data, on what legal basis, and at which stage—training, fine-tuning, inference or logging. This is no longer optional; it is a direct obligation under the GDPR, and you must be able to show your work.

The prompt is an analysis of 18 August 2026 of how the GDPR applies to AI training, web scraping and anonymisation, which argues that abstract appeals to innovation cannot override ordinary data protection duties. The analysis draws on the European Data Protection Board's Guidelines 03/2026 on web scraping and Guidelines 02/2026 on anonymisation, together with Opinion 28/2024. In our assessment, these clarifications close three persistent loopholes that organisations have relied on to avoid compliance: the claim that AI models fall outside the GDPR by nature, that publicly available data may be scraped without limit, and that synthetic or anonymised models automatically escape data protection law.

Which assumptions about AI and the GDPR no longer hold?

The EDPB has dismantled three claims that have shaped practice until now. First, AI models are not inherently anonymous. Through model queries, meaningful inferences or re-identification can occur, and as long as that remains possible, the model stays within the scope of the GDPR. Second, publicly available data carries no exemption. Web scraping involving personal data remains subject to the GDPR regardless of whether the source is public; organisations that scrape themselves or purchase pre-scraped datasets must document a legal basis, apply data minimisation and provide required information to data subjects. Third, synthetic or anonymised models do not automatically fall outside data protection law. The new anonymisation guidelines introduce a three-step test—No Record Isolation, No Linkage and No Inference—and models often remain within the GDPR because memorisation and inference attacks make re-identification possible.

What are the failure modes you must now guard against?

  • Undocumented legal basis — training or inference on personal data without a recorded lawful basis or demonstrable balancing of interests.
  • Unlawful provenance — using third-party models or datasets without due diligence on whether the training data itself was lawfully processed.
  • Model memorisation and inference — personal data reproduced or inferred from model outputs, making the model remain within GDPR scope despite claims of anonymisation.
  • Untraced data flows — personal data moving through training, fine-tuning, inference and logging stages without visibility over which controller or processor is responsible at each point.
  • Absent audit trails — no inspectable record of which verification steps, sources and corrections were applied to outputs in high-trust workflows.

Which concrete controls must you be able to demonstrate?

  1. Record the model and its purpose — document which model each workflow uses, which personal data it processes, and the lawful basis for that processing at each stage (training, fine-tuning, inference, logging).
  2. Apply the three-step anonymisation test — assess whether models and datasets meet No Record Isolation, No Linkage and No Inference; if they do not, arrange full GDPR compliance rather than claiming anonymity.
  3. Conduct a combined data protection and fundamental rights impact assessment — document the balancing of interests for each legal basis, particularly where legitimate interest is claimed for training.
  4. Register processing activities — enter each AI workflow into your record of processing activities (ROPA) with controller, processor and joint controller roles clearly assigned.
  5. Log and make inspectable — maintain audit trails of prompts, outputs and verification steps in workflows handling confidential or high-trust information, available for supervisory review.

How does the model memory itself change your compliance picture?

With traditional databases, personal data are traceable in discrete records. With AI, part of the risk shifts to what the model has memorised during training. The three-step anonymisation test forces you to assess not only the datasets you feed in, but also the model architecture itself and the release conditions before you can label a model as anonymous. This means your due diligence cannot stop at the dataset; it must extend to the model weights, the training procedure and the inference behaviour. If you use a third-party model, you cannot escape responsibility for understanding what it was trained on and whether that training was lawful.

What tooling can support compliance, and where does your judgement remain?

Verification layers can route tasks through selected independent models and make verification steps, corrections and sources visible for inspection. This supports review and control and gives insight into what happens; it is not a guarantee of correctness and does not automatically reduce hallucinations. Privacy-focused architectures can replace sensitive values with synthetic equivalents before processing and restore originals locally, sending only anonymised content onward. Audit logging can make verification steps inspectable during supervisory review. None of these tools removes your obligation to understand the legal basis for each processing step, to conduct the balancing of interests yourself, or to make the final judgement on whether a model or dataset falls within or outside the GDPR. The tools make your compliance verifiable; they do not make it automatic.

Sources: This article draws on reporting and guidance from EDPB, Alstonprivacy, Secureprivacy and Strac.

Elena Kovač

Written by

Elena Kovač

Follows EU policy as it turns from consultation into enforceable requirement.