AI en de AVG in 2026: waarom scraping en anonimisatie nu aantoonbaar moeten kloppen
De EDPB verduidelijkt met Guidelines 03/2026 en 02/2026 hoe de AVG geldt voor AI-training, webscraping en anonimisatie. Wat betekent dat voor uw AI-landschap?
U moet per AI-workflow aantonen welke persoonsgegevens erin verwerkt worden, op welke rechtmatige grondslag dat gebeurt, en wie daarvoor verantwoordelijk is. Die verplichting geldt nu, ongeacht of u zelf traint, modellen van derden inkoopt of datasets scrapt.
The prompt is an analysis of 18 August 2026 of webscraping, anonimisatie en rechtmatigheid in AI-training onder de AVG, which argues that the European Data Protection Board has closed three persistent loopholes in how the GDPR applies to AI model training and deployment. The analysis uses the EDPB's Guidelines 03/2026 on web scraping and Guidelines 02/2026 on anonymisation, alongside Opinion 28/2024, to show that abstract appeals to innovation do not override ordinary data protection duties. In our assessment, this means you can no longer treat AI systems as exempt from GDPR compliance; you must now document and verify your legal basis, your data sources, and your model's capacity to re-identify individuals before you can claim any exemption.
Welke drie aannames zijn nu onhoudbaar geworden?
The EDPB has systematically dismantled the reasoning that allowed many organisations to treat AI outside the GDPR's scope. First, AI models are rarely truly anonymous. Opinion 28/2024 establishes that queries to a model can produce meaningful inferences or even re-identification of individuals whose data was in the training set. As long as that is possible, the model remains within the GDPR's reach. Second, publicly available data carries no automatic exemption. Guidelines 03/2026 explicitly state that data being publicly accessible does not exempt web scraping from GDPR obligations. If personal data is involved, you must document a legal basis, apply data minimisation, and fulfil transparency duties. Third, synthetic or anonymised models do not automatically fall outside privacy law. The new anonymisation guidelines introduce a three-step test—No Record Isolation, No Linkage, and No Inference—that forces you to verify whether memorisation and inference attacks could allow re-identification. Most models fail this test and remain within the GDPR.
Welke GDPR-plichten gelden nu voor AI-training en -gebruik?
The core principles of the GDPR apply directly to AI workflows. You must establish who is the controller for each AI system, which personal data is used at which stage (training, fine-tuning, inference, or logging), and on which legal basis that processing occurs. Opinion 28/2024 makes clear that relying on "legitimate interest" as your legal basis for training requires substantially more rigorous justification, with a documented balancing test. You also cannot escape due diligence on data provenance. If you use models or datasets from third parties, you must verify the lawfulness of the training data they contain. Models trained on unlawfully processed personal data can propagate that harm downstream, and you bear responsibility for that chain.
The risk profile differs from traditional databases. In conventional storage, personal data sits in discrete records. In AI systems, part of the risk shifts to what the model has retained in its weights. The three-step anonymisation test forces you to evaluate not only datasets but also model architecture and model release procedures before you can claim anonymity.
Welke concrete controles moet u per AI-systeem kunnen aantonen?
- Documenteer rechtmatigheidsgrondslag en belangenafweging — record which legal basis applies to each workflow (training, inference, logging) and, if relying on legitimate interest, document the balancing test.
- Voer een gezamenlijke DPIA en FRIA uit — conduct a data protection impact assessment and, where required, a fundamental rights impact assessment before deploying the model.
- Registreer alle verwerkingen in het verwerkingsregister — maintain a record of all processing activities, including model name, purpose, data categories, retention, and controller role.
- Verifieer trainingsdata-herkomst — test whether training data was lawfully processed and whether the model passes the three-step anonymisation test before claiming any exemption.
- Maak prompts, outputs en logs inspecteerbaar — ensure that in high-trust workflows, audit trails of AI interactions are available for inspection and can be linked to the legal basis and controller decision.
Hoe vertaalt u deze richtsnoeren naar zichtbaarheid en controle?
The translation from guideline to practice requires visibility across three dimensions: which models and datasets run where, which GDPR roles (controller, processor, joint controller) apply to which processing, and what legal basis and balancing test are documented. You also need to know how prompts, outputs, and logs can be retrieved for audit. Verification tooling can support this. A privacy-focused verification layer can route tasks through selected independent models, make verification steps, corrections, differences, and sources visible for inspection, and help you see what is happening in your AI chain. This supports control and auditability; it is not a guarantee of correctness and does not automatically reduce hallucinations.
For anonymisation specifically, architectural approaches exist that replace sensitive document values with synthetic, session-bound equivalents before AI processing occurs, analyse the synthetic version, and restore original values locally afterward. The workflow is fail-closed: if privacy verification fails, the document is not sent. This is an architectural choice, not a promise about the degree of anonymisation or full GDPR compliance. The professional judgment remains yours.
Wat kunnen tools doen, en wat blijft uw eigen verantwoordelijkheid?
Verification and logging functions can make your compliance steps inspectable, which helps auditors see which verification steps and sources were used. Tooling can also make the data flow and role assignments visible across models and suppliers. What tooling cannot do is replace your professional judgment about whether a legal basis is sound, whether a balancing test is adequate, or whether a model's architecture truly prevents re-identification. The EDPB's new lines make clear that AI compliance under the GDPR is no longer a paper exercise in 2026. It is now a verifiable layer that must demonstrably work. You must be able to show your work, and you must be able to defend it.
Bronnen: Dit artikel is gebaseerd op berichtgeving en richtlijnen van EDPB, Alstonprivacy, Secureprivacy en Strac.
Geschreven door
Elena Kovač
Volgt EU-beleid op het moment dat het van consultatie naar handhaafbare eis gaat.