SecurityTechInsider AI security & governance
EN/ NL
Governance

GDPR responsibilities per phase of your generative-AI workflow

The EDPB web-scraping guidelines and the AP guidance bring generative AI under the GDPR. Here is how to divide responsibilities per workflow phase in concrete terms.

8 September 2026 4 min
Illustration for this article: GDPR responsibilities per phase of your generative-AI workflow. Poured concrete meeting brushed steel at a tight seam, the joint slightly misaligned.
Organisations must now document controller and processor roles and lawful bases for each phase of AI workflows, or face integrated GDPR and AI Act enforcement. Image: SecurityTechInsider — original editorial illustration

You must now document which role — controller or processor — your organisation holds in each phase of your generative-AI workflow, record the lawful basis for each phase, and demonstrate that division in writing to a regulator on demand.

The prompt is an analysis of 8 September 2026 of the application of GDPR principles to generative-AI processing at each workflow phase, which argues that the GDPR applies in full to scraping, training, inference and logging, and that the AI Act adds obligations on top rather than replacing data protection duties. The European Data Protection Board adopted guidelines on web scraping for generative AI in July 2026, followed by national authority guidance treating AI model development and deployment as ordinary GDPR processing. In our assessment, organisations can no longer treat GDPR compliance and AI Act compliance as separate workstreams. A single AI workflow touches both regimes at once, and a regulator reviewing your practices will examine them together.

Does the GDPR actually apply to my training data?

Yes. The EDPB has rejected a generic exemption for AI training. The principles of purpose limitation, transparency, data minimisation and accuracy apply before you even identify a lawful basis. If you rely on legitimate interest under Article 6 of the GDPR, that basis must pass a three-step test: you must identify a legitimate purpose, demonstrate that the processing is necessary to achieve it, and show that your interest outweighs the rights of the people whose data you are processing. This applies whether you scrape data yourself, outsource scraping to a third party, or purchase ready-made datasets or pre-trained models from a supplier. You cannot fully shift the GDPR responsibility onto a vendor.

Which failure modes and duties apply across the workflow?

  • Unlawful scraping without a stated basis — collection of personal data from the web without identifying a lawful basis under Article 6 of the GDPR.
  • Purpose creep in model training — using data collected for one purpose to train models for a different purpose without fresh consent or a compatible basis.
  • Inadequate due diligence on third-party data — purchasing datasets or models without contractual evidence that the supplier complied with the GDPR.
  • Opacity in inference and logging — deploying a model without documenting which personal data it processes, how long logs are retained, or who can access them.
  • Conflicting AI Act and GDPR obligations — treating high-risk AI system requirements and data protection duties as separate compliance tracks instead of integrated controls.

How do I divide controller and processor roles across phases?

For a concrete workflow — such as an internal knowledge assistant or AI-supported customer communication — you should document the following:

In the scraping phase, identify whether your organisation is the controller (you decide what data to collect and why) or a processor (you collect data on behalf of another organisation). In the training phase, record who decides which data enters the model and on what basis. In the inference phase, document who is responsible for the personal data the model processes when it generates outputs, and whether those outputs are logged. In the logging phase, establish who decides how long logs are kept and who can access them.

Record this division of roles in your contracts with any third parties and in a Data Protection Impact Assessment (DPIA). The audit rights and evidence obligations you write into AI contracts are the mechanism for making the due-diligence duty enforceable in practice.

What concrete controls must you be able to demonstrate?

  1. Document the model and its purpose — record which model each workflow uses, the lawful basis for the personal data it processes, and the controller and processor roles in writing.
  2. Maintain a processing inventory per phase — list the categories of personal data, the retention period, and the recipients at each stage from scraping through logging.
  3. Establish contractual audit rights — require any external scraper, trainer or model provider to grant you access to evidence of GDPR compliance and to respond to regulator requests.
  4. Implement a data minimisation checkpoint — before data enters training, verify that you are not collecting more data than necessary and that you have a stated purpose for each category.
  5. Log inference and output handling — record which personal data the model processes during inference, how long outputs are retained, and who accesses them.

What can tooling do, and what remains your responsibility?

Verification and privacy-enhancing tools can make your data flows and processing steps visible and can help you route sensitive data away from certain models. They can prepare processing on infrastructure within the EU and can block forwarding of personal data if a privacy check fails. None of this changes the underlying GDPR obligations or removes your accountability for the lawfulness of the processing. A tool can support your compliance work and can make your controls auditable, but it cannot guarantee compliance. The final professional judgement on whether your legal basis is sound, your processing is necessary, and your safeguards are adequate remains yours alone.

Sources: This article draws on reporting and guidance from European Data Protection Board, Licentium, Loyens & Loeff and Praxikon.

Marit Halversen

Written by

Marit Halversen

Covers AI governance and regulatory design, with a focus on how compliance obligations land on architecture rather than on paperwork.