The intersection of GDPR and AI training has become one of the most active areas of regulatory attention in 2025 and 2026. Organisations training AI models on data that includes personal information, and organisations deploying AI systems that generate outputs referencing identifiable individuals, face overlapping obligations that require careful analysis.

Does GDPR apply to AI training data?

Yes. Where training data contains personal data – defined as any information relating to an identified or identifiable natural person – GDPR applies in full to the collection, processing, and use of that data for training purposes. The fact that training may result in a model rather than a dataset does not exempt the training process from data protection obligations.

The CNIL, EDPB, and several national supervisory authorities have confirmed this position through guidance and enforcement. The Italian Garante's enforcement action against ChatGPT in 2023, based on the absence of a lawful basis for processing personal data in the training corpus, was the most prominent early signal of regulatory intent.

What legal basis applies to AI training on personal data?

Consent, legitimate interest, and public interest are the legal bases most commonly assessed in this context.

Consent as a legal basis faces significant practical obstacles: it must be freely given, specific, informed, and revocable. Obtaining valid consent from every individual whose personal data appears in a training corpus is generally not feasible for large-scale datasets.

Legitimate interest is the most commonly invoked basis, but it requires completion of a three-part test. The EDPB has noted that the balancing test is particularly demanding where training data was collected in a different context – for example, where social media posts are used to train a language model – because data subjects would not reasonably have expected their data to be used for that purpose.

Public interest applies primarily to research and academic contexts. Commercial AI development is unlikely to satisfy the public interest basis absent specific statutory authorisation.

What obligations apply where AI systems generate outputs about individuals?

Where an AI system generates outputs that constitute personal data – for example, a system that produces biographical summaries of named individuals – data protection obligations apply to both the output and the processing underlying it. Data subjects retain rights including the right of access, rectification, and objection.

The right not to be subject to solely automated decision-making under Article 22 applies where an AI system produces outputs that produce legal or similarly significant effects on individuals, without human involvement in the decision.

Transparency obligations under Articles 13 and 14 require that individuals are informed where their personal data has been processed to train an AI system, subject to the exception in Article 14(5)(b) where providing such information would involve disproportionate effort.

How does the EU AI Act interact with these GDPR obligations?

The two frameworks are additive. High-risk AI systems that process personal data must comply with both GDPR and the AI Act. The AI Act's data governance requirements – including quality criteria for training data and bias testing – overlap with GDPR's data minimisation and accuracy principles. A DPIA is likely to be required for AI systems that meet the nine EDPB criteria, which many AI training and deployment scenarios will.

What this means for your organisation

  • If your organisation trains AI models on data that includes personal information, GDPR applies to that training process. The fact that the output is a model rather than a dataset does not exempt the training from data protection obligations.
  • Legitimate interest is the most commonly used legal basis for AI training but it requires a completed, documented three-part test for each training activity. A generic legitimate interest claim will not withstand supervisory authority scrutiny.
  • AI systems that generate outputs about individuals – biographical summaries, risk scores, behavioural predictions – are subject to data protection obligations in respect of those outputs, including data subject rights of access and objection.
  • The GDPR and EU AI Act impose overlapping but distinct obligations on AI systems processing personal data. Compliance with one does not satisfy the other.

What you should do now

  1. Inventory every AI system your organisation uses that was trained on or processes personal data.
  2. For each training dataset, confirm the legal basis for processing and document the analysis – especially where legitimate interest is the basis.
  3. Assess whether any AI system generates outputs that constitute personal data and, if so, confirm data subject rights processes are in place.
  4. Determine whether a DPIA is required for high-risk AI processing activities, applying the nine EDPB criteria.
  5. Review your privacy notices to confirm they accurately describe AI-related processing and the legal bases relied upon.

How Priventia helps

Priventia's AI Systems Inventory module maps AI training and deployment obligations, linking GDPR and EU AI Act requirements in a single assessment. The platform detects DPIA triggers for AI processing activities and integrates AI governance with the broader compliance programme.