Executive conclusions
GDPR remains the gateway. There is no general “AI training” exception. Collection, corpus construction, pre-training, fine-tuning, evaluation, inference, logging, human decision-making and model improvement must each have a defined purpose, legal basis and data-protection design.
AI Act labels do not allocate GDPR responsibility. An AI Act provider can be a controller for pre-training, a processor for a customer’s prompts, and a separate controller for security telemetry or model improvement. A deployer is commonly—but not invariably—the controller for operational use.
Publicly accessible data remain protected. A webpage, social-media post, public register entry or photograph does not cease to be personal data merely because it can be accessed without authentication. “Manifestly made public” is a narrow special-category exception, not a general lawful basis.
The AI Act generally does not legalise personal-data processing. Article 2(7) preserves the GDPR, ePrivacy rules, the Law Enforcement Directive and the EU-institutions data-protection regime. The material exceptions are narrow: Article 10(5) for strictly necessary bias detection/correction in high-risk systems and Article 59 for tightly controlled sandbox processing in substantial public-interest projects.
High-risk and Article 22 analyses are independent. A system may be high-risk under the AI Act without making a solely automated significant decision, and a low-risk or transparency-only system may still trigger GDPR Article 22. A score can itself be a covered decision where the downstream decision-maker relies on it decisively.
Model anonymity is a fact question, not a label. A model is not anonymous merely because names were removed from the training set. Extraction, memorisation, singling-out, membership inference, linkability and the reasonably likely means available to the controller or another person must be assessed.
Special-category processing creates a structural tension. Bias testing may require protected-attribute data, while GDPR Article 9 restricts that processing. AI Act Article 10(5) is a narrow bridge for providers of high-risk systems; it is not a general permission for training, profiling or deployment.
AI transparency has several distinct layers. AI Act interaction or deepfake disclosure, provider-to-deployer technical information, GDPR privacy notices, Article 15/22 logic information, and AI Act Article 86 explanations answer different questions. One notice cannot be assumed to satisfy all of them.
High-risk delays do not create a compliance holiday. At the cut-off date, most AI Act provisions apply, while the amended timetable places the substantive Annex III high-risk duties on 2 December 2027 and product-embedded Article 6(1) duties on 2 August 2028. GDPR, equality, consumer, sectoral and already-applicable AI Act rules continue to govern now.
Enforcement can be parallel. Data-protection authorities, AI Act market-surveillance authorities, the AI Office, sector regulators, equality bodies, consumer authorities and courts can examine the same deployment through different legal tests. Contracts should allocate evidence and cooperation, not merely liability.
1. The legal architecture: two cumulative compliance stacks
1.1 The GDPR regulates data operations; the AI Act regulates practices, systems and models
The GDPR is technology-neutral. It applies whenever an entity within its territorial scope processes personal data, whether the processing uses a rules engine, statistical model, neural network, large language model, retrieval system or human workflow. The operative unit is the processing operation: obtaining data, labelling, structuring a corpus, training, validating, testing, embedding, retrieving, generating, scoring, retaining, disclosing, monitoring or deleting.
The AI Act asks different questions. It prohibits specified practices; classifies certain systems as high-risk by intended purpose; regulates general-purpose AI models; imposes transparency duties for direct interaction and synthetic content; and establishes provider, deployer, importer, distributor and product-manufacturer obligations. Its Article 2(7) expressly preserves EU data-protection and communications-privacy law, subject to the narrow mechanisms in Articles 10(5) and 59. The consequence is a dual test: the data processing must be lawful under data-protection law, and the AI practice or system must be lawful under the AI Act.
The two statutes can diverge. A factory-optimisation model may fall under product-safety or high-risk AI rules while using no personal data. Conversely, a customer-support chatbot may be outside the high-risk categories but still process names, account histories and communications under the full GDPR. “Not high-risk under the AI Act” is therefore not a privacy conclusion; “GDPR-compliant” is not an AI Act classification.
1.2 Actor labels do not map one-to-one
AI ecosystem position: Foundation-model developer / AI Act provider
- Typical GDPR position: Controller for source acquisition, corpus design and pre-training.
- Qualification: May act as processor for customer inference where it uses prompts only on documented instructions; may be a separate controller for abuse monitoring or model improvement.
AI ecosystem position: Enterprise customer / AI Act deployer
- Typical GDPR position: Usually controller for employee, applicant, customer or patient use.
- Qualification: May become joint controller where it co-determines a shared feedback/training purpose; may become an AI Act provider after rebranding, substantial modification or changed intended purpose.
AI ecosystem position: Application integrator
- Typical GDPR position: Processor, sub-processor, controller or joint controller depending design authority and reuse.
- Qualification: Integration decisions about features, retention, scoring thresholds and data reuse can be “essential means.”
AI ecosystem position: Cloud or hosting provider
- Typical GDPR position: Usually processor/sub-processor.
- Qualification: Independent security analytics, account intelligence or cross-customer product improvement can create a separate controller role.
AI ecosystem position: Data broker / corpus supplier
- Typical GDPR position: Usually controller for collection and disclosure.
- Qualification: The recipient needs its own lawful basis and due diligence; contractual warranties do not cure unlawful source processing.
AI ecosystem position: Human reviewer / employer / bank / authority
- Typical GDPR position: Part of the controller’s organisation or a separate controller.
- Qualification: A nominal human step does not prevent Article 22 where review is not genuine, informed and capable of changing the outcome.
A single statement that “the vendor is the processor” is usually inadequate. The agreement should allocate roles per purpose: service inference, prompt retention, safety review, fraud/abuse detection, support, benchmarking, telemetry, fine-tuning and model improvement. Any independent use requires a separate controller analysis and notice.
1.3 Personal data can exist at every AI lifecycle layer
Layer: Source data
- Examples: Web pages, customer records, CVs, call recordings, images, sensor data, public registers.
- Frequent misconception: “Public” or licensed data are not necessarily anonymous or lawfully reusable.
Layer: Labels and annotations
- Examples: Health condition, sentiment, ethnicity proxy, fraud marker, quality score.
- Frequent misconception: A label created by an annotator or model is still personal data if it relates to a person.
Layer: Derived features
- Examples: Embeddings, vectors, clusters, inferred interests, risk scores, facial templates.
- Frequent misconception: Pseudonymisation or numerical form does not remove GDPR protection.
Layer: Model parameters
- Examples: Weights or statistical representations that may memorise or permit extraction.
- Frequent misconception: Not every model is personal data, but anonymity requires evidence, not assertion.
Layer: Prompts and context
- Examples: Names, account data, legal files, medical histories, retrieved documents.
- Frequent misconception: A “no training” setting does not remove inference, logging, support or transfer processing.
Layer: Outputs
- Examples: Recommendations, rankings, summaries, generated biographies, probability scores.
- Frequent misconception: False or hallucinated information can still be personal data and engage accuracy and remedy duties.
Layer: Operational evidence
- Examples: Logs, human overrides, explanations, incident records, feedback.
- Frequent misconception: Mandatory accountability records need a legal basis, access controls and retention limits.
2. GDPR treatment across the AI lifecycle
2.1 Personal data, pseudonymisation and model anonymity
The threshold is whether information relates to an identified or identifiable natural person. Direct identifiers are not required. Identifiability can arise from linkage, singling out, inference, unique behaviour, an account identifier, a biometric template, or practical access to auxiliary data. Pseudonymised records and embeddings remain personal data where re-linking is reasonably possible.
For trained models, the correct question is not whether the model was intended to store records, but whether personal data can reasonably be extracted, reproduced, inferred or linked to a person, and whether the model can be used to single out or affect an individual. EDPB Opinion 28/2024 treats anonymity as case-specific and demanding: both identification and extraction risks must be insignificant in light of means reasonably likely to be used. The assessment should cover memorisation testing, membership-inference and model-inversion risk, prompt-based extraction, access level, attacker capability, auxiliary datasets, model release format and downstream controls.
An anonymous model can fall outside the GDPR as an object, while its deployment still processes personal data in prompts, retrieval records or outputs. Conversely, deleting the raw corpus does not make a model anonymous if it retains extractable personal information. Anonymisation also does not erase liability for unlawful collection or training that occurred earlier.
2.2 Purpose limitation and lawful basis must be assessed operation by operation
A controller should define a distinct purpose for each material operation. “Developing AI,” “innovation” or “improving the model” is generally too indeterminate to demonstrate purpose limitation, necessity or transparency. A purpose valid for delivering an answer to a user does not automatically justify retaining prompts, allowing reviewers to read them, using them to train another model, or combining them with data from other services.
Lifecycle operation: Source collection / scraping
- Possible Article 6 route: Legitimate interests, public task, legal obligation or consent depending context.
- Legal pressure points: Reasonable expectations, source restrictions, scale, Article 14 notice, children, special categories, objection and minimisation.
Lifecycle operation: Pre-training / fine-tuning
- Possible Article 6 route: Often legitimate interests is asserted; consent or statutory research/public-task bases may apply in narrower settings.
- Legal pressure points: Necessity is not established merely because more data improve performance; separate Article 9/10 gateways; compatibility of reuse.
Lifecycle operation: Service inference
- Possible Article 6 route: Contract necessity, legitimate interests, public task or legal obligation.
- Legal pressure points: Contract is limited to what is objectively necessary for the requested service; sensitive inputs need Article 9 condition.
Lifecycle operation: Human decision using output
- Possible Article 6 route: Basis for underlying decision: contract, law, public task or legitimate interests.
- Legal pressure points: Article 22, sector law, accuracy, discrimination, explanation, contestability.
Lifecycle operation: Security / abuse monitoring
- Possible Article 6 route: Legitimate interests or legal obligation.
- Legal pressure points: Scope, retention, access, cross-customer correlation and secondary use.
Lifecycle operation: Feedback / model improvement
- Possible Article 6 route: Separate consent or legitimate-interests analysis; sometimes legal obligation for safety monitoring.
- Legal pressure points: Do not bundle into service necessity; honour opt-outs and notices; distinguish mandatory safety evidence from product optimisation.
Lifecycle operation: Mandatory AI Act logs / incidents
- Possible Article 6 route: Article 6(1)(c) to the extent strictly necessary to comply with an applicable legal obligation.
- Legal pressure points: The obligation cannot bootstrap unrelated training or indefinite retention; minimise and segregate evidence.
Legitimate interests require a concrete interest, necessity and balancing. Commercial interests can qualify, but the controller must test less intrusive alternatives and weigh source context, scale, sensitivity, foreseeability, data-subject vulnerability, consequences and safeguards. Consent must be freely given, specific, informed and withdrawable; it is often unsuitable in employment, essential services and large-scale web collection. Contract necessity is interpreted narrowly: a processing operation is not necessary merely because it appears in terms and conditions or improves profitability.
Further use of existing customer, employee or public-sector data for training requires a compatibility analysis under Articles 5(1)(b) and 6(4), unless a new consent or specific legal basis applies. Relevant factors include the relationship between purposes, collection context, data nature, consequences and safeguards. Scientific-research provisions do not create a blanket AI R&D exemption; Article 89 safeguards and the applicable Union or Member State law remain necessary.
2.3 Publicly accessible and “manifestly public” data
Public accessibility affects expectations and balancing but does not remove GDPR protection. Search-engine visibility, a public social-media profile, an online directory or a public register does not itself authorise bulk collection and model training. The controller still needs an Article 6 basis, must assess purpose compatibility where applicable, and must address transparency, objection, minimisation, retention, security and data-subject rights.
For special-category data, Article 9(2)(e) applies only where the data subject has manifestly made the data public. The CJEU has rejected broad assumptions that any public statement or online activity opens all related sensitive data to unrestricted processing. The exception is intentional, item-specific and does not replace the Article 6 analysis. It also does not authorise processing sensitive inferences that the person did not themselves manifestly disclose.
2.4 Special categories, biometric data and criminal-offence data
Special-category processing requires both an Article 6 legal basis and an Article 9(2) condition. Relevant categories include health, racial or ethnic origin, political opinions, religious or philosophical beliefs, trade-union membership, genetic data, biometric data processed for unique identification, and sex-life or sexual-orientation data. A photograph is personal data but becomes special-category biometric data when subjected to specific technical processing for unique identification or authentication. Emotion, personality and socio-economic inferences may be highly intrusive even where they do not fall neatly within Article 9.
AI systems can create special-category data by inference. Deliberately deriving an ethnicity, health, political or sexual-orientation attribute for targeting, ranking or monitoring should be treated as Article 9 processing. Proxy variables can also produce unlawful discrimination even when the controller never records the protected characteristic.
Article 10 data about criminal convictions and offences may be processed only under official authority or where Union or Member State law provides appropriate safeguards. Private fraud or integrity systems must distinguish ordinary risk flags from allegations that effectively constitute offence data, and must examine national law.
2.5 Transparency: privacy notice, AI disclosure and data provenance
GDPR Articles 13 and 14 require information about the controller, purposes, legal bases, recipients, transfers, retention, rights and—where relevant—automated decision-making. Where personal data are not obtained from the individual, source and category information are central. Article 14(5)(b) can excuse individual notice where provision is impossible or involves disproportionate effort, but only after a documented assessment and appropriate measures, including making the information publicly available. It is not a general exemption for web-scale datasets.
The AI Act adds different disclosures. Article 50 generally requires persons to be informed when directly interacting with AI unless obvious in context; deployers of emotion-recognition or biometric-categorisation systems must inform exposed persons; and deepfakes or certain synthetic content require disclosure. These notices do not state why personal data are processed, which basis applies or how rights can be exercised, so they cannot replace a GDPR notice.
Transparency instrument: GDPR arts 13–14 notice
- Audience: Data subject
- What it must achieve: Explain personal-data processing, basis, sources, recipients, retention and rights.
- Not a substitute for: AI-use disclosure alone.
Transparency instrument: GDPR arts 13(2)(f), 14(2)(g), 15(1)(h)
- Audience: Data subject
- What it must achieve: Meaningful information about automated logic, significance and envisaged consequences.
- Not a substitute for: Source-code disclosure or generic marketing description.
Transparency instrument: AI Act art 13
- Audience: High-risk deployer
- What it must achieve: Provider instructions, capabilities, limits, oversight and performance information.
- Not a substitute for: Direct data-subject notice.
Transparency instrument: AI Act art 50
- Audience: Person interacting with or exposed to specified AI/synthetic content
- What it must achieve: Make AI involvement or manipulation apparent.
- Not a substitute for: Lawful basis, Article 22 safeguards or privacy notice.
Transparency instrument: AI Act art 86
- Audience: Affected person in a qualifying high-risk decision
- What it must achieve: Explain the AI system’s role and main elements of the decision, once applicable.
- Not a substitute for: Broader GDPR access/objection/rectification rights.
Transparency instrument: GPAI training-content summary
- Audience: Public and downstream ecosystem
- What it must achieve: Provide prescribed high-level transparency about training content.
- Not a substitute for: Individual Article 14 notice or proof of lawful collection.
2.6 Accuracy, rectification, erasure and model-level remedies
The accuracy principle applies to personal data in datasets, labels, scores and outputs, having regard to the processing purpose. A generative model’s false statement about a person can be personal data even though it was probabilistically generated. Controllers must provide workable routes to investigate, annotate, correct, suppress or prevent recurrence. The GDPR does not prescribe a single technical remedy such as retraining or machine unlearning; the chosen method must be effective in the circumstances and must address copies, retrieval stores, caches and downstream disclosures where required.
Erasure requests require a layer-specific response. Deleting a source record or prompt may be straightforward; removing influence from a trained model can be technically difficult. Difficulty is relevant to proportionality and method, but not a freestanding exemption. A controller should be able to explain whether the model is personal data, how the person can be located, whether extraction is possible, what suppression controls exist, and why the selected remedy is effective.
Trade secrets and intellectual property can limit the form of disclosure but do not justify a blanket refusal. In C‑203/22, the CJEU required intelligible information enabling the person to understand and challenge automated processing; an algorithm alone may be unintelligible, while a competent authority or court can balance protected secrets against rights.
2.7 Profiling and Article 22 automated decisions
Profiling is automated processing used to evaluate personal aspects. Article 22 adds a qualified right not to be subject to a decision based solely on automated processing, including profiling, that produces legal or similarly significant effects. The exceptions are necessity for a contract, authorisation by law with safeguards, or explicit consent. Even then, appropriate measures ordinarily include meaningful human intervention, an opportunity to express a view and a right to contest.
The CJEU’s SCHUFA judgment confirms that a score generated by one entity can itself be a decision where a third party relies on it strongly in deciding whether to contract. Splitting the pipeline between a scoring vendor and a bank, employer or insurer does not avoid Article 22. A nominal reviewer does not remove the “solely automated” character if the person lacks authority, time, information or genuine discretion to depart from the model.
AI Act high-risk classification and Article 22 overlap but are not coextensive. Recruitment ranking may be high-risk even where a panel makes the final decision. A customer-support bot can be non-high-risk yet trigger Article 22 if it autonomously terminates an account or denies a benefit. Both analyses should therefore appear as separate lines in the DPIA.
2.8 DPIA, privacy by design, security and international transfers
A DPIA is required where processing is likely to result in high risk. Statutory examples include systematic and extensive evaluation based on automated processing on which significant decisions are based, large-scale processing of special categories or offence data, and large-scale systematic monitoring of publicly accessible areas. National supervisory-authority lists add triggers. Many employment, credit, biometric, health, public-sector and behavioural-advertising AI uses therefore require a DPIA before deployment, irrespective of the AI Act’s delayed high-risk timetable.
Privacy by design should address more than de-identification. Relevant controls include source allow/deny lists, field-level minimisation, tokenisation or pseudonymisation, role-based access, tenant isolation, retrieval permission propagation, output filters, red-team testing for extraction, prompt-injection controls, secure evaluation environments, model-release controls, retention schedules, feedback segregation and documented deletion or suppression pathways.
Article 32 security analysis should include AI-specific attack and failure modes: model inversion, membership inference, training-data extraction, prompt leakage, cross-tenant retrieval, malicious fine-tuning, data poisoning, model theft and unsafe agent actions. AI Act robustness and cybersecurity duties complement, but do not replace, the GDPR risk-based test.
Sending prompts, logs, support data or model-access data to a non-EEA vendor, or allowing remote access from a third country, can be an international transfer. The controller must identify the transfer mechanism, conduct any required transfer-impact assessment, implement supplementary measures and ensure sub-processor transparency. A region selector is evidence, not proof, that all support, telemetry and failover processing remain in-region.
2.9 Processor, joint-controller and vendor-contract consequences
- Purpose-by-purpose role schedule. State who controls inference, safety review, abuse detection, support, analytics, fine-tuning and improvement.
- Instruction and reuse boundary. A processor may not convert customer prompts into its own training asset without a separate role, basis and notice.
- Sub-processors and locations. List model hosts, moderation services, support vendors and evaluation providers; control changes and transfers.
- Retention and deletion. Specify raw prompts, derived telemetry, safety samples, backups, embeddings, fine-tuned weights and account closure.
- Rights assistance. Require search, export, correction, suppression, deletion and explanation support with defined service levels.
- AI Act evidence. Obtain instructions, intended-purpose limits, performance by relevant groups, logging interfaces, human-oversight design, incident cooperation and conformity information.
- Change control. Treat model-version changes, new training uses, new data sources and changed decision thresholds as legal-review triggers.
- Security and incident routing. Coordinate GDPR breach, AI Act serious-incident, NIS2/DORA and sector notifications without assuming one report satisfies another.
3. The AI Act overlay
3.1 Application timetable at the cut-off date
At 18 August 2026, the prohibited-practice and AI-literacy provisions have applied since 2 February 2025; the general-purpose AI, governance and penalty framework has applied in material part since 2 August 2025, subject to model-specific transitions; and most remaining provisions, including Article 50 transparency, apply from 2 August 2026. The 2026 amendment (Regulation (EU) 2026/1744, in force since 27 July 2026) postponed the substantive obligations for Annex III high-risk systems to 2 December 2027 and for Article 6(1) product-embedded high-risk systems to 2 August 2028.
Date: 2 February 2025
- Material effect: Prohibited practices and AI literacy apply.
- Personal-data consequence: Prohibited facial-image scraping, specified biometric/emotion uses and manipulative practices cannot be justified by a GDPR basis.
Date: 2 August 2025
- Material effect: GPAI/governance/penalty provisions apply in material part, with transitional rules.
- Personal-data consequence: GPAI documentation, copyright policy, training-content summary and systemic-risk duties coexist with GDPR training lawfulness.
Date: 2 August 2026
- Material effect: Most remaining AI Act provisions apply, including Article 50 transparency.
- Personal-data consequence: AI disclosure is now an additional operational requirement; it does not replace privacy information.
Date: 2 December 2026
- Material effect: Newly inserted Article 5 prohibitions on AI systems that generate or manipulate non-consensual intimate images or child sexual abuse material apply (Regulation (EU) 2026/1744).
- Personal-data consequence: A GDPR legal basis or a disclosure label cannot cure this prohibited practice; realistic intimate-image outputs depicting identifiable persons are directly covered.
Date: 2 December 2027
- Material effect: Substantive Annex III high-risk system duties under the amended timetable.
- Personal-data consequence: Recruitment, credit, specified public-service and other Annex III systems require the full provider/deployer stack from this date, while GDPR applies already.
Date: 2 August 2028
- Material effect: Substantive Article 6(1) product-safety-component high-risk duties.
- Personal-data consequence: Medical-device and other regulated-product AI transitions need sector-specific planning; GDPR and product law apply in the interim.
3.2 Prohibited AI practices involving personal data
A valid GDPR legal basis cannot rescue an AI Act prohibited practice. Relevant prohibitions include manipulative or deceptive techniques causing significant harm; exploitation of vulnerabilities linked to age, disability or socio-economic situation; specified social scoring; criminal-risk assessment based solely on profiling or personality; untargeted scraping of internet or CCTV facial images to create or expand facial-recognition databases; specified biometric categorisation by highly sensitive traits; emotion inference in workplaces and educational institutions except narrow medical or safety uses; and real-time remote biometric identification by law enforcement in publicly accessible spaces outside tightly drawn exceptions. From 2 December 2026, Regulation (EU) 2026/1744 adds a further prohibition on AI systems that generate or manipulate realistic non-consensual intimate images, videos or audio of identifiable persons, or material depicting child sexual abuse.
The facial-image rule is especially significant for model developers. It is practice-based, not dependent on whether each image was public, licensed or accompanied by a GDPR basis. Targeted, lawful use of a limited image set is not automatically permitted; it still requires GDPR Article 6/9 analysis, AI Act classification and any national biometric law.
3.3 High-risk classification and data-protection relevance
High-risk classification turns primarily on intended purpose and sector. Annex III covers specified biometric uses; critical infrastructure; education and vocational training; recruitment, worker management and access to self-employment; access to essential private and public services, including creditworthiness and life/health insurance risk or pricing; law enforcement; migration and border control; and administration of justice and democratic processes. A narrow exception exists where a listed system does not pose a significant risk and performs only limited preparatory or procedural functions, but an Annex III system that profiles natural persons remains high-risk.
The classification matters to data protection because the required risk management, data governance, technical documentation, logging, transparency to deployers, human oversight, accuracy, robustness and cybersecurity can supply evidence for GDPR accountability. It does not itself establish Article 6 or 9 lawfulness. Conversely, a GDPR DPIA does not constitute AI Act conformity assessment.
3.4 Provider and deployer duties that intersect with personal data
AI Act duty: Risk management (art 9)
- Data-protection intersection: Fundamental-rights and foreseeable-misuse risks can include privacy, discrimination and security.
- Operational implication: Maintain a single risk register with separate legal tests and owners.
AI Act duty: Data governance (art 10)
- Data-protection intersection: Training/validation/test data quality, provenance, relevance, representativeness, errors and bias.
- Operational implication: Reconcile representativeness with GDPR minimisation and Article 9 restrictions.
AI Act duty: Technical documentation and logs (arts 11–12)
- Data-protection intersection: Records may contain personal data and support accountability.
- Operational implication: Define legal basis, access, retention and disclosure rules for evidence.
AI Act duty: Information to deployers (art 13)
- Data-protection intersection: Needed for controller DPIA, Article 22 explanation, security and rights response.
- Operational implication: Contract for sufficiently granular data and model information.
AI Act duty: Human oversight (art 14)
- Data-protection intersection: Supports, but does not automatically satisfy, meaningful human intervention under Article 22.
- Operational implication: Give reviewers competence, authority, time, uncertainty information and override capability.
AI Act duty: Accuracy, robustness, cybersecurity (art 15)
- Data-protection intersection: Overlaps GDPR accuracy/security but has system-performance focus.
- Operational implication: Test both individual-data accuracy and population-level performance/failure modes.
AI Act duty: Deployer duties (art 26)
- Data-protection intersection: Input data relevance, instruction compliance, monitoring, logs and worker information.
- Operational implication: Operational controller remains responsible for lawful input and use.
AI Act duty: FRIA (art 27)
- Data-protection intersection: Broader than a DPIA; includes equality, due process, safety and other Charter rights.
- Operational implication: Combine documentation where efficient, but preserve both statutory scopes.
3.5 Article 10(5): narrow special-category processing for bias controls
Article 10(5) is the AI Act’s specific route for providers of high-risk systems to process special-category data where strictly necessary to ensure bias detection and correction. It is not a general training basis and does not cover unrelated profiling, product personalisation, marketing or deployer decision-making. The provider should still identify an Article 6 basis; Article 10(5) is best treated as the specific Article 9 gateway for the defined bias-control purpose.
The conditions require a necessity showing and safeguards. The provider should document why the objective cannot be effectively achieved with synthetic or anonymised data, isolate the protected attributes, prevent reuse and onward access, apply state-of-the-art security and strict access controls, keep records of the necessity and delete the special-category data when the bias-control purpose is complete or retention is no longer justified. A deployer conducting its own equality audit needs a separate Article 9 route, often under employment or substantial-public-interest law at Member State level.
3.6 Article 59 regulatory sandboxes
Article 59 permits further processing in an AI regulatory sandbox of personal data lawfully collected for other purposes for development, training and testing of specified public-interest AI systems, subject to stringent conditions. The route is limited to the sandbox, substantial public-interest objectives and supervised safeguards; it is not a general commercial re-use exception. The project must avoid using the sandbox processing to take measures or decisions affecting participants, limit access and disclosure, monitor risks, retain auditable records and delete data when the authorised work ends.
The sandbox route does not displace the rest of the GDPR. Transparency, security, data-subject rights, governance, any required DPIA and the limits of the sandbox plan remain relevant. A project outside an approved sandbox cannot invoke Article 59 by analogy.
3.7 General-purpose AI and training data
GPAI providers must maintain technical documentation, provide information to downstream system providers, implement a copyright-compliance policy and publish a sufficiently detailed summary of training content using the prescribed framework. Providers of models with systemic risk have additional evaluation, risk-management, incident and cybersecurity obligations.
These duties improve ecosystem transparency but do not create a GDPR basis for acquiring or training on personal data. A copyright text-and-data-mining exception and an AI Act training-content summary address different protected interests. A dataset can be copyright-lawful yet violate the GDPR, or privacy-lawful yet infringe copyright or database rights.
Open-source status also does not create a GDPR exemption. AI Act relief for certain freely and openly licensed models or components is limited and does not apply uniformly to prohibited practices, systemic-risk models or downstream high-risk systems. A public model release can increase the reasonably likely means of extraction and therefore weaken an anonymity claim.
3.8 Fundamental-rights impact assessment and explanation
Article 27 requires a fundamental-rights impact assessment for specified deployers of high-risk systems, notably bodies governed by public law, private entities providing public services, and deployers of the covered credit and life/health insurance systems. The FRIA examines affected groups, duration and frequency, specific fundamental-rights risks, human oversight and mitigation. It can be integrated with a DPIA, but the FRIA is broader and the DPIA remains legally distinct.
Article 86 provides an explanation right for persons affected by specified high-risk decisions once the relevant provisions apply. It focuses on the role of the high-risk system and the main elements of the decision. GDPR access, transparency, rectification, objection and Article 22 safeguards may apply independently and earlier.
4. Use-case analysis
4.1 Web-scale pre-training of a foundation model
The developer ordinarily determines the sources, corpus design, exclusions, model objective and training method and is therefore a controller. Personal data can appear in webpages, forums, code repositories, books, images, public records and metadata. Removing obvious identifiers does not necessarily anonymise context or prevent memorisation.
Legitimate interests may be argued, but no categorical presumption exists. The developer must articulate the interest, show that the particular data and scale are necessary, and balance impact and expectations. High-risk factors include children, special categories, intimate content, doxxing, stale or inaccurate information, paywalled or access-controlled sources, opt-out signals, facial images, broad retention and public release of model weights. Article 14 transparency and rights channels need a practical design.
The AI Act adds GPAI documentation, copyright policy and training-content summary duties, and may add systemic-risk controls. It separately prohibits untargeted scraping of facial images to create or expand facial-recognition databases. Copyright text-and-data-mining permissions do not answer GDPR lawfulness.
- Minimum controls. Source taxonomy and provenance; high-risk source exclusions; special-category and child-data filters; facial-image rule; rights-reservation handling; Article 14 public notice; objections and removals; memorisation/extraction testing; model-release risk assessment; retention and deletion evidence.
- Residual uncertainty. Whether the trained model is anonymous and whether large-scale legitimate interests survive balancing are fact-specific and regulator-dependent.
4.2 Enterprise LLM API or copilot with customer data
Where an enterprise sends prompts solely to obtain a response and the vendor uses them only on documented instructions, the enterprise is usually controller and the vendor processor for inference. That classification changes if the vendor uses content for its own model improvement, cross-customer analytics, product research or independent account intelligence.
The enterprise should prohibit unnecessary personal or confidential data, use technical data-loss-prevention controls, segregate tenants, restrict support access, define prompt/log retention and verify sub-processors and transfer locations. A “training off” toggle must be tested against safety sampling, abuse monitoring, human review, backups and telemetry.
The AI Act classification depends on the copilot’s intended purpose. A drafting assistant may be transparency-only or non-high-risk; the same model integrated into recruitment, credit, health or public-benefit decisions can become high-risk. A user-facing chatbot generally needs Article 50 disclosure.
4.3 Retrieval-augmented generation over internal records
RAG does not avoid GDPR because documents are not placed in model weights. Chunking, indexing, embeddings, vector search, query logs and generated outputs are processing. Access entitlements from the source repository must propagate into retrieval; otherwise the system can create a new unauthorised disclosure channel.
The controller should test whether the original purpose permits AI-assisted retrieval, minimise indexed fields, separate sensitive collections, apply document-level permissions, manage stale records and define deletion propagation from source to chunks, vectors, caches and evaluation logs. Generated answers should cite source records where feasible and distinguish retrieved facts from model inference.
A RAG assistant becomes legally different when used to decide employment, benefits, credit or care. The same technical stack can move from ordinary workplace tooling to Annex III high-risk use and Article 22 exposure solely because the intended purpose and decision pathway change.
4.4 Recruitment, worker management and workplace monitoring
AI used to advertise jobs selectively, screen CVs, rank applicants, analyse interviews, allocate tasks, evaluate performance or recommend termination is a high-impact GDPR use and falls within Annex III high-risk categories where specified. The high-risk duties are delayed under the current timetable, but GDPR, equality, labour and already-applicable AI Act prohibitions govern immediately.
Employee or candidate consent is usually weak because of power imbalance and lack of a genuine alternative. The employer must identify another Article 6 basis and, for special-category or biometric data, an Article 9 condition grounded in Union or Member State law. Automated rejection or effective determinative ranking may trigger Article 22. Human reviewers require authority, training, sufficient time, access to relevant evidence and a recorded ability to depart from the score.
Emotion recognition in the workplace is prohibited except narrow medical or safety uses. Personality, voice, facial or behavioural analysis also creates accuracy, discrimination and scientific-validity concerns. Works-council, information-and-consultation and national employee-monitoring rules can be stricter than the GDPR. For digital labour platforms, the Platform Work Directive adds specific limits and transparency/human-review rules as Member States transpose it.
- Evidence to preserve. Job-related validity, subgroup performance, feature relevance, adverse-impact testing, accessibility, reviewer overrides, reasons for decisions, candidate notices and appeal outcomes.
- Stop condition. Do not deploy workplace emotion inference, or a system whose vendor cannot support meaningful explanation and contestation.
4.5 Creditworthiness, insurance and fraud controls
AI used to evaluate creditworthiness or establish a credit score is an Annex III high-risk use, subject to the limited statutory exclusions. AI used for risk assessment or pricing in life and health insurance is also covered. Specified deployers must complete a FRIA when the relevant high-risk provisions apply. Sectoral credit, insurance and equality rules continue independently.
SCHUFA is central: a vendor-generated score can be an Article 22 decision where the lender gives it a determining role. A lender therefore needs to understand the score’s factors, data sources, limits and uncertainty and cannot outsource the ability to explain or contest. C‑203/22 reinforces the need for intelligible, case-specific logic information.
Fraud detection may fall outside the specific Annex III creditworthiness category, but it remains subject to GDPR and can trigger Article 22 if it automatically blocks transactions, closes accounts or generates similarly significant effects. Legal-obligation and legitimate-interest bases must be distinguished; criminal-offence data and AML secrecy rules require national-law analysis. Bias against protected or proxy groups remains possible even without explicit special-category data.
4.6 Customer service, personalised offers, advertising and recommender systems
A service chatbot processing account data usually relies on contract necessity or legitimate interests for the conversation, with Article 50 AI disclosure and GDPR notice. Escalation to a human does not cure an unlawful basis or excessive retention. If the bot autonomously denies refunds, terminates service or changes access to an essential service, Article 22 and possibly high-risk classification must be reassessed.
Personalised advertising and recommender systems add ePrivacy and DSA layers. Access to or storage of information on a device often requires consent under national ePrivacy implementation. The DSA prohibits profiling-based advertising using special-category data, prohibits profiling-based advertising to known minors, prohibits specified dark patterns and requires recommender-system transparency. GDPR Article 21 gives an absolute objection to direct-marketing processing.
An AI-generated persuasion strategy can also engage the AI Act’s manipulation and vulnerability prohibitions, consumer-law rules on unfair or misleading practices, and equality law. The AI Act prohibition has a significant-harm threshold; consumer law can intervene below that threshold.
4.7 Biometrics, facial recognition and emotion systems
Biometric identification or authentication requires Article 6 and, where used for unique identification, Article 9. Necessity and proportionality are demanding because alternatives such as badges or device-based credentials may be less intrusive. Explicit consent can be invalid where access to work or an essential service is conditional and no equivalent alternative exists.
The AI Act prohibits untargeted facial-image scraping for recognition databases and specified biometric categorisation by sensitive traits. Remote biometric identification in publicly accessible spaces for law enforcement is subject to a specific prohibition and narrow exceptions. Other biometric systems can be high-risk and remain subject to national constitutional, police, employment and surveillance law.
Emotion-recognition claims deserve separate scrutiny. Workplace and education uses are generally prohibited. In other settings, scientific validity, demographic performance, transparency, fairness and the difference between observable expression and internal emotional state should be tested rather than assumed.
4.8 Healthcare, clinical decision support and medical devices
Health data require an Article 6 basis and an Article 9 condition, commonly healthcare provision, public health, research under applicable law, or explicit consent in suitable contexts. Patient consent to treatment is not automatically GDPR consent for training. Secondary use of clinical records for model development requires purpose, legal-basis and research-law analysis.
Clinical AI can be software as a medical device or a safety component under the Medical Devices or In Vitro Diagnostic Medical Devices Regulations. Product-embedded AI high-risk duties follow the amended 2 August 2028 timetable, while medical-device quality, clinical evaluation, vigilance, GDPR and professional standards apply independently. A model’s population-level accuracy does not answer individual-data accuracy, subgroup safety or clinical utility.
Human oversight must be clinically meaningful. The clinician needs relevant inputs, confidence and limitations, not a bare recommendation. Automation bias, alert fatigue, drift and off-label use belong in both safety and data-protection assessments.
4.9 Law enforcement, migration and border systems
Processing by competent authorities for prevention, investigation, detection or prosecution of criminal offences is generally governed by Directive (EU) 2016/680 and national implementing law rather than the GDPR. Other authority processing can remain under the GDPR or the EU-institutions regime. A precise purpose and institutional-role analysis is essential.
The AI Act prohibits or tightly restricts specified predictive-policing, biometric and emotion practices and classifies many law-enforcement, migration and border systems as high-risk. Legal authority, necessity, proportionality, data quality, independent oversight, logging, disclosure and effective remedies are central. A procurement contract cannot supply the legal authority for coercive public-sector processing.
4.10 Generative outputs about identifiable persons and deepfakes
A generated biography, accusation, health statement, profile, voice clone or image can be personal data where it relates to an identifiable person. The controller must address accuracy, source provenance, lawful disclosure, rectification and erasure. Defamation, personality rights, image and voice rights, copyright, consumer and criminal law may apply in parallel.
AI Act deepfake disclosure addresses deception but does not legalise use of the person’s data or likeness. From 2 December 2026, generating or manipulating realistic non-consensual intimate images of an identifiable person is itself a prohibited practice under Article 5 as amended by Regulation (EU) 2026/1744, independently of any disclosure. A disclosure label does not cure lack of Article 6/9 basis, infringement of personality rights or harmful publication. Platforms hosting user-generated synthetic content may also have DSA notice-and-action and risk-management obligations.
Operational remedies include source correction, retrieval suppression, output blocking for verified false claims, provenance markers, account sanctions, model patching or retraining where necessary, and notice to recipients. The correct remedy depends on the controller’s role and technical layer.
4.11 Bias testing and synthetic data
Synthetic data are useful but not automatically anonymous. If records can be linked back, reproduce rare combinations, preserve membership information or are generated from a small population, the GDPR can still apply. The controller should test singling out, linkability, inference and leakage rather than rely on the label “synthetic.”
For high-risk-system provider bias testing, Article 10(5) can permit strictly necessary special-category processing where synthetic or anonymised data cannot achieve the objective. Outside that route, equality monitoring may rely on national employment or substantial-public-interest law. Protected attributes used for an audit should be segregated from production decision features and governed to prevent adverse use.
4.12 Consolidated use-case matrix
Use case: Web pre-training
- GDPR / data law: Controller; public data still personal; Art 14; LI balancing; Art 9 filters; anonymity evidence.
- AI Act: GPAI duties; facial-image scraping ban; systemic-risk duties if relevant.
- Adjacent law: Copyright/TDM, database rights, source terms.
- Baseline risk: Critical
Use case: Enterprise LLM API
- GDPR / data law: Customer controller/vendor processor per purpose; Art 28; transfers; prompt minimisation.
- AI Act: Art 50 if user-facing; classification by intended use.
- Adjacent law: Confidentiality, privilege, sector outsourcing.
- Baseline risk: Medium–High
Use case: Internal RAG
- GDPR / data law: Embeddings and logs are personal; purpose compatibility; permissions; deletion propagation.
- AI Act: May become high-risk through decision use.
- Adjacent law: Cybersecurity, records management.
- Baseline risk: High if sensitive
Use case: Recruitment/HR
- GDPR / data law: DPIA; weak consent; Art 22; Art 9/proxies; candidate rights.
- AI Act: Annex III high-risk; workplace emotion ban.
- Adjacent law: Equality, labour, works councils, platform-work rules.
- Baseline risk: Critical
Use case: Credit/insurance
- GDPR / data law: Art 22; explanations; accuracy; Art 9/10; FRIA/DPIA.
- AI Act: Annex III high-risk; FRIA for specified deployers.
- Adjacent law: Credit, insurance, equality, consumer law.
- Baseline risk: Critical
Use case: Fraud/AML
- GDPR / data law: LI or legal obligation; Art 22 if blocking; offence data.
- AI Act: Classification depends on intended purpose; not automatically exempt.
- Adjacent law: AML, banking secrecy, DORA.
- Baseline risk: High
Use case: Chatbot/customer service
- GDPR / data law: Contract/LI; conversation notice; retention; escalation.
- AI Act: Art 50; high-risk only if used for covered decisions.
- Adjacent law: ePrivacy, consumer law.
- Baseline risk: Medium
Use case: Ads/recommenders
- GDPR / data law: Profiling, consent for tracking, Art 21, children/special data.
- AI Act: Manipulation/vulnerability rules; transparency.
- Adjacent law: DSA, ePrivacy, consumer law.
- Baseline risk: High
Use case: Biometrics/emotion
- GDPR / data law: Arts 6/9; necessity; bystanders; accuracy.
- AI Act: Multiple prohibitions; otherwise high-risk.
- Adjacent law: National surveillance/employment/constitutional law.
- Baseline risk: Critical
Use case: Clinical AI
- GDPR / data law: Health basis; DPIA; data quality; transfers; patient rights.
- AI Act: Product high-risk timetable; oversight and robustness.
- Adjacent law: MDR/IVDR, professional liability, health law.
- Baseline risk: Critical
Use case: Law enforcement/migration
- GDPR / data law: LED or GDPR/EUDPR; statutory authority; sensitive/offence data.
- AI Act: Prohibitions and Annex III high-risk.
- Adjacent law: Constitutional, criminal-procedure, asylum law.
- Baseline risk: Critical
Use case: Deepfakes/person outputs
- GDPR / data law: Accuracy, lawful disclosure, rectification/erasure.
- AI Act: Art 50 disclosure; provider marking duties.
- Adjacent law: Defamation, personality, copyright, DSA.
- Baseline risk: High
5. Additional legal angles
5.1 Equality and fundamental-rights law
EU and national equality rules prohibit direct and indirect discrimination in employment, goods and services, social protection and other covered fields. AI can discriminate through historical labels, measurement error, selection bias, proxies, feedback loops, accessibility failures or differential error rates. An AI Act conformity assessment is not a defence to discrimination; a GDPR legal basis is not a defence either.
The Charter protects private life, personal data, non-discrimination and effective judicial protection. Public authorities must also satisfy legality, necessity and proportionality. The AI Act FRIA operationalises part of this analysis for specified high-risk deployers, but courts and equality bodies can apply broader standards.
5.2 Employment and platform-work law
National employment law can regulate monitoring, consultation, collective bargaining, workplace surveillance, dismissal evidence and occupational safety. These rules are often more specific than the GDPR and can require works-council agreement or employee-representative consultation before an AI deployment.
Directive (EU) 2024/2831 on platform work contains algorithmic-management rules, including limits on processing certain worker data, DPIA-related requirements, transparency, human oversight and review. At the cut-off date Member State transposition remains material, so local implementation should be checked rather than assuming uniform direct effect.
5.3 ePrivacy and communications confidentiality
The ePrivacy Directive and national implementing laws govern confidentiality of communications, traffic/location data, terminal access and direct marketing. AI assistants that read email, listen to calls, access device identifiers, set SDK storage or analyse messaging content can trigger these rules in addition to the GDPR. Article 50 AI disclosure does not constitute ePrivacy consent.
Voice assistants and meeting transcribers also process data of bystanders and non-account holders. The product design should provide prominent recording/transcription signals, local controls, retention limits and a lawful basis for each participant category. Employment and wiretap/recording laws can be stricter.
5.4 Digital Services Act and consumer law
Online-platform AI may be subject to DSA rules on dark patterns, advertising transparency, recommender systems, minors and systemic-risk mitigation. Profiling-based advertising using GDPR special-category data is prohibited, as is profiling-based advertising to known minors. Very large platforms and search engines have broader duties concerning fundamental-rights and societal risks.
The Unfair Commercial Practices Directive, Consumer Rights Directive and unfair-terms rules address misleading claims, omissions, aggressive persuasion, hidden commercial intent, subscription traps and imbalanced terms. Consumer-law intervention does not require the AI Act’s significant-harm threshold. Marketing a system as objective, unbiased, private or “human-reviewed” without substantiation can itself create liability.
5.5 Copyright, database rights and trade secrets
The DSM Copyright Directive contains text-and-data-mining exceptions for research organisations and cultural-heritage institutions and a broader exception where access is lawful and rights have not been reserved in the required manner. The scope of those exceptions, contract terms, database rights and national implementation must be assessed separately from GDPR. Neither licensing nor a TDM exception supplies a personal-data legal basis.
Trade secrets protect model architecture, weights, feature engineering and business rules, but rights and regulator access must be balanced. Explanation can often be provided through factors, ranges, relative influence, test scenarios and decision logic without disclosing source code. Procurement should require enough technical information to comply with C‑203/22, DPIA and AI Act duties.
5.6 Data Act, data-sharing frameworks and public-sector data
The Data Act can create access and sharing rights for connected-product and related-service data, but it does not override the GDPR or create a basis for disclosing personal data to a third party. A recipient training an AI model still needs a lawful basis, purpose limitation, transparency and rights mechanisms. The Data Governance Act and Open Data framework likewise preserve protections for personal, confidential and intellectual-property-protected data.
5.7 Product safety, cybersecurity and civil liability
AI embedded in medical devices, machinery, vehicles or other regulated products can be subject to both sector product-safety law and the AI Act. NIS2 and DORA can add governance, resilience, incident-reporting and third-party-risk duties for covered entities. A single incident can therefore require parallel assessment as a GDPR personal-data breach, an AI Act serious incident, a sector safety event and a cyber incident.
Civil exposure can arise under GDPR Article 82, national contract and tort law, consumer law, equality law and product liability. Directive (EU) 2024/2853 modernises product liability for software and digital products; Member State transposition and its temporal application must be checked. The Commission formally withdrew its proposal for an AI Liability Directive, with the withdrawal published in the Official Journal on 6 October 2025; fault-based AI liability therefore remains a matter of Member State law.
5.8 Competition and market power
Data access, model tying, exclusivity, interoperability restrictions and cross-service data combination can raise competition and Digital Markets Act issues. The CJEU has confirmed that competition authorities may consider GDPR compliance where relevant to abuse analysis, while not replacing data-protection authorities. A dominant AI ecosystem should not assume that consent obtained through market power is freely given or that data combination is contractually necessary.
5.9 Territorial scope and cross-border mismatches
The GDPR applies through establishment, offering goods or services to persons in the Union, or monitoring their behaviour in the Union. The AI Act has a different market-and-output orientation, including providers placing systems or GPAI models on the Union market, deployers in the Union, and certain non-EU providers or deployers where output is used in the Union. One regime may therefore apply when the other does not.
Non-EU providers may require a GDPR representative and/or an AI Act authorised representative, depending on the facts. These are different roles. For EEA States outside the EU, confirm AI Act incorporation and national transitional arrangements.
6. Enforcement, remedies and allocation of exposure
6.1 Parallel authorities and legal tests
Regime / authority: GDPR — supervisory authority and courts
- Principal questions: Lawful basis, fairness, transparency, rights, Article 22, DPIA, security, transfers, roles.
- Typical measures: Orders to stop/restrict processing, erase or rectify data, provide information, suspend transfers, administrative fines; Article 82 damages.
Regime / authority: AI Act — market-surveillance authority / AI Office
- Principal questions: Prohibited practice, classification, conformity, provider/deployer duties, GPAI/systemic risk, transparency.
- Typical measures: Corrective action, withdrawal/recall or restriction, information demands, administrative fines; GPAI supervision by the AI Office.
Regime / authority: Equality / employment bodies and courts
- Principal questions: Direct/indirect discrimination, accessibility, labour consultation, dismissal or monitoring legality.
- Typical measures: Injunction, compensation, burden-shifting, reinstatement or labour remedies under national law.
Regime / authority: Consumer / DSA authorities and Commission
- Principal questions: Deception, dark patterns, ads, recommender transparency, platform systemic risk.
- Typical measures: Cessation orders, platform remedies, fines and collective redress.
Regime / authority: Sector regulators / notified bodies
- Principal questions: Clinical safety, credit, insurance, communications, financial resilience, product conformity.
- Typical measures: Authorisation conditions, remediation, reporting, product action and sector sanctions.
Regime / authority: Civil courts
- Principal questions: Contract, negligence, product defect, personality/defamation, IP, damages and causation.
- Typical measures: Damages, injunctions, declarations, evidence preservation and contractual relief.
The same facts can support separate proceedings because the protected interests and statutory elements differ. Authorities and courts must still respect proportionality, procedural fairness and EU principles concerning cumulative punitive sanctions. Operationally, an organisation should not assume that clearance by a data-protection authority binds a market-surveillance, equality, consumer or sector regulator.
6.2 Indicative maximum administrative fines
Under the GDPR, the upper tier is the greater of EUR 20 million or 4% of worldwide annual turnover; the lower tier is the greater of EUR 10 million or 2%, subject to the statutory criteria. Under the AI Act, prohibited-practice breaches can reach EUR 35 million or 7%; other operator-obligation breaches can reach EUR 15 million or 3%; supplying incorrect, incomplete or misleading information can reach EUR 7.5 million or 1%; and GPAI provider fines can reach the higher of EUR 15 million or 3% of worldwide annual turnover. SME rules and the precise calculation require Article 99/101 analysis.
Financial penalties are not the principal operational risk. A prohibition, processing ban, model withdrawal, deletion order, transfer suspension, procurement exclusion, adverse equality finding or inability to explain decisions can remove the system from use and force retraining or redesign.
6.3 Evidence and causation
Logs, dataset provenance, model cards, validation results, version histories, human overrides and incident records can be decisive. At the same time, those records may contain personal data and trade secrets. Evidence design should therefore define what is logged, why, who can access it, how long it is retained and how it can be produced to regulators or courts without unnecessary disclosure.
7. Integrated compliance protocol for an AI use case
- Inventory the system and versions. Record the model, application, vendor, deployment environment, intended purpose, affected population, decision consequences and modification history.
- Map the data by lifecycle stage. Identify sources, categories, data subjects, labels, embeddings, prompts, retrieved context, outputs, logs, feedback, checkpoints and deletion paths.
- Allocate legal roles by purpose. Map AI Act provider/deployer/importer/distributor and GDPR controller/joint-controller/processor positions separately.
- Classify the AI Act use. Test prohibited practices first, then high-risk categories/exemptions, GPAI duties, Article 50 transparency, value-chain and substantial-modification rules.
- Establish Article 6 and Article 9/10 routes. Do not combine training, service, security and improvement into one basis. Record necessity, balancing, national-law conditions and objections.
- Test purpose compatibility and source legality. Review original notices, licences, confidentiality, ePrivacy, copyright/TDM, database rights and source restrictions.
- Complete DPIA and, where required, FRIA. Keep privacy, discrimination, due process, accessibility, safety, cyber and misuse risks distinct while using one evidence base.
- Design rights and explanations. Create notice, access, correction, objection, erasure, contest and human-review procedures; contract for vendor support.
- Implement data and model controls. Minimise, segregate, protect and retain data by layer; test extraction, leakage, drift, subgroup performance and foreseeable misuse.
- Engineer meaningful human oversight. Define reviewer competence, authority, time, information, escalation and override documentation; measure automation bias and override rates.
- Contract and transfer correctly. Use Article 28/26 terms, sub-processor controls, locations, no-training boundaries, AI Act documentation, audits, incident cooperation and change control.
- Monitor and re-assess. Trigger review for new data sources, model versions, fine-tuning, changed thresholds, new populations, new purposes, vendor terms, incidents or regulatory changes.
7.1 Go/no-go red flags
Red flag: Untargeted scraping of facial images for a recognition database
- Why it is a stop or escalation condition: Potential AI Act prohibited practice; public availability or consent from a subset does not cure the practice.
Red flag: Workplace or education emotion inference outside medical/safety exception
- Why it is a stop or escalation condition: AI Act prohibition; also serious employment, equality and validity concerns.
Red flag: No articulated Article 6 basis per lifecycle purpose
- Why it is a stop or escalation condition: The AI Act does not provide a general training or deployment basis.
Red flag: Special-category or offence data without a specific gateway
- Why it is a stop or escalation condition: Articles 9/10 prohibit processing absent a valid condition and safeguards.
Red flag: Vendor cannot explain sources, factors, retention, reuse or sub-processors
- Why it is a stop or escalation condition: Controller cannot complete DPIA, transparency, rights or AI Act supply-chain duties.
Red flag: Human reviewer cannot depart from the score
- Why it is a stop or escalation condition: Article 22 risk and ineffective AI Act human oversight.
Red flag: Model anonymity claimed without extraction/linkability testing
- Why it is a stop or escalation condition: Pseudonymisation or deleted source data do not prove anonymity.
Red flag: Production decision system has no appeal or correction route
- Why it is a stop or escalation condition: Conflicts with Article 22 safeguards, accuracy, due process and sector law.
Red flag: Training or improvement hidden inside service terms
- Why it is a stop or escalation condition: Purpose, transparency, consent and fairness defects; processor may become controller.
Red flag: High-impact deployment proceeds before DPIA/FRIA and worker consultation
- Why it is a stop or escalation condition: Assessments must precede processing/use; delayed AI Act high-risk dates do not delay GDPR or labour duties.
7.2 Minimum procurement questions
- Which datasets and source categories were used, and what exclusion, rights-reservation and special-category controls were applied?
- Is the model itself treated as anonymous, pseudonymous or personal data, and what extraction or membership-inference testing supports that position?
- For each customer-data use, is the vendor processor, joint controller or separate controller? Is model improvement disabled by contract and architecture?
- What prompts, outputs, safety samples, telemetry and support records are retained, where, for how long, and by which sub-processors?
- Can the vendor provide data-subject search, export, correction, suppression, erasure and intelligible decision-explanation assistance?
- What performance, subgroup, drift, robustness, cybersecurity and foreseeable-misuse evidence is available for the intended population?
- What model or policy changes can occur without notice, and what changes trigger re-validation, conformity work or a new DPIA?
- How are incidents routed across GDPR, AI Act, NIS2/DORA, product-safety and sectoral reporting?
8. Areas of legal uncertainty and monitoring
Issue: Whether model parameters are personal data
- Current position: Fact-specific. EDPB Opinion 28/2024 requires a robust anonymity analysis focused on identification and extraction risk.
- Monitoring action: Re-test after model release, fine-tuning, new attacks, expanded access or new auxiliary data.
Issue: Legitimate interests for web-scale training
- Current position: Not categorically excluded, but requires a granular three-part test and safeguards; public accessibility is not permission.
- Monitoring action: Monitor EDPB and DPA final guidance, enforcement and CJEU litigation.
Issue: Article 14 at dataset scale
- Current position: Disproportionate-effort exception is possible but narrow and safeguard-dependent.
- Monitoring action: Maintain evidence for numbers, age, risk, notice design and rights channels.
Issue: Rights against trained models
- Current position: No single mandatory technical method; remedy must be effective and role-specific.
- Monitoring action: Track regulator decisions on unlearning, suppression, model access and proof of anonymity.
Issue: Role allocation in multi-party AI stacks
- Current position: Purpose-by-purpose factual analysis; contract labels are not decisive.
- Monitoring action: Review new telemetry, improvement, feedback and co-design features.
Issue: AI Act standards and delayed high-risk implementation
- Current position: Current application dates are reflected here; technical standards, templates and guidance continue to mature.
- Monitoring action: Use consolidated EUR-Lex text and official implementation sources before each release.
Issue: Member State special-category, employment and biometric law
- Current position: Material national variation under GDPR arts 9–10 and labour/constitutional law.
- Monitoring action: Add local-law annex for every deployment country.
Issue: EEA incorporation
- Current position: GDPR is established in the EEA; AI Act applicability outside EU depends on incorporation and national steps.
- Monitoring action: Confirm for Norway, Iceland and Liechtenstein.
Issue: July 2026 EDPB consultation materials
- Current position: Consultation-stage web-scraping/anonymisation materials are not treated as final authority in this memorandum.
- Monitoring action: Replace directional points with final text when adopted.
9. Conclusion
Personal-data processing by AI should be governed as a sequence of legally distinct operations inside a wider system-risk framework. The GDPR decides whether each operation involving a person’s data is lawful, fair, transparent, limited, accurate, secure and contestable. The AI Act decides whether the practice is prohibited, whether the system or model is subject to risk-tier obligations, and what evidence, oversight and disclosures are required. Neither analysis can be collapsed into the other.
The practical rule is to classify first and deploy second. Establish the purpose and data pathway; allocate controller and provider roles; identify Article 6 and 9/10 authority; complete DPIA and FRIA where required; design meaningful human review and explanations; obtain supply-chain evidence; and preserve rights and incident response at every data layer. The highest risk is not “AI” in the abstract, but an unexamined shift of purpose or responsibility between source collection, training, inference and consequential decision-making.
