Security
Balancing innovation and privacy: LLMs under GDPR
How to use LLMs under GDPR: lawful basis, data minimization, DPIAs, erasure and data residency, and where self-hosting and PII detection help.

In short
You can use LLMs under GDPR if you have a lawful basis for the personal data involved, minimize what reaches the model, run a DPIA where the risk is high, and can honor access and erasure requests. Most personal data enters through prompts, retrieved documents, memory and logs, not only training data. Processing in approved regions or on your own infrastructure makes the rest easier to prove.
Yes, you can use large language models under the GDPR, as long as you can show a lawful basis for the personal data they process, keep that data to a minimum and honor people's rights over it. In practice, that means mapping every place personal data enters an LLM system, controlling where it is processed, and keeping records that prove what happened. None of this blocks innovation, but it has to be designed in from the first pilot.
Does GDPR apply to LLMs?
The GDPR applies whenever you process personal data of people in the EU, whether your company is established in the EU or offers services to people there. An LLM system processes personal data as soon as a prompt, a retrieved document or an output contains information about an identifiable person. It makes no difference that the model "only predicts text".
The stakes are real. Under Article 83 of the GDPR, the most serious infringements can be fined up to €20 million or 4% of worldwide annual turnover, whichever is higher.
Where does personal data enter an LLM system?
Most teams think about training data first. In an enterprise deployment, far more personal data flows through the system at run time.
| Where it enters | Example | The GDPR questions it raises |
|---|---|---|
| Training and fine-tuning data | Support transcripts used to fine-tune a model | Lawful basis, compatible purpose, what the model retains |
| Prompts | A customer types account details into a chat | Minimization, retention, transfer to a model provider |
| Retrieved documents (RAG) | An HR assistant retrieves an employee's file | Access control, purpose limitation |
| Uploaded files | A loan applicant uploads bank statements | Special categories of data, retention |
| Agent memory | An assistant remembers a user's preferences | Transparency, erasure |
| Logs and traces | Full prompts and outputs kept for debugging | Retention, access, security |
| Outputs | A generated summary about a named person | Accuracy, automated decision-making |
Map each of these for your use case before you pick a model. The map becomes the backbone of your records of processing and your impact assessment.
Which GDPR principles matter most for LLMs?
Six parts of the regulation come up in almost every LLM project:
- Lawful basis (Article 6). Consent is only one of six bases. Many enterprise uses rely on contract or legitimate interests, which require a documented balancing test. Health, biometric and other special categories of data (Article 9) need an additional condition.
- Purpose limitation and data minimization (Article 5). Use personal data only for the purpose you stated, and send the model only what it needs for the task.
- Storage limitation (Article 5). Set retention periods for prompts, agent memory and traces, not only for source systems.
- Security (Article 32). Encrypt data, restrict access by role and pseudonymize where you can.
- Automated decisions (Article 22). Decisions with legal or similarly significant effects, such as a credit refusal, need meaningful human involvement and safeguards.
- Data protection impact assessments (Article 35). A DPIA is required when processing is likely to result in a high risk to people. Large-scale processing of customer or employee data with a new technology often qualifies.
Transparency runs through all of them. People should know when an AI system processes their data and why, which usually means updating privacy notices before launch.
What have EU regulators said about AI models?
In December 2024, the European Data Protection Board published Opinion 28/2024 on personal data in AI models. Three points from it matter for anyone building on LLMs:
- Anonymity is assessed case by case. For a model to count as anonymous, it should be very unlikely both to identify the people whose data trained it and to let someone extract that data through queries.
- Legitimate interest can work, with a three-step test. You identify the interest, show the processing is necessary for it, and balance it against people's rights, including what they would reasonably expect.
- Unlawful training can follow the model. A model developed with unlawfully processed personal data can affect the lawfulness of its deployment, unless the model has been duly anonymized.
For deployers, the practical takeaway is due diligence: ask your model provider how it trained the model and what it does with your data.
How does the EU AI Act fit in?
The EU AI Act entered into force on August 1, 2024 and applies alongside the GDPR, not instead of it. The AI Act sets obligations by an AI system's risk level, while the GDPR keeps governing the personal data those systems process. Our guide to the EU AI Act covers its timeline and requirements.
Should you use private LLMs or hosted APIs?
Where the model runs decides who else processes your data and where. None of the options makes you compliant by itself; you still need a lawful basis, minimization and a way to handle rights requests.
| Option | Where data is processed | What to put in place |
|---|---|---|
| Hosted model API | The provider's infrastructure and regions | A data processing agreement, transfer safeguards, retention terms |
| Hosted API with regional data residency | The provider's infrastructure in your chosen region | The same agreement, plus a check of which services stay in region |
| Open-weight model on your infrastructure | Your cloud account or data center | Your own security, patching and capacity planning |
Many enterprises combine them: an open-weight model in their own environment for the most sensitive data, and an approved hosted model for the rest.
A GDPR checklist for LLM projects
- Map the data flows for training, prompts, retrieval, files, memory, logs and outputs.
- Choose and document a lawful basis for each purpose, with a balancing test where you rely on legitimate interests.
- Run a DPIA before launch when the processing is likely to be high risk, and revisit it when the use case changes.
- Minimize what reaches the model. Remove or mask personal data the task does not need.
- Set retention periods for prompts, agent memory and traces, and enforce them.
- Control where processing happens, with processor agreements and transfer safeguards for any provider outside your region.
- Build rights handling that can find and erase a person's data everywhere it was copied, including memory and logs.
- Keep a person in the loop for decisions with legal or similarly significant effects.
- Train the people who build and operate the system on these rules.
- Keep records of processing activities and of what each AI system did.
How does Dynamiq support GDPR programs?
Dynamiq is built for teams that must answer these questions for an auditor, and the platform covers SOC 2, HIPAA and GDPR, with a DPA available.
- Processing where you choose. Run Dynamiq in our cloud or self-hosted in your own AWS, Azure, GCP, IBM Cloud or OpenShift environment, on any Kubernetes cluster or on-premises; see where Dynamiq runs.
- Your choice of model. Use the model providers your data protection officer approved from the SDK, call the most-used models through one AI Gateway endpoint, or run open models on managed inference.
- PII detection before the model. Guardrail nodes detect PII and prompt injection at the start of a workflow, so you decide what happens to flagged input.
- Memory you can erase. Agent memory is scoped by user and session and can be deleted per user when a request comes in.
- A record of every run. Each run is traced with its inputs, outputs and steps, which gives your DPIA and your auditors real evidence; see observability.
The security and trust center covers our compliance, data protection and sub-processors in more detail.
FAQ
Can you put personal data in LLM prompts under GDPR?
Yes, if you have a lawful basis for that processing, the data is necessary for the task, and the model provider is bound by a processing agreement with suitable safeguards. Minimize first: mask or remove personal data the model does not need, and set a retention period for stored prompts.
Do you need a DPIA for an LLM project?
You need one when the processing is likely to result in a high risk to people, which is common for large-scale processing of customer, patient or employee data with a new technology. Even when it is not strictly required, a DPIA is the clearest way to document your decisions.
Does the right to erasure apply to LLMs?
Yes. People can ask you to erase their personal data, and that includes copies in prompts, retrieved documents, agent memory and logs. Removing data from a trained model's weights is much harder, which is one reason to keep personal data out of training sets and rely on retrieval instead.
Is a self-hosted LLM automatically GDPR compliant?
No. Self-hosting controls where data is processed and removes a third party from the chain, but you still need a lawful basis, minimization, retention rules, security and a way to handle rights requests.
Does the EU AI Act replace GDPR for AI systems?
No. The AI Act regulates AI systems by risk level, and the GDPR continues to govern any personal data those systems process. Most enterprise AI projects in the EU have to satisfy both.


