EU data residency and GDPR for AI workloads: a practical guide
GDPR compliance for AI workloads is often treated as a legal question handled separately from infrastructure decisions, but many of the practical requirements are determined by where and how a system is built, not only by a data processing agreement. This post covers the infrastructure-level decisions that affect GDPR compliance for RAG and LLM deployments.
Data residency is more than a region setting
Choosing a data center located within the EU, such as Frankfurt, Amsterdam, Helsinki, or Paris, addresses where data is stored at rest, but residency also covers where data is processed, where backups live, and where logs and telemetry are sent. A system can store primary data in the EU while still sending logs or metrics to a processor outside the EU, which undermines the residency claim.
Review the full data flow for a deployment, including monitoring, error tracking, and any third-party API calls made during inference, such as an external embedding or moderation service. Each of these is a potential path for data to leave the intended region.
Personal data inside training and retrieval content
Documents used for RAG or fine-tuning frequently contain personal data, whether or not that was the original intent of the document. Treat any corpus containing customer names, contact details, or other identifying information as personal data under GDPR, with corresponding obligations around lawful basis, retention, and deletion.
For fine-tuning specifically, consider whether personal data needs to be present in training examples at all, or whether it can be anonymized or replaced with synthetic equivalents. Removing personal data from training data eliminates the harder compliance questions around a data subject's rights with respect to a trained model.
The right to erasure and model weights
A data subject's right to erasure is straightforward to satisfy for data stored in a database, but it is genuinely difficult to satisfy for data embedded into fine-tuned model weights, since there is no reliable way to remove a specific individual's influence from a trained model without retraining.
The practical mitigation is to avoid training on personal data directly where possible, and where it cannot be avoided, to maintain the ability to retrain or discard affected models on a defined timeline when an erasure request is received.
Processor agreements and subprocessors
Any infrastructure or model provider involved in processing personal data is a processor or subprocessor under GDPR and requires an appropriate data processing agreement. This applies to the underlying GPU cloud provider, not only to the application vendor, and it applies to any third-party model API used, even briefly, during a pipeline.
Maintain a current list of all subprocessors touching personal data in an AI pipeline. This list is frequently requested during vendor security reviews and is far easier to produce if maintained continuously rather than reconstructed on demand.
- Map every third-party service touching data in the pipeline, including monitoring and logging tools
- Confirm a data processing agreement exists for each one
- Reassess this list whenever a new integration or model provider is added
Building compliance into the pipeline, not around it
Retrofitting GDPR compliance onto an existing pipeline is significantly more expensive than designing for it from the start, particularly around data lineage tracking needed to respond to access and erasure requests. Instrument the pipeline to record where each piece of data came from and where it has been used from the outset.
Choosing infrastructure with EU data residency by default, rather than as an optional configuration, removes one recurring source of compliance risk and simplifies the data flow review described earlier in this post.
Want this applied to your data?
We scope private AI projects in one call and start with a pilot you can evaluate.