A sheep named Maria and a collie labelled Governance illustrate GDPR identity records beside AI similarity patterns.
About 13 mins
Share:

Key Takeaways

  • GDPR was built around identifiable records. AI was built around statistical patterns. Most of the tension between GDPR and AI stems from this fundamental difference, not from the law simply lagging behind technology.
  • Once personal data has been absorbed into a trained model, it no longer exists as a separate record that can simply be accessed, modified, or erased.
  • Many AI harms arise from learned statistical patterns rather than the processing of an identifiable person’s record, making traditional GDPR rights difficult to apply.
  • Questions about erasure, traceability, explainability, accountability, and legal standing often reflect the same underlying mismatch between organising information around identity and organising it around similarity.
  • The real question isn’t how to adapt the GDPR to AI, but whether a framework built around identifiable individuals can fully address systems organised around similarity.

There’s a familiar story being told about GDPR and AI, and chances are you’ve come across it more than once. It goes something like this: GDPR is an established legal framework, AI is a new and complicated technology, and the two are simply going through an adjustment period. Give regulators a little more time, let the courts clarify a few grey areas, publish some additional guidance, and everything will eventually fall into place. Put simply, it’s the regulatory equivalent of two people from different cultures learning to live together, with some initial friction, but nothing that goodwill, enough patience, and open communication can’t fix.

Unfortunately, I don’t think that story is true. The tension between GDPR and AI isn’t the result of a legal framework struggling to keep pace with a new technology but the result of a clash between two fundamentally different ways of organising information: GDPR organises information around identity, AI organises it around similarity.

That’s not a technical mismatch that better engineering, another court decision, or more regulatory guidance will smooth over. It’s a conceptual divide that runs through almost every unresolved question in AI regulation.

GDPR’s right of access, right to erasure, right to rectification, and even the very concept of a data subject make sense only in a world where information is organised around named, identifiable people. AI, by contrast, works by organising information around statistical patterns derived from things that resemble one another. These aren’t two dialects of the same language but two entirely different languages, and pretending otherwise is why so much AI compliance work feels like forcing a translation between concepts that simply don’t have an equivalent.

In many professional circles, discussions about GDPR and AI often begin with questions about consent, transparency, or erasure. But this blog post takes a different approach. Because before we can understand why those rights seem to break down when applied to AI, we first need to examine the assumptions behind GDPR and the AI Act. Once those assumptions are clear, the rest of the debate looks very different.

What Organising by Identity Actually Means

Let’s start with the world GDPR was written for, because it’s worth being precise about what that world actually looked like, rather than what we assume it looked like in hindsight.

Picture a typical CRM system, the sort of thing thousands of organisations across Europe have been using for decades. Someone called Maria signs up for a newsletter. Her name goes into a field, and her email address goes into another. Somewhere there is a timestamp recording when she gave consent, and perhaps another field noting which marketing campaign brought her in. Together, those fields form a single record that represents Maria inside the system.

Now, here’s the important part. That record never becomes anything other than itself. Throughout its entire lifecycle, it remains attached to the person it refers to, in this case, Maria. If the company wants to know what data it holds about Maria, someone can easily retrieve her record. If Maria asks for it to be modified or deleted, the record can be updated or removed accordingly.

Therefore, organising information around identity means that every piece of information has an identifiable owner. GDPR’s entire architecture is built around this assumption. GDPR’s individual rights depend on the idea that personal data exists as discrete records linked to identifiable individuals. If you know who the person is, you know where his or her data is, which means you can inspect, modify, or delete it. It’s a remarkably coherent model for organising information, and for decades of enterprise software, it was also an accurate one.

What Organising by Similarity Actually Produces

Now forget the CRM entirely, and let’s see what actually happens when an AI model is trained on personal data.

Maria’s data, along with other people’s data, enters a training process. Unlike in the CRM, however, nothing is filed under anyone’s name, which means the model doesn’t learn, “This record belongs to Maria”. Instead, it makes tiny adjustments to millions of internal parameters, each nudged very slightly by the data it has been trained on, gradually learning the statistical relationships within that data. Therefore, Maria’s contribution to the model isn’t a record sitting somewhere waiting to be accessed, modified, or erased, but a collection of subtle influences spread across a vast mathematical structure and indistinguishable from the influences of everyone else who was statistically similar to her.

This is, of course, a simplification. In real life, AI models can sometimes memorise or reveal parts of their training data, which is precisely why techniques like machine unlearning, privacy-preserving training, and defences against membership inference attacks have all become active areas of research. These exceptions, however, don’t alter the broader point developed here: AI organises information around learned statistical relationships rather than identifiable records. Understanding the distinction between preserving information about Maria as an identifiable individual and absorbing patterns she contributed to alongside countless others matters a lot.

That’s because, when an AI model generates an output, it isn’t retrieving information about Maria but generating a response based on the patterns distributed across its own internal representations. Not even the engineers who built the model could point to a specific folder and say “That’s Maria’s data, separate from everyone else’s.” Not because that data is hidden, but because it no longer exists as a separate thing.

Two Consequences of Organising Information Around Similarity

The distinction between organising information around identity and organising it around similarity has consequences that extend well beyond the training process. Two of them help explain why familiar GDPR assumptions become difficult to apply to AI.

The first consequence is irreversibility. GDPR’s rights largely assume that processing can be undone. Modify the record, withdraw consent, delete the data, and the effects of that processing disappear. Machine learning doesn’t work like that. Once Maria’s data has nudged a model’s internal parameters during training, removing her original data afterwards doesn’t remove the nudge itself. The model has already learned from it. Fully reversing that may require retraining the model from scratch. Although machine unlearning techniques can sometimes be used to remove the influence of specific training data without full retraining, they remain technically challenging and are not yet a general solution in practice.

The second consequence is the loss of record-level traceability. Traditional software, the kind GDPR was written to sit alongside, behaves according to rules that programmers explicitly wrote. Since the same input produces the same output every time, if something goes wrong, you can usually trace it back to the line of code that caused it.

But AI systems work differently, with their behaviour emerging from statistical relationships learned during training rather than rules written by a programmer. It’s worth being precise here, though, because the issue isn’t that AI systems are unpredictable in some mystical way. If you run an AI model at a fixed configuration on fixed hardware, the outputs can be almost identical across runs. What’s actually different is traceability, and this is where the distinction between organising information around identity and organising it around similarity becomes important again.

In the CRM, if Maria is unfairly denied something because of how her data was processed, a regulator can trace that outcome back to her specific record and inspect exactly what happened to it. That’s because the outcome can be easily linked to the data.

But because AI organises information differently, if Maria is unfairly denied something by an AI model, there is no relevant record to inspect because the output wasn’t produced from her data as a separate, identifiable record.

The reason why the traditional accountability model breaks down is that it assumes an outcome can always be traced back to the processing of an identifiable person’s data. But when information is organised around statistical patterns learned from the data of many individuals, that assumption no longer holds.

One More Consequence: No Identity, No Plaintiff

The previous sections explained why familiar GDPR assumptions become difficult to apply to AI. But the distinction between organising information around identity and organising it around similarity has another consequence: it changes the way we think about who those rights are designed to protect.

GDPR’s rights are addressed to an identified or identifiable natural person. That wording is fundamental to the way the regulation works. Before someone can exercise a right, whether that’s access, correction, objection, or erasure, there has to be a recognisable relationship between that person and a particular piece of data. In other words, there has to be an identifiable person and a corresponding record to which the person’s legal rights attach.

Many AI harms don’t fit that structure. Consider a model that consistently produces less favourable outcomes for people whose names, backgrounds, or other characteristics resemble a particular group. Nobody deliberately decided to disadvantage any one individual. The model simply learned that people who resemble a particular statistical pattern should be treated in a certain way. The harm is real and can appear in hiring decisions, credit scoring, content moderation, recommendation systems, or even in the way a chatbot responds.

A well-known example is Amazon’s experimental recruitment system, which the company eventually abandoned after discovering that it consistently downgraded applications from women. The model wasn’t explicitly programmed to favour men. It simply learned statistical patterns from a decade of historical recruitment data, much of which reflected a male-dominated workforce. The disadvantage emerged from those learned patterns rather than from an explicit rule targeting a particular type of applicant.

This is the essence of “no identity, no plaintiff”. Because AI increasingly produces harms that arise from statistical patterns rather than identifiable records, the difficulty isn’t only identifying the person affected; it’s also that the harm itself didn’t arise from the processing of that person’s identifiable record but from the model’s application of a statistical pattern.

To me, this is the most fundamental and least discussed consequence of the distinction between organising information around identity and organising it around similarity.

Where This Actually Leaves Us

By this point, it should be clear why the tension between GDPR and AI isn’t simply a temporary compliance problem or the result of regulation struggling to keep pace with technology. Neither is it a case of GDPR needing revision in the same way a piece of software needs a patch.

At the same time, none of this means GDPR is irrelevant to AI. GDPR remains an essential framework for governing the processing of personal data, but one that was built for a world where information belongs to identifiable people and can therefore be found, inspected, corrected, and erased on that basis. By contrast, AI was built to learn from patterns that emerge across many people rather than from the record of any one individual.

This distinction isn’t merely theoretical, as it changes the way we should think about AI regulation. Much of the current discussion assumes that the challenge is to extend familiar GDPR concepts a little further so they can accommodate AI systems. Since the tension is structural rather than a temporary mismatch, that approach isn’t just incomplete; it’s misdirected. The recurring debates about erasure, accountability, explainability, group harms, and even legal standing are often treated as separate compliance challenges. However, viewed through the distinction between organising information around identity and organising it around similarity, they look like different expressions of the same underlying mismatch.

The “no identity, no plaintiff” problem illustrates why this matters. A regulatory framework built around individual rights assumes that harm will eventually produce an identifiable person who can exercise those rights. But harms that emerge from statistical patterns don’t always produce such a person. An AI system may consistently disadvantage an entire category of people without ever creating the kind of identifiable, record-based harm on which the GDPR’s individual rights are built. That isn’t a hypothetical edge case but a natural consequence of building accountability on information organised around identity while AI organises information around similarity.

The question, then, isn’t how the GDPR can be adapted to fit AI more neatly but whether a regulatory framework built around identifiable individuals can ever fully address systems that organise information around similarity. Since that may never be entirely possible, the more important question is what a regulatory framework designed around similarity would actually need to look like.

That question matters because regulation shapes how organisations design, deploy, and audit AI systems. If the underlying assumptions are wrong, compliance alone won’t necessarily address the harms AI actually creates.

Extra Sources and Further Reading

  • Legal and Human Rights Issues of AI: Gaps, Challenges and Vulnerabilities – ScienceDirect
    https://www.sciencedirect.com/science/article/pii/S2666659620300056
    This article provides a broad overview of the main legal and human rights challenges raised by AI, including transparency, bias, discrimination, privacy, accountability, liability, and cybersecurity. arguing that existing legal frameworks need continual adaptation.
  • Processing of Synthetic Data in AI Development for Healthcare and the Definition of Personal Data in EU Law – Oxford Academic
    https://academic.oup.com/ijlit/article/doi/10.1093/ijlit/eaag002/8504180?login=false
    This paper examines whether synthetic data generated from real health data should be treated as personal data under the GDPR. It combines legal analysis with an empirical study of membership inference attacks to assess re-identification risks and concludes that, although synthetic data is generally likely to be anonymous, legal uncertainty remains over when identification is considered “reasonably likely”.
  • Lessons from GDPR for AI Policymaking – University of Pennsylvania
    https://laweconcenter.org/wp-content/uploads/2023/09/SSRN-id4528698.pdf
    This paper argues that, while the GDPR provides useful regulatory experience, AI presents greater challenges due to its technical complexity, reliance on self-regulation, and the difficulty of effective enforcement, meaning the GDPR cannot simply be used as a blueprint for AI regulation.
  • The Challenges of the GDPR in the Era of Artificial Intelligence – e-Publica https://revistas.rcaap.pt/epublica/article/download/46039/31014/219479
    This article examines the continuing role of the GDPR in the age of AI, focusing on areas where the Regulation and the AI Act complement each other while also creating legal tensions. It argues that the GDPR remains essential for protecting individuals’ digital rights, but that its interpretation—and in some respects the broader regulatory framework—must evolve to address the challenges posed by AI.

admin

[atlasvoice]
Transform Your Business with NexusJump Data & AI Tips
To get you started, over the next few days we will send you a series of seven data and AI tips.

Great! We’ve received your information.