A sheep passes through an AI-labelled shearing machine between 1968 and 2030, illustrating changing privacy risk over time.
About 14 mins
Share:

Key Takeaways

  • Data protection law is built around a single decisive moment — typically when data is collected and consent is given — with the assumption that the risk can be understood and assessed once and for all.
  • AI systems break that assumption in two directions: 1) data can become more revealing after the fact, once a new capability that nobody anticipated at collection and consent emerges; and 2) data keeps becoming more revealing as the AI system continues to observe, infer, and connect information long after the original processing decision was made.
  • The EU AI Act is already moving towards a more dynamic approach to risk assessment, requiring organisations to consider the period and frequency of use for certain high-risk AI systems and to update the assessment if relevant elements change or become outdated.
  • Because AI has exposed that the risk attached to data can change long after the original processing decision was made, we don’t necessarily need another privacy right or another compliance requirement, but a change in how we think about the privacy risk itself.

Somewhere, right now, there’s probably a photo of you that you’ve completely forgotten about. Maybe a conference badge shot from 2016, a tagged photo from a friend’s wedding, or a LinkedIn headshot you replaced years ago but never deleted from wherever it first appeared. When you posted that photo, there was nothing particularly concerning about it from a data protection perspective. It was just a face, sitting on the internet like millions of other faces.

But here’s the important part: although the photo hasn’t changed, everything around it did, including the technology, the ways data can be linked, and what can now be done with something that once seemed almost completely harmless. This raises a question we should spend more time reflecting on when we talk about data protection and AI: what happens when the privacy risk of a piece of data doesn’t really exist when it’s collected, but that changes years later because the technology, and even the entire world around that data, has changed?

I think this question exposes something truly fundamental, namely the fact that the privacy risk of data can change over time, sometimes long after it was collected and processed.

Why the Point of Collection Matters

Data protection law places a lot of emphasis on the moment when processing begins. Before an organisation can process someone’s personal data, it needs to have a lawful basis for doing so. Basically, it needs to know what it’s going to use the data for and whether that use is likely to create risks for the people whose data is being processed. If the processing is likely to result in a high risk, a data protection impact assessment should be carried out before the processing begins. In other words, organisations are expected to think about the processing, its purpose, and the risks it may create before setting anything in motion.

The point of collection is generally treated as the best place to assess risks because that’s when a person has the opportunity to say yes or no. The problem is that this approach assumes that all the risks can be understood when the data is collected, even those that may only emerge in the future, when the data could reveal something different or be used in circumstances that weren’t foreseeable at the time. But this assumption doesn’t hold in AI systems. AI systems are where things start to come apart, and they come apart in two directions that are rarely discussed as part of the same problem.

Two Ways Privacy Risk Moves

The first one is backward-looking and happens when something changes after the data has already been collected. Sometimes, data that seemed perfectly ordinary at the time it was collected can suddenly become much more revealing because of a new capability that didn’t exist then. Nobody necessarily did anything wrong, nor did the organisation holding it suddenly become careless. What changed was everything around the data, whether that was a new technology or a new way of processing it that turned what once seemed harmless into something that could identify, infer information about, or expose a person to risk.

The second one is forward-looking and happens when the system keeps running. This is where AI creates a somewhat different problem because nothing has to change outside the system for the risk to grow. For example, an agentic AI tool can observe a person over repeated interactions, build up a profile of them, connect new information with things it already knows, and keep refining what it thinks it knows about them. The longer this goes on, the more information the system can accumulate and the more it can potentially infer. The UK’s Information Commissioner’s Office (ICO) made a similar point in its 2026 Tech Futures report on agentic AI, noting that these systems can process significant and growing amounts of personal information simply because of how they operate.

Taken together, these two points show the privacy risk in AI isn’t something you can assess once and then leave behind at the point of collection. That’s because the level of risk can change later, sometimes gradually as a system keeps running and building on what it already knows, and other times quite suddenly when a new technology or capability appears. Either way, the risk may no longer look anything like it did when the original assessment was made.

An Example That Proves Both Points

What I’ve described above may sound like a theoretical problem, but it isn’t. The story of Clearview AI shows what this looks like in practice, and it brings both sides of the problem together.

Clearview AI is a company that built a facial recognition database by scraping photographs from across the open web, including social media profiles, tagged photos, and professional pages, basically anything with a face that its crawlers could find. None of those photos would normally raise any particular data protection concerns. Simply put, a profile photo that you put online for friends, colleagues, or potential employers to see isn’t particularly sensitive or unusual personal data.

But when a New York Times investigation brought Clearview’s practices to public attention in January 2020, it became clear that the company had found a way to make those ordinary, publicly available photos do something they were never meant to do. The company had developed a facial recognition system through which authorised users could upload a “probe” image — for example, a photograph of an unidentified person — and search it against Clearview’s database of publicly available images of people who had no idea were even in it, something that raises obvious questions under the GDPR. The system returned potential matches and links to the webpages where those images appeared.

European regulators then started looking into what Clearview was doing, and what happened next is important to this argument because the story didn’t end with the original collection of the photos.

Hamburg’s data protection authority and Italy’s Garante found Clearview’s processing unlawful in 2021 and 2022. Italy and Greece each imposed the maximum GDPR fine of €20 million in 2022. France’s CNIL also fined the company €20 million that year, followed by another €5.2 million in 2023 after Clearview failed to comply with an order to delete the data of European citizens and stop collecting more of it. The UK’s ICO originally fined Clearview £7.5 million in 2022, although that decision was overturned on appeal in 2023 on a jurisdictional issue relating specifically to how foreign law enforcement agencies use the data of UK residents, rather than because the underlying scraping was found to be acceptable. In September 2024, the Dutch data protection authority imposed a €30.5 million fine, with another €5.1 million added for continued non-compliance.

Currently, Clearview says its service is available only to vetted government agencies and contractors working for those agencies. It also claims the results are only investigative leads, not definitive identifications, and that users are required to independently verify them.

The legal issues here are important, of course, but putting those aside for a moment, this example shows something else. Photos posted in the early 2010s, when the idea of searching for someone by their face across a huge database was far less developed, later became part of exactly that kind of system as the technology caught up. The people in those photos hadn’t done anything new, they hadn’t suddenly disclosed more information about themselves, and the photos hadn’t changed either. What changed was what technology could do with them, and that was enough to turn what had once been fairly ordinary personal data into the subject of regulatory action across several countries.

But there’s another part to the Clearview story, and in some ways I think it’s even more important for what we’re dealing with now: by 2025, the company still hadn’t paid the European fines imposed on it and, according to European regulators, hadn’t stopped collecting photographs either, meaning this wasn’t one event that regulators could investigate, fine the company for, and then close the file on. Since the collection itself keeps going and every new photo scraped creates the same problem again, the risk doesn’t just continue; it keeps getting bigger.

The Law Is Starting to Shift

The Clearview example also raises a more practical question: If privacy risk can change after data has been collected and can continue to grow as an AI system keeps operating, how can a person meaningfully give consent to risks that may not even be knowable at the time they’re asked to agree?

Regulators have noticed this, along with many of the other issues mentioned above, and the EU AI Act already contains some recognition that the risks associated with AI systems may need to be considered beyond the point at which it’s first put into use. For example, Article 27 requires deployers of certain high-risk AI systems to assess their potential impact on fundamental rights, including privacy and data protection, before the systems are put into use.

One of the things that assessment has to cover is the period and frequency of use, which suggests that how long and how often a system is used can matter when it comes to risk. Additionally, if any element of the assessment changes or is no longer up to date while the system is being used, the deployer has to update the information. While that doesn’t amount to a general requirement to reassess AI systems continuously, it does show that the law recognises that an assessment made at the beginning of a system’s use may not remain accurate indefinitely.

The UK’s ICO has been thinking along similar lines. In its work on agentic AI, the ICO has pointed out that these systems can process significant and growing amounts of personal information simply through the way they operate. The more information a system accumulates, the more connections it may be able to make between different pieces of information, and the more it may be able to infer about a person or use that information in ways that weren’t possible from any single piece of data on its own. That means organisations need to think about necessity and data minimisation not only when the system is first introduced, but also as it continues to run. That, I think, should be the starting point, even if the surrounding compliance culture — DPIAs, consent flows, purpose limitation, and so on — still treats the original processing decision as the moment when the question is settled.

What’s still missing, I think, is putting these things together. Most of the guidance and commentary I’ve come across treats retroactive re-identification, the accumulation of information by agentic systems, and privacy risks created by AI-generated inferences as separate problems, each with its own set of controls. But what if they’re actually different versions of the same problem? If the risk attached to data can change over time, sometimes because a new capability appears and other times because a system keeps learning and building on what it already knows, then perhaps the problem isn’t that we need another right or another clause in an already crowded framework, but that we need to rethink the idea that privacy risk can be assessed once and then filed away.

What This Means for How We Assess Risk

As we’ve been able to see, privacy risk can dramatically change over time. A photo or a piece of data posted online or put at the disposal of a company years ago can easily end up in a pile of data that feeds an AI system. From there, it can escalate as different pieces of information are connected, revealing things that weren’t apparent from the original data, things that were never supposed to become public, and eventually even putting that person at risk of harm.

Furthermore, a data protection impact assessment completed when an AI system is launched and never looked at again can end up answering a question that has changed considerably since the assessment was made.

The practical implication may not be significant, and may not mean organisations need to throw out their existing compliance processes. It does mean, however, that certain changes should be treated as a reason to go back and look at the assessment again rather than assuming that what was true when the system was launched is still true today.

A significant improvement in the general capabilities of AI systems that could potentially be used with an organisation’s older datasets is one example. A system moving from a narrow, well-defined use case into something broader and more open-ended, like the kind of drift the ICO has flagged in relation to agentic systems, is another. And sometimes, it may simply be the passage of enough time for the assumptions you made when you first assessed the risk to no longer hold. That alone may be a good reason to take another look.

None of this fits neatly into a checklist, but that’s rather the point. A one-time approach to data protection is comforting because it has a clear end point. You assess the risk, document it, and move on. But if AI systems can change the level of risk over time, then we have to accept that some risk assessments may not have an end point. Treating them as if they do could itself create a risk that wasn’t there when the original assessment was made.

That’s probably not the easiest thing to build into an organisation’s processes, and it’s even harder to explain to a board that would much rather have a box that can be ticked once. But data that was safe yesterday isn’t a guarantee that it will still be safe tomorrow. And the sooner that becomes an ordinary assumption when we think about AI systems — rather than something we only think about after something goes wrong — the fewer Clearview cases there will be to write about.

Extra Sources and Further Reading

 

admin

[atlasvoice]
Transform Your Business with NexusJump Data & AI Tips
To get you started, over the next few days we will send you a series of seven data and AI tips.

Great! We’ve received your information.