Pseudonymization vs anonymization under the GDPR: what changes for licensing

Pseudonymised data usually remains personal data under the GDPR because a key or other information can re-link it, while truly anonymous data falls outside the regulation. Recital 26 sets the test: whether anyone could identify the person by means reasonably likely to be used. Counsel decides which side a dataset sits on.

Is pseudonymised data still personal data under the GDPR?

Usually yes, and anonymous data usually is not. Pseudonymisation replaces direct identifiers with codes while a way to re-link them still exists, so the data remains personal data in the hands of whoever can re-link it. Anonymisation means the person is no longer identifiable by any means reasonably likely to be used, and then the GDPR does not apply to that data. The line between the two is a legal and technical judgment, not a label.

The GDPR text on EUR-Lex sets the vocabulary. Article 4 defines personal data and pseudonymisation, and Recital 26 explains that the regulation does not apply to anonymous information, while pseudonymised data that could be attributed to a person with additional information should be treated as identifiable. It also tells controllers to account for all means reasonably likely to be used to identify someone.

This is general information, not legal, tax or financial advice. Confirm with your own counsel before acting.

Side-by-side: pseudonymisation and anonymisation

QuestionPseudonymisationAnonymisation
What changesDirect identifiers are replaced with tokens or codesIdentifiability is removed so the person cannot be singled out, linked or inferred
Can it be reversed?Yes, if the key or other additional information existsNot by any means reasonably likely to be used
Personal data under the GDPR?Generally yes for the party that can re-linkNot personal data once truly anonymous
Who holds the key?The company or a trusted separate partyNo key should exist
Common techniquesTokenisation, hashing with secret salt, consistent ID replacementAggregation, generalisation, removal of rare combinations, noise
Typical riskTreating tokens as anonymousOverstating anonymity when free text still names people
Use case in licensingKeeps account and ticket relationships intactSafest for outside release, but loses some structure

What the EDPB and the courts add

The European Data Protection Board published Guidelines 01/2025 on pseudonymisation, first released for public consultation in 2025. Check the EDPB site for the current status and final text. They explain how pseudonymisation can reduce risk without removing data from the GDPR's scope, and they stress that pseudonymised data stays personal data when the means to attribute it still exist.

The EU's Court of Justice has also considered whether pseudonymised data is personal data from the recipient's point of view, not only the sender's. The reasoning turns on whether the recipient has reasonably likely means to identify the person. Case law is moving, so counsel should check the latest ruling and its limits before anyone relies on a recipient-perspective argument. This page does not rely on any specific case citation.

Workplace-record examples

RecordPseudonymised treatmentAnonymised treatmentResidual risk
Helpdesk ticket from a named EU employeeName replaced by "Requester 118"; mapping kept by the companyName removed, role and country generalised, free text redactedFree text can still name colleagues
Slack thread with direct messagesUsernames tokenised across the corpusThread shortened and speakers merged into rolesWriting style can reveal an author
Sales call transcriptCaller and agent tokens; account ID keptAccount ID removed; only outcomes retainedRare deal details can single out a customer
Engineering pull requestAuthor handle tokenisedAuthor field droppedCommit history in public repos can re-identify
Approval workflow logApprover ID tokenisedAggregated by approval type and time bucketSmall teams make approvers obvious

Decision rule: the re-linking test

Ask one question of every dataset: who, anywhere in the chain, could re-link a record to a person using means reasonably likely to be used? If the company itself can, the data is pseudonymised at best. If a buyer could combine it with another dataset to single someone out, treat it as personal data too. Only when the honest answer is nobody does the anonymous label start to hold.

Questions to ask before choosing a method

  1. Where are the people in these records located, and do any sit in the EU?
  2. Who would hold the re-linking key, and could a buyer ever obtain it?
  3. Have names been removed from free text and attachments, or only from structured fields?
  4. Which fields would a determined outsider combine to single someone out?
  5. Do customer contracts or employee notices speak to reuse?

Mistakes to avoid

MistakeWhy it hurtsFix
Calling tokenised data "anonymous"It is usually still personal data where a key existsUse the word "pseudonymised" until counsel concludes otherwise
Scrubbing only structured fieldsFree text keeps names and detailsScan bodies, notes and attachments
Keeping the key next to the dataRe-linking becomes trivialSeparate the key and restrict access

Why a US company should care

The GDPR can reach non-EU organizations that offer goods or services to people in the EU or monitor their behavior. If a US company's records include EU customers, employees or contacts, its data can carry GDPR questions even though the company sits in the United States. A related question is whether the EU AI Act touches a US licensor, answered in does the EU AI Act apply to US companies.

California has its own vocabulary, with distinct treatment of deidentified information; see whether licensing is a sale under the CCPA and the employee data exemption guide. Workplace messages raise ownership and access questions too: can an employer export Slack and Teams DMs and who owns documents employees create.

When pseudonymisation wins and when anonymisation wins

  • Choose pseudonymisation when buyers need to follow an account, ticket or deal through time, the company keeps the key, and counsel is comfortable that the data stays within a controlled legal framework.
  • Choose anonymisation when the data leaves the company's control and the structure can survive aggregation or generalisation.
  • Choose neither yet if the company cannot say where EU personal data sits, or if free text and attachments have not been scanned.

How SourceX fits

Partners are not asked to decide whether a dataset is pseudonymised or anonymous. That call sits with the company and its counsel, who also say whether EU personal data is present. The chosen method is then written into the de-identification requirements agreed before work begins.

Next step

Use the company fit checker, a preliminary non-binding screen, and read how it works. Then register as a partner to introduce a US company.

Common questions

Is hashed data anonymous?

Not usually. Hashing a name or email produces a consistent identifier that can often be matched against known values, so regulators generally treat hashed identifiers as pseudonymous. A secret salt held separately improves the technique, but the data still counts as personal data for anyone able to re-link it. Counsel should assess the specific method.

Does the GDPR apply to a US company with no EU office?

It can. The regulation applies to organizations outside the EU that offer goods or services to people in the EU or monitor their behavior, so a US company with EU customers, users or employees may be in scope. The company's counsel should check the facts rather than assume either way.

Can pseudonymised data be licensed?

It can be handled lawfully if the licensor has a valid basis and the license terms control onward use, but it remains regulated personal data for anyone who can re-link it. Many companies prefer stronger transformation or exclusion of EU personal data. This is a question for the company's privacy counsel.

What makes anonymisation fail in business records?

Free text and rare combinations. Names inside ticket bodies, unusual job titles in small teams, and distinctive deal details can single out a person even after identifier fields are removed. Testing on a sample, including a search for residual names, is how companies find those gaps before release.

Does a partner need to understand this to refer a company?

No. Partners make introductions and give basic fit information, and never export, upload or describe confidential records. It helps to know the vocabulary so you do not promise that data is anonymous. Rights review and redaction requirements are handled between the company and SourceX.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment