Pseudonymization vs anonymization under the GDPR: what changes for licensing
Pseudonymised data usually remains personal data under the GDPR because a key or other information can re-link it, while truly anonymous data falls outside the regulation. Recital 26 sets the test: whether anyone could identify the person by means reasonably likely to be used. Counsel decides which side a dataset sits on.
Is pseudonymised data still personal data under the GDPR?
Usually yes, and anonymous data usually is not. Pseudonymisation replaces direct identifiers with codes while a way to re-link them still exists, so the data remains personal data in the hands of whoever can re-link it. Anonymisation means the person is no longer identifiable by any means reasonably likely to be used, and then the GDPR does not apply to that data. The line between the two is a legal and technical judgment, not a label.
The GDPR text on EUR-Lex sets the vocabulary. Article 4 defines personal data and pseudonymisation, and Recital 26 explains that the regulation does not apply to anonymous information, while pseudonymised data that could be attributed to a person with additional information should be treated as identifiable. It also tells controllers to account for all means reasonably likely to be used to identify someone.
This is general information, not legal, tax or financial advice. Confirm with your own counsel before acting.
Side-by-side: pseudonymisation and anonymisation
| Question | Pseudonymisation | Anonymisation |
|---|---|---|
| What changes | Direct identifiers are replaced with tokens or codes | Identifiability is removed so the person cannot be singled out, linked or inferred |
| Can it be reversed? | Yes, if the key or other additional information exists | Not by any means reasonably likely to be used |
| Personal data under the GDPR? | Generally yes for the party that can re-link | Not personal data once truly anonymous |
| Who holds the key? | The company or a trusted separate party | No key should exist |
| Common techniques | Tokenisation, hashing with secret salt, consistent ID replacement | Aggregation, generalisation, removal of rare combinations, noise |
| Typical risk | Treating tokens as anonymous | Overstating anonymity when free text still names people |
| Use case in licensing | Keeps account and ticket relationships intact | Safest for outside release, but loses some structure |
What the EDPB and the courts add
The European Data Protection Board published Guidelines 01/2025 on pseudonymisation, first released for public consultation in 2025. Check the EDPB site for the current status and final text. They explain how pseudonymisation can reduce risk without removing data from the GDPR's scope, and they stress that pseudonymised data stays personal data when the means to attribute it still exist.
The EU's Court of Justice has also considered whether pseudonymised data is personal data from the recipient's point of view, not only the sender's. The reasoning turns on whether the recipient has reasonably likely means to identify the person. Case law is moving, so counsel should check the latest ruling and its limits before anyone relies on a recipient-perspective argument. This page does not rely on any specific case citation.
Workplace-record examples
| Record | Pseudonymised treatment | Anonymised treatment | Residual risk |
|---|---|---|---|
| Helpdesk ticket from a named EU employee | Name replaced by "Requester 118"; mapping kept by the company | Name removed, role and country generalised, free text redacted | Free text can still name colleagues |
| Slack thread with direct messages | Usernames tokenised across the corpus | Thread shortened and speakers merged into roles | Writing style can reveal an author |
| Sales call transcript | Caller and agent tokens; account ID kept | Account ID removed; only outcomes retained | Rare deal details can single out a customer |
| Engineering pull request | Author handle tokenised | Author field dropped | Commit history in public repos can re-identify |
| Approval workflow log | Approver ID tokenised | Aggregated by approval type and time bucket | Small teams make approvers obvious |
Decision rule: the re-linking test
Ask one question of every dataset: who, anywhere in the chain, could re-link a record to a person using means reasonably likely to be used? If the company itself can, the data is pseudonymised at best. If a buyer could combine it with another dataset to single someone out, treat it as personal data too. Only when the honest answer is nobody does the anonymous label start to hold.
Questions to ask before choosing a method
- Where are the people in these records located, and do any sit in the EU?
- Who would hold the re-linking key, and could a buyer ever obtain it?
- Have names been removed from free text and attachments, or only from structured fields?
- Which fields would a determined outsider combine to single someone out?
- Do customer contracts or employee notices speak to reuse?
Mistakes to avoid
| Mistake | Why it hurts | Fix |
|---|---|---|
| Calling tokenised data "anonymous" | It is usually still personal data where a key exists | Use the word "pseudonymised" until counsel concludes otherwise |
| Scrubbing only structured fields | Free text keeps names and details | Scan bodies, notes and attachments |
| Keeping the key next to the data | Re-linking becomes trivial | Separate the key and restrict access |
Why a US company should care
The GDPR can reach non-EU organizations that offer goods or services to people in the EU or monitor their behavior. If a US company's records include EU customers, employees or contacts, its data can carry GDPR questions even though the company sits in the United States. A related question is whether the EU AI Act touches a US licensor, answered in does the EU AI Act apply to US companies.
California has its own vocabulary, with distinct treatment of deidentified information; see whether licensing is a sale under the CCPA and the employee data exemption guide. Workplace messages raise ownership and access questions too: can an employer export Slack and Teams DMs and who owns documents employees create.
When pseudonymisation wins and when anonymisation wins
- Choose pseudonymisation when buyers need to follow an account, ticket or deal through time, the company keeps the key, and counsel is comfortable that the data stays within a controlled legal framework.
- Choose anonymisation when the data leaves the company's control and the structure can survive aggregation or generalisation.
- Choose neither yet if the company cannot say where EU personal data sits, or if free text and attachments have not been scanned.
How SourceX fits
Partners are not asked to decide whether a dataset is pseudonymised or anonymous. That call sits with the company and its counsel, who also say whether EU personal data is present. The chosen method is then written into the de-identification requirements agreed before work begins.
Next step
Use the company fit checker, a preliminary non-binding screen, and read how it works. Then register as a partner to introduce a US company.
Common questions
Is hashed data anonymous?
Not usually. Hashing a name or email produces a consistent identifier that can often be matched against known values, so regulators generally treat hashed identifiers as pseudonymous. A secret salt held separately improves the technique, but the data still counts as personal data for anyone able to re-link it. Counsel should assess the specific method.
Does the GDPR apply to a US company with no EU office?
It can. The regulation applies to organizations outside the EU that offer goods or services to people in the EU or monitor their behavior, so a US company with EU customers, users or employees may be in scope. The company's counsel should check the facts rather than assume either way.
Can pseudonymised data be licensed?
It can be handled lawfully if the licensor has a valid basis and the license terms control onward use, but it remains regulated personal data for anyone who can re-link it. Many companies prefer stronger transformation or exclusion of EU personal data. This is a question for the company's privacy counsel.
What makes anonymisation fail in business records?
Free text and rare combinations. Names inside ticket bodies, unusual job titles in small teams, and distinctive deal details can single out a person even after identifier fields are removed. Testing on a sample, including a search for residual names, is how companies find those gaps before release.
Does a partner need to understand this to refer a company?
No. Partners make introductions and give basic fit information, and never export, upload or describe confidential records. It helps to know the vocabulary so you do not promise that data is anonymous. Rights review and redaction requirements are handled between the company and SourceX.
Related pages
- Does the EU AI Act apply to a US company that licenses data?
- Is licensing company records a 'sale' under the CCPA?
- CCPA employee data exemption expired: what it means for licensing workplace records
- Can an employer export Slack and Teams DMs for data licensing?
- Who owns documents employees create? Work made for hire explained
- Check Company Fit for Data Licensing
Free resources
- Business succession planning assessment — Ten questions on successor, transition and documentation.
- NPV calculator — Net present value with a discounted cash flow table.
- Time value of money calculator — Future and present value with optional regular payments.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment