GDPR further processing: can operational data be reused for AI training?

GDPR purpose limitation means personal data collected for one purpose can be reused for AI training only if the new purpose is compatible under Article 6(4) or has its own lawful basis. Because that is hard to show for mailboxes and tickets, US companies often exclude or anonymize EU records first.

Can operational data be reused for AI training under GDPR?

It depends on why the data was first collected and on what the person was told. The short rule: a new purpose must either be compatible with the original one under the test in Article 6(4) of the GDPR, or rest on its own legal basis, such as consent. Licensing employee or customer records to an AI developer is rarely an obvious fit for the first path.

That is why many US companies exclude EU records, or anonymize them, before they license. This is general information, not legal, tax or financial advice. Confirm with your own counsel or data protection adviser before acting.

What does purpose limitation require?

GDPR principles say personal data must be collected for specified, explicit and legitimate purposes and not further processed in a way that is incompatible with them. A company that gathered mailbox data to run its business has a stated purpose. Handing a copy to a third party to train models is a different activity.

Three points follow from the text.

  • The purpose is judged at collection, based on what was disclosed and what the person could reasonably expect.
  • "Further processing" by someone else, such as a licensee, still has to be justified.
  • Compatibility is assessed case by case. There is no list of automatically permitted reuses that covers AI training.

What are the Article 6(4) compatibility factors?

When the new purpose is not based on consent, Article 6(4) asks the controller to take into account, among other things, the factors in the table. The wording below paraphrases the regulation; read the official text for the exact language.

FactorThe question to askHow AI-training licensing usually scores
Link between purposesIs the new use close to the original one?Often weak: running a business and training a third party's model are far apart
Context of collectionWhat would the person have expected, given their relationship with the company?Employees and customers rarely expect their messages to leave the company
Nature of the dataDoes it include special categories or criminal-offense data?Any such data weighs strongly against compatibility
Consequences for peopleWhat could happen to them because of the new use?Depends on whether the output could identify or affect them
SafeguardsIs the data encrypted, pseudonymized or otherwise protected?Strong safeguards help but do not cure a weak link

If a company cannot argue that the factors point to compatibility, the fallback is a fresh lawful basis, which for large mailbox or ticket archives is hard to obtain person by person.

How this plays out in common situations

SituationWhat to checkTypical outcome to confirm with counsel
EU-based employee mailboxesEmployee notice, works-council or national employment rules, purposes disclosedExclude from the license, or anonymize rigorously
Customer support tickets from EU contactsPrivacy notice at sign-up, ticket content, free-text identifiersExclude by customer location, or de-identify and document the method
Shared Slack or Teams channels with EU membersWorkspace membership, guests, connected channelsCarve out channels and messages by membership
Sales call recordings with EU prospectsRecording notice, consent, voice as personal dataUsually exclude; see the page on employee voices in licensed recordings
Truly anonymous aggregatesWhether anyone can be identified by any reasonably likely meansMay fall outside the regime if the standard is met, which is a high bar

The wider framing of how buyers read this issue is in the guide to EU AI Act Article 10 data governance questions.

Why anonymization is not a shortcut

The GDPR applies to personal data, and pseudonymized data remains personal data. Replacing names with tokens while keeping the mapping or leaving rare facts in free text does not take records outside the regulation. Genuine anonymization means no one can be identified by means reasonably likely to be used, and that bar is demanding for unstructured text such as email and chat.

For that reason, many practitioners treat a location-based carve-out as the simpler path. It is not perfect either, because EU residents can appear in US systems, but it narrows the problem.

Disclosure and documentation good practice

Whatever route the company takes, a written record helps. Keep a short file that states the original purposes disclosed to employees and customers, the location filters used to exclude EU and UK records, the date the filters ran, and who approved them. If a buyer or regulator later asks how an EU record ended up in a dataset, that file is the first thing counsel will want.

Update employee and customer notices going forward so that future disclosures to licensees are described plainly. A notice written today does not repair past collection, but it improves the position for records created from now on.

How buyers look at this

AI developers and data buyers ask suppliers how rights were established, which records are in scope and which jurisdictions are excluded. A clean, documented exclusion is easier to review than an argument that a compatibility test was met. Related contract points, such as what a buyer may retain after a term ends, are covered in the page on the residuals clause.

If you want the wider background on why data provenance matters, see what AI training data is.

What a referral partner should and should not do

Partners introduce a company and give basic fit information. They do not review, export or describe confidential records, and they do not give legal advice. If a sponsor asks whether EU data can be included, the honest answer is that SourceX and the company's counsel decide that during rights review, and that exclusion is the usual starting point.

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. The reward is paid only after the buyer pays and SourceX receives its fee; no reward is guaranteed.

Questions to put to your counsel

  • Which of our notices described the purposes for our email, chat and ticket data?
  • Which records relate to people in the EU or UK, and can we separate them reliably?
  • Would an anonymization method satisfy the standard, and who documents it?
  • Do any special categories of data appear in free text?
  • What do we promise a buyer about exclusions, and how do we back it up?

Next step

Use the company fit checker for a preliminary, non-binding screen and read how the process works. If you know a US company with 50+ full-time employees at peak (contractors excluded) and years of records, register as a partner to introduce it.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Is Article 6(4) the only route to reuse personal data for AI training?

No. A company can also rely on consent or on another lawful basis that fits the new purpose. In practice, consent from every employee or customer in a large archive is hard to obtain, which is why exclusion or genuine anonymization is more common for EU records.

Does pseudonymizing names make the data compatible for reuse?

Pseudonymization is a safeguard that counts in the compatibility analysis, but pseudonymized data is still personal data. It helps, though it does not by itself make a different purpose compatible. Counsel should assess the whole set of factors rather than treat tokenization as a fix.

Do US companies need to worry about this if all their staff are in the US?

Often much less. If a company has no EU establishment and no EU-based people in its records, there may be little EU personal data to analyze. The check is whether any mailbox, channel, ticket or customer record relates to people in the EU or UK.

Can an EU data subject object after a license is signed?

Individuals can have rights over their personal data, and a license does not remove them. That is a reason buyers and suppliers prefer to exclude EU records from the dataset up front rather than manage requests against delivered data later.

Who decides whether EU records can be included in a SourceX license?

The company decides with its own counsel, working with SourceX during rights review. Partners do not make or influence that call. Nothing is binding until the company agrees price and terms and signs, and redaction and exclusion rules are agreed before any work begins.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment