GDPR further processing: can operational data be reused for AI training?
GDPR purpose limitation means personal data collected for one purpose can be reused for AI training only if the new purpose is compatible under Article 6(4) or has its own lawful basis. Because that is hard to show for mailboxes and tickets, US companies often exclude or anonymize EU records first.
Can operational data be reused for AI training under GDPR?
It depends on why the data was first collected and on what the person was told. The short rule: a new purpose must either be compatible with the original one under the test in Article 6(4) of the GDPR, or rest on its own legal basis, such as consent. Licensing employee or customer records to an AI developer is rarely an obvious fit for the first path.
That is why many US companies exclude EU records, or anonymize them, before they license. This is general information, not legal, tax or financial advice. Confirm with your own counsel or data protection adviser before acting.
What does purpose limitation require?
GDPR principles say personal data must be collected for specified, explicit and legitimate purposes and not further processed in a way that is incompatible with them. A company that gathered mailbox data to run its business has a stated purpose. Handing a copy to a third party to train models is a different activity.
Three points follow from the text.
- The purpose is judged at collection, based on what was disclosed and what the person could reasonably expect.
- "Further processing" by someone else, such as a licensee, still has to be justified.
- Compatibility is assessed case by case. There is no list of automatically permitted reuses that covers AI training.
What are the Article 6(4) compatibility factors?
When the new purpose is not based on consent, Article 6(4) asks the controller to take into account, among other things, the factors in the table. The wording below paraphrases the regulation; read the official text for the exact language.
| Factor | The question to ask | How AI-training licensing usually scores |
|---|---|---|
| Link between purposes | Is the new use close to the original one? | Often weak: running a business and training a third party's model are far apart |
| Context of collection | What would the person have expected, given their relationship with the company? | Employees and customers rarely expect their messages to leave the company |
| Nature of the data | Does it include special categories or criminal-offense data? | Any such data weighs strongly against compatibility |
| Consequences for people | What could happen to them because of the new use? | Depends on whether the output could identify or affect them |
| Safeguards | Is the data encrypted, pseudonymized or otherwise protected? | Strong safeguards help but do not cure a weak link |
If a company cannot argue that the factors point to compatibility, the fallback is a fresh lawful basis, which for large mailbox or ticket archives is hard to obtain person by person.
How this plays out in common situations
| Situation | What to check | Typical outcome to confirm with counsel |
|---|---|---|
| EU-based employee mailboxes | Employee notice, works-council or national employment rules, purposes disclosed | Exclude from the license, or anonymize rigorously |
| Customer support tickets from EU contacts | Privacy notice at sign-up, ticket content, free-text identifiers | Exclude by customer location, or de-identify and document the method |
| Shared Slack or Teams channels with EU members | Workspace membership, guests, connected channels | Carve out channels and messages by membership |
| Sales call recordings with EU prospects | Recording notice, consent, voice as personal data | Usually exclude; see the page on employee voices in licensed recordings |
| Truly anonymous aggregates | Whether anyone can be identified by any reasonably likely means | May fall outside the regime if the standard is met, which is a high bar |
The wider framing of how buyers read this issue is in the guide to EU AI Act Article 10 data governance questions.
Why anonymization is not a shortcut
The GDPR applies to personal data, and pseudonymized data remains personal data. Replacing names with tokens while keeping the mapping or leaving rare facts in free text does not take records outside the regulation. Genuine anonymization means no one can be identified by means reasonably likely to be used, and that bar is demanding for unstructured text such as email and chat.
For that reason, many practitioners treat a location-based carve-out as the simpler path. It is not perfect either, because EU residents can appear in US systems, but it narrows the problem.
Disclosure and documentation good practice
Whatever route the company takes, a written record helps. Keep a short file that states the original purposes disclosed to employees and customers, the location filters used to exclude EU and UK records, the date the filters ran, and who approved them. If a buyer or regulator later asks how an EU record ended up in a dataset, that file is the first thing counsel will want.
Update employee and customer notices going forward so that future disclosures to licensees are described plainly. A notice written today does not repair past collection, but it improves the position for records created from now on.
How buyers look at this
AI developers and data buyers ask suppliers how rights were established, which records are in scope and which jurisdictions are excluded. A clean, documented exclusion is easier to review than an argument that a compatibility test was met. Related contract points, such as what a buyer may retain after a term ends, are covered in the page on the residuals clause.
If you want the wider background on why data provenance matters, see what AI training data is.
What a referral partner should and should not do
Partners introduce a company and give basic fit information. They do not review, export or describe confidential records, and they do not give legal advice. If a sponsor asks whether EU data can be included, the honest answer is that SourceX and the company's counsel decide that during rights review, and that exclusion is the usual starting point.
Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. The reward is paid only after the buyer pays and SourceX receives its fee; no reward is guaranteed.
Questions to put to your counsel
- Which of our notices described the purposes for our email, chat and ticket data?
- Which records relate to people in the EU or UK, and can we separate them reliably?
- Would an anonymization method satisfy the standard, and who documents it?
- Do any special categories of data appear in free text?
- What do we promise a buyer about exclusions, and how do we back it up?
Next step
Use the company fit checker for a preliminary, non-binding screen and read how the process works. If you know a US company with 50+ full-time employees at peak (contractors excluded) and years of records, register as a partner to introduce it.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
Is Article 6(4) the only route to reuse personal data for AI training?
No. A company can also rely on consent or on another lawful basis that fits the new purpose. In practice, consent from every employee or customer in a large archive is hard to obtain, which is why exclusion or genuine anonymization is more common for EU records.
Does pseudonymizing names make the data compatible for reuse?
Pseudonymization is a safeguard that counts in the compatibility analysis, but pseudonymized data is still personal data. It helps, though it does not by itself make a different purpose compatible. Counsel should assess the whole set of factors rather than treat tokenization as a fix.
Do US companies need to worry about this if all their staff are in the US?
Often much less. If a company has no EU establishment and no EU-based people in its records, there may be little EU personal data to analyze. The check is whether any mailbox, channel, ticket or customer record relates to people in the EU or UK.
Can an EU data subject object after a license is signed?
Individuals can have rights over their personal data, and a license does not remove them. That is a reason buyers and suppliers prefer to exclude EU records from the dataset up front rather than manage requests against delivered data later.
Who decides whether EU records can be included in a SourceX license?
The company decides with its own counsel, working with SourceX during rights review. Partners do not make or influence that call. Nothing is binding until the company agrees price and terms and signs, and redaction and exclusion rules are agreed before any work begins.
Related pages
- Employee voices in licensed recordings: what publicity and voice laws mean
- EU AI Act Article 10: the data governance questions buyers ask suppliers
- What is a residuals clause, and why does it matter in a data license?
- What is AI training data?
- Check Company Fit for Data Licensing
- How SourceX US company data referrals work
Free resources
- Business exit readiness assessment — A preliminary exit readiness score and checklist for advisors.
- SDE vs EBITDA calculator — Seller's discretionary earnings next to market-rate EBITDA.
- IRR calculator — Internal rate of return on annual cash flows.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment