What data should be excluded from AI training? A default exclusion list
By default, exclude HR and medical files, privileged legal threads, payment card data, government ID numbers, passwords and keys, client-restricted material and employees' personal messages from AI training data. Treat protected health information and financial institution customer records as excluded unless counsel confirms a lawful route, then de-identify everything that remains in scope.
Why start from a default exclusion list
The safest way to scope workplace records for AI training is to exclude sensitive categories by default and add back only what counsel approves. Start with HR and medical files, privileged legal threads, payment card data, government ID numbers, credentials, client-restricted material and personal messages, then strip the remaining identifying detail through de-identification.
A default list does two jobs. It speeds up scoping, because each system owner knows what to filter before any export is planned. And it answers the first objection most owners raise, which is whether licensing means handing over everything. It does not: in a SourceX process the company and SourceX agree the redaction and de-identification requirements before work starts, and nothing is handed over until an agreement is signed and the company authorizes delivery.
The default exclusion checklist
Tick each item once you have confirmed where it lives and how it will be filtered out.
People and HR
- Personnel files, performance reviews, disciplinary and investigation records.
- Medical, leave, disability, accommodation and workers' compensation records.
- Background check and drug test results.
- Payroll records with bank details, and benefits enrollment data.
Legal and privileged material
- Email and chat threads with outside or in-house counsel seeking or giving legal advice.
- Litigation files, legal hold notices and settlement negotiations.
- Board materials marked privileged or confidential.
Financial and payment data
- Full card numbers, card security codes, and bank account and routing numbers, wherever they appear in tickets, emails or call transcripts.
- Customer financial records held by a business that counts as a financial institution.
Identifiers and secrets
- Social Security, driver's license, passport and other government ID numbers.
- Passwords, API keys, tokens and private keys, including any left in code repositories, tickets and chat.
Client and third-party material
- Records the company holds on behalf of clients, such as an outsourcer's or agency's client data.
- Documents under confidentiality agreements that bar reuse.
- Licensed third-party content: research subscriptions, purchased reports and stock media.
Personal and sensitive communications
- Personal messages in work email and chat.
- Direct messages about health, family or personal finances.
- Information about children.
- Precise location history.
Which categories have a lawful route, and which never do
| Category | Default | Possible route, confirmed by counsel |
|---|---|---|
| Protected health information | Exclude | De-identification by Expert Determination or Safe Harbor, or a valid authorization |
| Financial institution customer data | Exclude | Only within GLBA notice and opt-out limits; rarely practical for a training license |
| Card numbers and security codes | Exclude | None; remove from every source |
| Government IDs and credentials | Exclude | None; remove from every source |
| Client-owned records | Exclude | The client's written consent |
| Privileged legal threads | Exclude | None for training use |
| HR files | Exclude | Aggregated, non-identifying workflow metadata only |
| Personal messages | Exclude | None |
| Business records with incidental names | Include after de-identification | Redaction rules agreed before work begins |
On health data, the HHS guidance on de-identification describes two methods: Expert Determination, where a qualified expert documents that the risk of re-identification is very small, and Safe Harbor, which removes 18 specified identifiers where there is no actual knowledge that the remainder could identify someone. Information de-identified by either method is no longer protected health information under the Privacy Rule. The HIPAA and AI training data guide goes further.
On financial data, the FTC's Gramm-Leach-Bliley Act guidance explains that the Privacy Rule requires notices to customers about information-sharing practices, and opt-out rights before sharing with certain nonaffiliated third parties. The guide to the FTC Safeguards Rule for CPA firms shows how widely the financial institution label can reach. State rules are moving too; see state AI laws in 2026 that touch training data.
This is general information, not legal, tax or financial advice. Confirm with your own counsel, tax adviser or professional body before acting.
How to use the results
| Finding | Meaning | What to do next |
|---|---|---|
| Excluded categories sit in separate systems or folders | Filtering is straightforward | Move on to the data inventory |
| Sensitive details are scattered through free text | A redaction pass is needed | Agree redaction rules and spot-check samples internally |
| Exclusions remove most of a system's content | That system adds little | Drop it and focus on cleaner systems |
| Client data cannot be separated from the company's own | Rights are unclear | Leave that system out of scope |
| Secrets found in repositories | A security issue first, a licensing issue second | Rotate the credentials before anything else |
Red flags that stop a scope
- Pressure to send everything and clean it up later.
- Records that are mainly consumer personal data with no licensing basis.
- Datasets already licensed to someone else for AI training.
- Content generated with AI in order to sell it.
- Archives deleted, or nobody able to export them.
How partners can use this list
Share the checklist with an owner who worries that licensing means exposing staff or clients. It shows that exclusions come first and that the company keeps control of scope. Do not collect samples, screenshots or exports yourself; a partner's role is the introduction plus a few basic facts about fit. The company fit checker offers a preliminary, non-binding read, and the explainer on what AI training data is gives owners useful background.
Next step
When an owner is comfortable with the exclusion-first approach, register as a partner to introduce them, or send them straight to sourcex.si/apply. Each stage from qualification to payment is set out on how it works.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
If records are de-identified, do the excluded categories still need removing?
Yes, for most of them. De-identification removes identifiers from records that are otherwise in scope; it does not make privileged advice, card numbers, credentials or client-owned material appropriate to license. Treat exclusion as the first filter and de-identification as the second. Health information is the main category where a recognized de-identification method can bring records back into scope.
Should all email be excluded from AI training data?
No. Business email is often among the most valuable records because it shows how decisions were made and work was coordinated. Exclude personal messages, privileged threads, HR matters and mailboxes that mainly serve clients, and remove identifiers from the rest. Many companies start with shared team mailboxes and project threads rather than individual inboxes.
Can source code be licensed once secrets are removed?
Often, once three checks pass: secrets are stripped and rotated, the code was written by employees or under written assignment, and it contains no client-owned code or third-party code under licenses that restrict reuse. Pull requests, review comments and linked tickets add the context that makes engineering records useful for training and evaluation.
Who decides the final exclusion list?
The company does, with its counsel, and the list is agreed with SourceX before any work begins. Partners have no role in this step. The default list on this page is a starting point: companies in regulated industries usually add categories, and some restore items once counsel confirms a lawful route such as recognized de-identification.
Does excluding sensitive data leave anything worth licensing?
Usually yes, for companies with deep operational history. The value for AI developers sits in how work gets done: tickets and resolutions, project histories, engineering reviews, approvals and their outcomes. Those records rarely depend on HR files, card numbers or personal messages, so excluding them reduces risk without removing most of the useful material.
Related pages
- HIPAA and AI training data: what the rules allow and what stays out of a license
- FTC Safeguards Rule for CPA firms: what it means for client records and referrals
- State AI laws in 2026 that touch AI training data: what to check before licensing
- Check Company Fit for Data Licensing
- What is AI training data?
- How SourceX US company data referrals work
Free resources
- MCP ROI calculator — Estimate hours saved, implied savings and first-year ROI from MCP.
- Business exit readiness assessment — A preliminary exit readiness score and checklist for advisors.
- SDE vs EBITDA calculator — Seller's discretionary earnings next to market-rate EBITDA.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment