How to redact PII from email archives at scale before licensing them

To redact PII from emails at scale, agree a written specification with the company first, then process each message in parts: headers, body, signature, quoted replies and attachments. Replace every person and identifier with the same token across the whole archive, deduplicate threads before review, and sign off only after a sampled manual check passes.

The short answer: treat every email as five documents

Redacting PII from an email archive at scale means splitting each message into headers, body, signature, quoted history and attachments, redacting each part with rules suited to it, and replacing every person and identifier with the same token wherever it appears. A single reply-all can carry a customer's mobile number three times: in a signature, deep in the quoted thread and on an attached invoice.

For an owner weighing a license, the reassuring point is that de-identification is a defined, written process agreed with the company before work begins. On SourceX deals, no email leaves the company before a license agreement has been signed by all parties and the company has given its go-ahead.

Email earns the effort because it records how decisions are made, escalated and resolved across a business; why business email archives are valuable for AI explains what buyers look for.

What has to be decided before processing starts?

  • Scope. Which mailboxes, departments and years are in, and which are excluded outright. Counsel threads, HR investigations, board correspondence and personal, non-business mail are the usual exclusions.
  • The privacy regimes in play. US and EU definitions of personal data differ, as the PII vs personal data comparison explains. The GDPR can apply to organizations outside the EU that offer goods or services to, or monitor the behavior of, people in the EU, so correspondence with EU contacts raises questions even for a US company. Where messages carry health information from a HIPAA context, HHS recognizes two de-identification methods: Safe Harbor, which removes 18 specified identifiers, and Expert Determination, where a qualified expert documents that re-identification risk is very small. California raises its own question, covered in whether licensing counts as a CCPA sale.
  • Security duties. A covered financial institution under the FTC's Safeguards Rule must maintain a written information security program that includes encrypting customer information in transit over external networks and at rest, so a working copy of an archive holding customer information needs the same protection. The Safeguards Rule guide for CPA firms shows how that plays out for one profession.
  • The specification. Entity list, token scheme, what stays, how each attachment type is treated, and the pass mark for the sample check.

Which parts of an email carry PII?

PartTypical PIIHandling rule
Address headers (From, To, Cc, Bcc, Reply-To)Names and addresses of every participantMap each address to a person token; keep a label for internal, customer or vendor
Routing headers and message IDsIP addresses, server names, device detailsDrop routing chains; issue new IDs that keep threads linked
Subject lineCustomer names, claim or order numbersSame rules as the body; redact once and reuse across the thread
BodyNames, phone numbers, addresses, account numbers, dates of birthPattern, entity-model and dictionary detection
Signature blockName, title, direct line, mobile, office address, profile linksDetect the block, replace contact details, keep the job title if the spec allows
Quoted replies and forwardsEarlier messages, often from external partiesRedact each unique segment once, then reuse the result
AttachmentsInvoices, resumes, contracts, spreadsheets, scanned formsExtract text, OCR images, or exclude file types by rule
Calendar invitesAttendee lists, dial-in numbers, meeting linksTreat as headers plus body

How to redact an email archive, step by step

  1. Export and freeze a copy. Use the mail platform's eDiscovery or archive tools to export the agreed mailboxes in a standard format such as PST, MBOX or EML. Record message, mailbox and attachment counts so nothing is lost or added silently.
  2. Apply exclusions first. Remove excluded mailboxes, folders and correspondents before any redaction, so excluded material is never processed at all.
  3. Thread and deduplicate. Group messages into conversations and separate quoted history from new text. A thirty-message thread may add one new paragraph per reply; redacting each unique segment once is faster and more consistent than reprocessing the same quoted text again and again.
  4. Build the identity map. Merge header addresses, the employee directory and CRM contacts into one list of people and organizations, with name variants such as "Dana Ruiz", "Dana", "D. Ruiz" and "druiz@". Each person gets one token for the whole archive.
  5. Detect identifiers in the text. Run pattern rules for phone, card, account and ID formats; an entity model for names and places the map does not know; the identity map as a dictionary; and signature-block detection, since signatures follow a predictable shape at the end of each message.
  6. Handle attachments by type. Extract text from documents, OCR scanned PDFs and images, and apply the same rules. Where a file type is high risk and low value, such as exported customer lists, exclude it by rule instead of redacting it.
  7. Replace, then rebuild. Swap identifiers for tokens while keeping paragraph breaks, quote levels and timestamps, so a reader can still follow who said what and when.
  8. Check a stratified sample by hand. Draw messages across departments, years, thread lengths and attachment types. Reviewers log every miss by entity type; if misses exceed the pass mark in the specification, fix the rule and reprocess the full archive.
  9. Close out with a log. Write down which version of the specification and tooling ran, how the sample was drawn, what it found and who signed off. The identity map stays with the company: it is the key that reverses the tokens and never travels with the dataset.

Why do consistent tokens matter?

Consistent tokens keep an archive useful. If one customer is [PERSON-0412] in March and [PERSON-0977] in June, a model can no longer see that a single relationship ran through a dozen threads, a pricing dispute and a renewal.

The business content survives: who chased whom, about which invoice, and when. That is the part buyers training agents on real work care about.

Common mistakes and fixes

MistakeWhat goes wrongBetter approach
Redacting only the newest message in each threadQuoted history repeats the original identifiersSplit quoted text, redact each unique segment, propagate
A fresh token every time a name appearsRelationships and threads stop making senseOne archive-wide identity map
Ignoring Bcc, Reply-To and routing headersAddresses and IP numbers survive in fields nobody readsRebuild headers from the identity map; drop routing data
Skipping scanned attachmentsImage-only PDFs pass text filters untouchedOCR them or exclude the file type
Blacking out every capitalized wordProduct names, job titles and places that carry meaning disappearThe specification lists what stays; test for over-redaction in the sample
Deleting identifiers instead of tokenizingSentences lose their subject and thread structure breaksA typed token in place of each identifier
Shipping the identity map with the dataAnyone holding it can reverse the redactionThe map stays with the company

Illustrative example: a freight brokerage's sales and operations mail

Illustrative and fictional: a freight brokerage with 220 full-time employees agrees to license eight years of sales and operations email. Its specification excludes legal, HR and board mailboxes, tokenizes people, carriers and shippers, keeps lane descriptions, quote discussions and load outcomes, and excludes spreadsheet attachments.

Deduplication shrinks the text to process considerably, because most of each message is quoted history. The first sample shows driver mobile numbers surviving inside scanned rate confirmations; the team adds OCR for that document type, reprocesses the archive, and the second sample meets the pass mark.

What this means for owners and the partners who introduce them

For an owner, the takeaway is control. The company sets the exclusions, approves the specification and signs nothing until price and terms work; its archive stays its own, because the data is licensed, not sold.

Partners never export, forward or describe emails. They introduce the company and share basic fit facts. The baseline is a company based in the US that had 50+ full-time employees at peak (contractors excluded), has operated for years with records to show for it, holds the rights to what it would license and has an authorized sponsor, for example its CEO or CFO. A fit check for companies gives a quick, non-binding first screen, and how SourceX referrals work lays out the stages. Call recordings bring a different set of problems, covered in how to redact PII from call transcripts.

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. Rewards are payable only after the buyer pays and SourceX receives its fee, and no reward is guaranteed.

This is general information, not legal, tax or financial advice. Confirm privacy and security obligations with your own counsel before acting.

Next step

If an owner you know has years of business email and fits the baseline, register as a partner and introduce them. The company can also file its own application at sourcex.si/apply.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Should job titles and company names stay in redacted emails?

Often yes, if the specification allows it. Job titles show who approves, escalates and decides, which is part of what makes email useful for training agents. The company's own name and product names usually stay too. Customer, supplier and partner organization names are typically tokenized, because they can identify a relationship and may be covered by confidentiality terms in the underlying contracts.

Can tokenized emails be re-identified?

Sometimes, if enough context survives. A unique combination of role, location, dates and events can point to one person even without a name. That is why the specification considers indirect identifiers, why the sample check looks for them, and why the identity map that links tokens to real people stays with the company and never travels with the dataset.

How should spreadsheet attachments be handled?

With caution. Spreadsheets attached to email are often exports of customer lists, payroll or pricing, with personal data in every row and little workflow context. Many specifications exclude them by default and include only named exceptions, such as standard order templates, after their columns have been reviewed. Documents and PDFs that describe work usually carry more value for less risk.

Do newsletters, auto-replies and system notifications need to stay in the archive?

Usually not. Marketing newsletters, out-of-office replies and automated alerts add volume without showing how people work, and some system notifications contain customer records in bulk. Filtering them out early reduces the text to redact and review. Some notifications, such as approval requests from a workflow tool, may be worth keeping if the specification treats them like any other message.

Who keeps the map between tokens and real names?

The company. The identity map is the key that reverses the redaction, so it stays inside the company's controlled environment, under its own retention policy, and is not delivered with the dataset. Keeping it lets the company answer later questions, such as confirming that a particular person's messages were excluded, without exposing identities to anyone else.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment