De-identifying free text under HIPAA: emails, notes and tickets
De-identify free text under HIPAA by classifying sources, excluding patient-facing queues, running entity detection, replacing identifiers consistently, reviewing samples by hand, and choosing Safe Harbor or Expert Determination. Free text fails quietly because identifiers hide in sentences. SourceX agrees redaction requirements with the company before any work begins.
Why free text is where de-identification fails quietly
Free text fails because identifiers hide in sentences: a name in the middle of a note, a date written as "last Tuesday after the March 3rd visit", a hospital room in a ticket. Structured fields can be dropped by column; prose has to be read. A workable workflow for emails, notes and tickets has three stages: automated detection, human review of a sample, and, where the data stays high-risk, a written expert determination.
HHS describes the two routes in its guidance on methods for de-identification. Under Expert Determination, a qualified expert determines and documents that the risk of re-identification is very small. Under Safe Harbor, 18 specified identifiers are removed and the organization has no actual knowledge that what remains could identify a person. Health information de-identified by either method is no longer protected health information under the Privacy Rule.
This is general information, not legal, tax or financial advice. Confirm with your own counsel or privacy officer before acting.
Prerequisites before any text is touched
- The company has confirmed whether it is a HIPAA covered entity, a business associate or neither. A healthcare administration company with no patient data may not be in scope at all.
- An owner is named for the project: a privacy or compliance lead, not only IT.
- The systems in scope are listed (email, ticketing, chat, call-note fields) with date ranges.
- Counsel has said whether de-identification, patient authorization or exclusion is the route for each record type.
- A working copy exists. The live system is never edited.
Partners do none of this work. They check only that the company has a person who will own it. The records themselves never pass through a referral partner.
How to de-identify free text, step by step
- Classify the text sources. Separate clinical notes, patient emails, billing correspondence, operational tickets and internal chat. Each has a different identifier profile, and some may be excluded entirely.
- Cut scope before you scrub. Drop whole channels, mailboxes and ticket queues that are patient-facing if the plan is to license only operational material. Exclusion is cheaper and safer than removal.
- Run automated detection. Use pattern rules for phone numbers, emails, record numbers and dates, plus a named-entity tool for names, places and organizations. Expect misses on misspellings, nicknames and abbreviations.
- Replace, do not just delete. Swap identifiers for consistent placeholders such as {PERSON_1}, so a thread still reads as one conversation. Consistency across a thread matters for usefulness to AI buyers.
- Handle dates deliberately. Safe Harbor treats all date elements except year as identifiers, so dates in prose need shifting or generalizing, and any text where age or timing reveals more needs a second look.
- Review a sample by hand. Have someone who knows the business read a random sample from each source, scoring leaks per hundred records. Set the acceptable miss rate with the expert or counsel up front.
- Check for actual knowledge. If a reviewer notices that a note describes a recognizable local event or a very rare condition, the Safe Harbor condition on actual knowledge is in question. Escalate it.
- Decide Safe Harbor or Expert Determination. Pure Safe Harbor on free text is hard to defend. Many teams commission an expert to assess residual risk across the whole corpus and document the conclusion.
- Freeze and log. Record the tool versions, rules, reviewer sign-offs and sample results so the process can be shown later.
See how HIPAA expert determination works for the process, report and choosing an expert.
Common mistakes with unstructured text
| Mistake | Why it hurts | Fix |
|---|---|---|
| Relying on find-and-replace of a name list | Misses nicknames, typos and names not on the list | Use entity detection plus manual sampling |
| Removing names but leaving exact dates | Dates other than year are identifiers under Safe Harbor | Shift or generalize dates in prose and metadata |
| Scrubbing the body but not the subject line, attachments or file names | Identifiers sit in metadata | Include headers, attachment names and embedded images in scope |
| Treating a pass rate on the first batch as proof | Different sources leak differently | Sample each source separately |
| Skipping the "actual knowledge" check | A rare condition plus a small town can identify a person | Add a reviewer escalation rule |
| Assuming ticket text is not health information | Support tickets from clinics or patients can contain it | Ask counsel to classify each queue |
Operational tickets are easier than clinical text but share the pattern. The support ticket anonymization guide covers the field-by-field approach, and the CPRA sensitive information explainer covers the California layer on message contents.
Illustrative example
Illustrative: a fictional 180-person medical billing services firm keeps eight years of support tickets from client practices. Tickets are written by billing staff and sometimes quote patient names, visit dates and claim numbers. The firm's counsel decides to exclude the queue labelled "patient disputes" completely, run detection on the other queues, replace names and numbers with consistent placeholders, and engage an expert to assess the remaining corpus. The firm also checks its client contracts, because some tickets belong to the practices. Only after those steps does the firm talk about price and terms.
Red flags that stop the project
- The corpus is mostly clinical notes and no one can say who authorized their use.
- Reviewers keep finding names after the second pass, which means the tool is not fit for this text.
- The company's client contracts say the notes belong to the client practices.
- No one is willing to sign off on the sample results.
Any one of these means the records are not ready, and the honest answer to the owner is "not yet".
Who owns what
| Task | Owner | Partner role |
|---|---|---|
| HIPAA status and route decision | Company counsel or privacy officer | None |
| Detection and redaction | Company, with SourceX workflow support | None |
| Expert determination | Independent qualified expert engaged for the company | None |
| Introduction and basic fit information | Referral partner | Makes the introduction only |
SourceX agrees de-identification and redaction requirements with the company before any work begins, and data is delivered only after an executed agreement and the company's authorization. Records that are mainly protected health information without HIPAA authorization or de-identification are a red flag for the program.
Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. The reward is paid only after the buyer pays and SourceX receives its fee; an introduction, meeting or signed agreement alone does not trigger payment, and no reward is guaranteed.
Next step
Use the company fit checker for a first screen, read how SourceX referrals work, and then register as a partner to introduce a healthcare-adjacent company whose records are mainly administrative.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
Is Safe Harbor enough for emails and clinical notes?
Safe Harbor requires removing 18 specified identifiers and having no actual knowledge that the rest could identify someone. In free text, that is difficult to demonstrate because identifiers sit inside sentences. Many organizations choose Expert Determination for prose-heavy corpora. Counsel or a privacy officer should decide which route fits the specific records.
Can automated tools alone de-identify free text?
Tools find most patterned identifiers, but they miss nicknames, misspellings and context that points to a person. Use them as the first pass, then review a sample from each source and record the miss rate. The acceptable threshold is a decision for the company's expert or counsel, not a default setting.
Do all healthcare company records count as protected health information?
No. HIPAA applies to covered entities and business associates and to health information they hold. A healthcare administration business may hold mostly operational records such as vendor contracts and internal SOPs. Counsel should classify each record type. Records that are mainly PHI without authorization or de-identification are a red flag for the program.
What does an expert determination produce?
The HHS guidance says a qualified expert determines that the risk of re-identification is very small and documents the methods and results. The practical output is a written determination tied to a specific dataset and conditions. A different dataset or a new release usually needs a fresh look.
Does a referral partner need to review any of this text?
No. Partners make introductions and give basic fit information, and they never export, upload or describe confidential records. If an owner offers to send samples, decline politely and point them to SourceX, where requirements are agreed with the company before any work starts.
Related pages
Free resources
- MOIC calculator — Multiple on invested capital from realized and unrealized value.
- PDF bank statement to CSV converter — Turn Chase, Bank of America or Wells Fargo PDF statements into CSV, privately in your browser.
- Client data licensing eligibility checker — A transparent preliminary screen for one company.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment