Quasi-identifiers and re-identification risk in workplace data
Quasi-identifiers are ordinary fields, such as job title, office, project name and timestamp, that identify a person when combined. Removing names alone does not de-identify emails or tickets. Companies reduce the risk by generalizing precise values, suppressing rare combinations and testing group sizes before any dataset is licensed.
What is a quasi-identifier, and why does it matter before licensing?
A quasi-identifier is a piece of information that does not name anyone on its own but can point to a person when combined with other pieces. A job title, an office location, a project codename and a timestamp are all ordinary examples in workplace records.
The risk is not the obvious fields. A company that strips names and email addresses from a ticket archive can still leave behind enough context for a determined reader to work out who wrote a message. That is why a buyer's counsel, and the company's own, look at combinations of fields and not just a list of banned columns.
For a referral partner the practical point is simple: you never touch the records, but you can tell an owner that "we removed the names" is the start of the conversation, not the end of it.
How do quasi-identifiers work in emails, tickets and chat?
Re-identification usually happens by linkage: a reader matches a few attributes in the licensed dataset against something they already know or can look up. The classic textbook trio is a postal code, a birth date and a gender. Workplace data has its own trio.
| Quasi-identifier | Where it appears | Why it can single someone out |
|---|---|---|
| Job title plus team | Email signatures, ticket assignee fields, Slack channel names | A "VP of Claims Automation" may be one person in the whole company |
| Office or site | Meeting invites, shipping tickets, badge or door logs | A three-person branch makes every message from that site traceable |
| Project or customer codename | Subject lines, Jira epics, CRM opportunity names | Public press or a customer's own announcement links the codename to named staff |
| Timestamp | Message headers, call logs, ticket history | Exact times match shift rosters, calendar entries or a known incident |
| Rare vocabulary | Free-text notes and replies | Distinctive phrasing, product nicknames or a regional dialect act like a signature |
| Tool and access context | Approver fields, workflow routing, admin actions | Only one person may hold an approval role for a given region |
The HHS de-identification guidance is a useful reference point even outside healthcare, because it shows the two ways regulators think about the problem. Safe Harbor removes a fixed list of 18 identifiers, which include dates and small geographic units. Expert Determination has a qualified expert document that the risk of re-identification is very small. Most business datasets are not health records, so neither method is mandated for them, but the reasoning carries over: a list of fields to remove is a floor, and a risk judgment about the whole dataset is the ceiling.
This is general information, not legal, tax or financial advice. Confirm with your own counsel, tax adviser or professional body before acting.
How do generalization and suppression reduce the risk?
Two techniques do most of the work, and they trade usefulness for safety in different ways.
- Generalization replaces a precise value with a broader one. An exact timestamp becomes a date, a date becomes a month, a branch address becomes a region, "Senior Claims Automation Lead" becomes "senior operations role".
- Suppression removes a value or a whole record when it is too rare to generalize. A message from the only person at a two-person site can be dropped or have its site field blanked.
A third step, consistent pseudonymization, replaces each real identifier with a stable token so conversations still read as conversations. It keeps thread structure, which buyers value, but the mapping key must stay with the company and out of the delivered data.
How far to go is a decision made between the company and the buyer, with the company's counsel involved. Over-generalizing destroys the sequence of who-said-what-when that makes workplace records useful. Under-generalizing leaves linkage risk. At SourceX, de-identification and redaction requirements are agreed with the company before any work begins, and data is delivered only after an executed agreement and the company's authorization.
A quick screen: the group-size test
A simple rule of thumb helps non-lawyers see the problem: for any combination of fields that stays in the data, ask how many people share it. If the answer is one or two, treat the combination as identifying.
- Pick the three or four fields you plan to keep, such as department, site, role level and month.
- Count how many distinct people could produce each combination across the archive.
- Flag every combination shared by only a handful of people.
- Generalize one field in the flagged combinations, for example site to region, and count again.
- Suppress what is still too rare, and record the rule you used.
Illustrative: a fictional 180-person engineering firm keeps the fields "office", "role" and "week". In its Denver office only one project coordinator exists, so every message with that combination belongs to her. Changing "office" to "region" and "week" to "month" puts her in a group with several other coordinators, and the rare Denver-only messages are suppressed.
Mistakes that leave records identifiable
| Mistake | Why it hurts | Fix |
|---|---|---|
| Removing names but keeping signatures and footers | Signatures carry title, phone, office and sometimes a photo | Strip signature blocks or replace them with a role level |
| Keeping exact timestamps | Times match calendars and incident logs | Round to date or month |
| Leaving codenames in subject lines | Press releases and customer posts link codenames to staff | Replace with category labels |
| Treating attachments as out of scope | PDFs and slides hold the same names and dates | Include attachments in the same review |
| Deleting the rare records after sampling | Counts were taken before suppression, so groups look bigger than they are | Re-run the group-size count after every change |
What does this mean for a referral partner?
You do not review, describe or handle any records. Your job is the introduction and basic fit information. What you can do is lower the company's anxiety by explaining that identification risk is a known, manageable step handled during the data inventory and terms stage, and that a contract can also bind the buyer: the no-re-identification clauses explainer covers who is liable when a licensee tries to link records back to people.
Two neighbouring topics come up in the same conversations. A residuals clause affects what a counterparty may keep in memory, and field-of-use restrictions limit what a buyer may do with the data. For companies in financial services, whether GLBA covers business customers is a separate, earlier question.
Rewards work the same way as for any referral: the partner earns 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, up to $100,000 per referred company, paid only after the buyer pays and SourceX receives its fee. No reward is guaranteed.
Next step
If you know a US company with 50+ full-time employees at peak (contractors excluded), years of records and an owner who would consider a license, run the company fit checker for a preliminary screen and read how the process works. Then register as a partner to make the introduction.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
Is removing names and email addresses enough to de-identify workplace records?
Usually not. Signatures, job titles, office names, project codenames, timestamps and distinctive wording can still point to one person. Reviewers look at the combination of fields left in the data and ask how many people share each combination, then generalize or suppress anything too rare.
What is the difference between generalization and suppression?
Generalization makes a value less precise, for example turning an exact time into a month or an address into a region. Suppression removes a value or a whole record when it is too unusual to generalize safely. Most projects use both, plus consistent tokens so threads remain readable.
Does the HIPAA Safe Harbor list apply to ordinary business emails?
Not as a legal requirement, since HIPAA covers protected health information. It is still a helpful benchmark because it names dates, small geographic units and other identifiers. Treat it as a reference point, then decide with counsel what standard fits the company's own records and contracts.
Who decides how much to generalize before a dataset is licensed?
The company decides, with its counsel, and agrees the requirements with the licensing platform and the buyer before any work begins. Too little generalization leaves linkage risk, while too much strips the sequence and context that make business records useful for training and evaluation.
Can a contract replace technical de-identification?
No, they work together. Technical steps lower the chance that anyone can single out a person, and contract terms such as a ban on re-identification and limits on use add a legal remedy if a licensee tries. Companies normally want both before data leaves their control.
Related pages
- No-re-identification clauses in data licenses and who is liable
- What is a residuals clause, and why does it matter in a data license?
- Field-of-use restrictions in data licenses explained
- Does GLBA cover business customers' data, and what can a lender license?
- Check Company Fit for Data Licensing
- How SourceX US company data referrals work
Free resources
- MCP ROI calculator — Estimate hours saved, implied savings and first-year ROI from MCP.
- Business exit readiness assessment — A preliminary exit readiness score and checklist for advisors.
- SDE vs EBITDA calculator — Seller's discretionary earnings next to market-rate EBITDA.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment