Quasi-identifiers and re-identification risk in workplace data

Quasi-identifiers are ordinary fields, such as job title, office, project name and timestamp, that identify a person when combined. Removing names alone does not de-identify emails or tickets. Companies reduce the risk by generalizing precise values, suppressing rare combinations and testing group sizes before any dataset is licensed.

What is a quasi-identifier, and why does it matter before licensing?

A quasi-identifier is a piece of information that does not name anyone on its own but can point to a person when combined with other pieces. A job title, an office location, a project codename and a timestamp are all ordinary examples in workplace records.

The risk is not the obvious fields. A company that strips names and email addresses from a ticket archive can still leave behind enough context for a determined reader to work out who wrote a message. That is why a buyer's counsel, and the company's own, look at combinations of fields and not just a list of banned columns.

For a referral partner the practical point is simple: you never touch the records, but you can tell an owner that "we removed the names" is the start of the conversation, not the end of it.

How do quasi-identifiers work in emails, tickets and chat?

Re-identification usually happens by linkage: a reader matches a few attributes in the licensed dataset against something they already know or can look up. The classic textbook trio is a postal code, a birth date and a gender. Workplace data has its own trio.

Quasi-identifierWhere it appearsWhy it can single someone out
Job title plus teamEmail signatures, ticket assignee fields, Slack channel namesA "VP of Claims Automation" may be one person in the whole company
Office or siteMeeting invites, shipping tickets, badge or door logsA three-person branch makes every message from that site traceable
Project or customer codenameSubject lines, Jira epics, CRM opportunity namesPublic press or a customer's own announcement links the codename to named staff
TimestampMessage headers, call logs, ticket historyExact times match shift rosters, calendar entries or a known incident
Rare vocabularyFree-text notes and repliesDistinctive phrasing, product nicknames or a regional dialect act like a signature
Tool and access contextApprover fields, workflow routing, admin actionsOnly one person may hold an approval role for a given region

The HHS de-identification guidance is a useful reference point even outside healthcare, because it shows the two ways regulators think about the problem. Safe Harbor removes a fixed list of 18 identifiers, which include dates and small geographic units. Expert Determination has a qualified expert document that the risk of re-identification is very small. Most business datasets are not health records, so neither method is mandated for them, but the reasoning carries over: a list of fields to remove is a floor, and a risk judgment about the whole dataset is the ceiling.

This is general information, not legal, tax or financial advice. Confirm with your own counsel, tax adviser or professional body before acting.

How do generalization and suppression reduce the risk?

Two techniques do most of the work, and they trade usefulness for safety in different ways.

  • Generalization replaces a precise value with a broader one. An exact timestamp becomes a date, a date becomes a month, a branch address becomes a region, "Senior Claims Automation Lead" becomes "senior operations role".
  • Suppression removes a value or a whole record when it is too rare to generalize. A message from the only person at a two-person site can be dropped or have its site field blanked.

A third step, consistent pseudonymization, replaces each real identifier with a stable token so conversations still read as conversations. It keeps thread structure, which buyers value, but the mapping key must stay with the company and out of the delivered data.

How far to go is a decision made between the company and the buyer, with the company's counsel involved. Over-generalizing destroys the sequence of who-said-what-when that makes workplace records useful. Under-generalizing leaves linkage risk. At SourceX, de-identification and redaction requirements are agreed with the company before any work begins, and data is delivered only after an executed agreement and the company's authorization.

A quick screen: the group-size test

A simple rule of thumb helps non-lawyers see the problem: for any combination of fields that stays in the data, ask how many people share it. If the answer is one or two, treat the combination as identifying.

  1. Pick the three or four fields you plan to keep, such as department, site, role level and month.
  2. Count how many distinct people could produce each combination across the archive.
  3. Flag every combination shared by only a handful of people.
  4. Generalize one field in the flagged combinations, for example site to region, and count again.
  5. Suppress what is still too rare, and record the rule you used.

Illustrative: a fictional 180-person engineering firm keeps the fields "office", "role" and "week". In its Denver office only one project coordinator exists, so every message with that combination belongs to her. Changing "office" to "region" and "week" to "month" puts her in a group with several other coordinators, and the rare Denver-only messages are suppressed.

Mistakes that leave records identifiable

MistakeWhy it hurtsFix
Removing names but keeping signatures and footersSignatures carry title, phone, office and sometimes a photoStrip signature blocks or replace them with a role level
Keeping exact timestampsTimes match calendars and incident logsRound to date or month
Leaving codenames in subject linesPress releases and customer posts link codenames to staffReplace with category labels
Treating attachments as out of scopePDFs and slides hold the same names and datesInclude attachments in the same review
Deleting the rare records after samplingCounts were taken before suppression, so groups look bigger than they areRe-run the group-size count after every change

What does this mean for a referral partner?

You do not review, describe or handle any records. Your job is the introduction and basic fit information. What you can do is lower the company's anxiety by explaining that identification risk is a known, manageable step handled during the data inventory and terms stage, and that a contract can also bind the buyer: the no-re-identification clauses explainer covers who is liable when a licensee tries to link records back to people.

Two neighbouring topics come up in the same conversations. A residuals clause affects what a counterparty may keep in memory, and field-of-use restrictions limit what a buyer may do with the data. For companies in financial services, whether GLBA covers business customers is a separate, earlier question.

Rewards work the same way as for any referral: the partner earns 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, up to $100,000 per referred company, paid only after the buyer pays and SourceX receives its fee. No reward is guaranteed.

Next step

If you know a US company with 50+ full-time employees at peak (contractors excluded), years of records and an owner who would consider a license, run the company fit checker for a preliminary screen and read how the process works. Then register as a partner to make the introduction.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Is removing names and email addresses enough to de-identify workplace records?

Usually not. Signatures, job titles, office names, project codenames, timestamps and distinctive wording can still point to one person. Reviewers look at the combination of fields left in the data and ask how many people share each combination, then generalize or suppress anything too rare.

What is the difference between generalization and suppression?

Generalization makes a value less precise, for example turning an exact time into a month or an address into a region. Suppression removes a value or a whole record when it is too unusual to generalize safely. Most projects use both, plus consistent tokens so threads remain readable.

Does the HIPAA Safe Harbor list apply to ordinary business emails?

Not as a legal requirement, since HIPAA covers protected health information. It is still a helpful benchmark because it names dates, small geographic units and other identifiers. Treat it as a reference point, then decide with counsel what standard fits the company's own records and contracts.

Who decides how much to generalize before a dataset is licensed?

The company decides, with its counsel, and agrees the requirements with the licensing platform and the buyer before any work begins. Too little generalization leaves linkage risk, while too much strips the sequence and context that make business records useful for training and evaluation.

Can a contract replace technical de-identification?

No, they work together. Technical steps lower the chance that anyone can single out a person, and contract terms such as a ban on re-identification and limits on use add a legal remedy if a licensee tries. Companies normally want both before data leaves their control.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment