How to redact PII from call transcripts and audio before licensing them

To redact PII from call transcripts, agree a written redaction specification first, then remove identifiers from the text and the matching audio segments together, normalize spoken numbers so card, account and phone digits are caught, replace speaker names with role tokens, and hand-check a stratified sample of calls before anything leaves the company.

The short answer: redact the words and the audio together

Redacting PII from a call transcript means removing every identifier from the text and from the matching stretch of audio, then proving the result with a sample that people check by hand. A transcript that shows [DOB-1] while the recording still plays the caller saying "June fourth, seventy-nine" has not been redacted at all.

The aim is to strip identity while keeping the shape of the conversation: who spoke when, what the caller wanted, what the agent tried and how the call ended. That structure is what makes support, billing and sales calls useful to AI developers training and evaluating agents that handle real customer work. Over-redaction destroys it as surely as under-redaction leaks it.

With SourceX, the company and SourceX agree the de-identification and redaction requirements before any work begins, and recordings or transcripts move only after an executed agreement and the company's authorization. The steps below show what a sound specification covers, so owners and the partners who introduce them know what to expect.

What needs to be settled before anyone redacts a call?

Redaction cleans up recordings the company was entitled to make and is entitled to license. It cannot fix either problem after the fact.

  • Lawful capture. The federal Wiretap Act permits a private party to record a call it takes part in, or one where a party gave prior consent, unless the purpose is criminal or tortious. Some states demand every party's consent; California's Penal Code section 632 is one example for confidential communications. Confirm what notice callers heard, year by year.
  • Ownership of the calls. A contact center answering calls for clients may be bound by those clients' contracts; the guide to confidentiality clause use restrictions explains what to look for.
  • A written redaction specification: entity list, replacement tokens, audio treatment, what stays in, and the pass mark for quality checks.
  • An archive inventory: lines or queues, years covered, file formats, whether recordings are mono or two-channel, and which speech-to-text engine produced any stored text.
  • A controlled workspace where recordings are processed inside the company's environment or one it has approved.

The overview of how company data is anonymized before AI licensing places call redaction inside the wider de-identification process.

Audio bleeping or transcript redaction: which does the dataset need?

It depends on what the buyer will train. A text-only deliverable needs transcript redaction; anything that ships audio needs both, with identical spans removed from each.

DeliverableWhat is removedMain residual riskTypical use
Redacted transcript onlyIdentifier spans replaced with typed tokensIdentifiers the speech engine misheard and wrote as ordinary wordsConversation flow, intent and resolution modeling
Audio with masked spansThe same time ranges replaced with silence or a tone of equal lengthThe voice itself, plus background speechSpeech recognition and voice agent training
Aligned transcript and audioMatching spans in both, linked by word timestampsTimestamp drift that leaves a syllable or digit audibleMultimodal agent training and evaluation
Derived records onlyRaw speech dropped; intents, summaries, dispositions and handle times keptFree-text summaries that quote the callerEvaluation sets and workflow analysis

Two details matter. Replace a span with silence or a tone of the same duration instead of cutting it out, so every later timestamp in the call stays valid. And masking words does not anonymize a voice; the guide to Illinois BIPA and call recordings covers when voice data can raise biometric questions.

How to redact a call transcript, step by step

  1. Define entity classes and tokens. Typical classes are person names, phone numbers, email addresses, street addresses, dates of birth, account, policy and order numbers, payment card data, government ID numbers, security answers and health details. Give each a typed token such as [CUSTOMER-1] or [ACCOUNT-1], numbered consistently within the call, and list what stays: product names, the company's own name, service dates and outcome codes.
  2. Start from the best transcript available. Re-transcribe with word-level timestamps and speaker diarization if the stored text lacks them. Two-channel recordings, with agent and caller on separate tracks, make speaker attribution far more reliable than diarizing a mixed mono file.
  3. Normalize speech before detection. Convert spoken numbers to digits ("four one one one" becomes 4111, "double five" becomes 55, "oh" becomes 0), join digit runs that a pause split across two segments, and collapse spelled-out words ("S as in Sam") into the word they spell.
  4. Detect in layers. Combine pattern rules for structured numbers, with a Luhn check for card numbers, a named-entity model for names and places, context triggers ("my date of birth is", "the last four are", "my mother's maiden name") and a lookup of the caller's own details from the CRM record linked to the call, which catches unusual names a model misses.
  5. Replace speaker labels with roles. Diarization output and contact-center exports often label speakers by agent name or login. Rename them AGENT and CUSTOMER, or CALLER-1 and CALLER-2 for transfers and conference calls, so no real name survives in the speaker field.
  6. Carry the spans into the audio. Map each redacted word range back to its timestamps, pad a short margin on both sides to absorb drift, and mask that range in the audio file. Then run speech-to-text again on the masked audio: any identifier it still recognizes is a miss.
  7. Clean metadata and side files. File names often contain the caller's phone number or an account ID; call detail records carry ANI, agent IDs and queue notes; QA scorecards and wrap-up notes hold free text. Each needs the same specification.
  8. Hand-check a stratified sample. Draw calls across every queue, year, call length and transcription-confidence band, including transfers and calls with long holds. Reviewers listen and read, log each miss by entity class, and the team fixes the rule and reprocesses the whole archive, not just the sampled calls.
  9. Keep a redaction log. Record the specification version, tool versions, sample design, miss counts and sign-off, so the company can show what was done if a buyer or its own counsel asks.

Where do numbers read aloud slip through?

Spoken numbers are the most frequent miss because detection rules are written for typed text. Test the sample for each of these patterns:

How it is saidWhat the transcript often showsFix
Card or account number read in groups with pausesDigits split across two segments, or "for" and "to" instead of 4 and 2Join adjacent segments and map homophones before matching
Repeated digits"double five", "triple seven"Expand repetitions during normalization
Agent reads the number back to confirmThe same digits in a second speaker turn, sometimes partialRedact both turns and partial repeats such as "ending in"
Email address spelled out"j dot ruiz at example dot com"Rebuild the address from its spoken parts, then match
Date of birth in casual form"third of March, eighty-two"Parse spoken date forms, not just numeric dates
Keypad entry read back by the phone systemAn account number inside an IVR promptCheck IVR segments and keypad logs, not only live speech

Card numbers deserve their own controls rather than one line in the entity list; PCI DSS and call recordings explains how card data is handled in recorded calls. Email archives fail in different places, such as quoted threads and signature blocks, which the guide on redacting PII from email archives covers.

Common mistakes and how to fix them

MistakeWhy it hurtsFix
Redacting the text and shipping the original audioThe recording still carries every identifierMask audio from the same spans, then re-transcribe to verify
One generic [REDACTED] token for everythingA model cannot tell a name from an account number, and the dialogue stops making senseTyped, numbered tokens that stay consistent within each call
Removing outcomes along with identitiesDispositions, refund decisions and escalation notes are what buyers valueList in the specification what must stay
Trusting the vendor's confidence scoresPoor audio produces confident errorsSample by confidence band and listen
Sampling only clean, short callsMisses cluster in long, noisy and transferred callsStratify the sample across all of these
Forgetting file names and exportsPhone numbers and account IDs leak through metadataApply the specification to every field that ships

Illustrative example: a home services dispatch line

Illustrative and fictional: a home services company with 160 full-time employees holds seven years of two-channel recordings from its dispatch and billing lines. Its specification covers names, phone numbers, street addresses, gate and alarm codes, card data and technician employee IDs, and keeps job types, appointment outcomes and callback reasons.

The first QA sample finds that gate codes read aloud pass straight through, because no rule expects a short code after the word "gate". The team adds context triggers for "gate", "alarm" and "lockbox", reprocesses the full archive and draws a fresh sample, which meets the pass mark. The redaction log records both rounds.

How partners should talk about call recordings

Partners introduce the company and share basic fit information; they never request, export or listen to recordings. A useful first conversation stays at the level of the archive, not its contents.

Fair questions for a partner: how many years of recordings exist, which platform stores them, whether callers hear a recording notice, and whether the calls are the company's own or handled for clients. To qualify, the company should be US-based, with 50+ full-time employees at peak (contractors excluded), a documented operating history going back several years, clear rights to the recordings and an owner or executive willing to sponsor the deal. The company fit checker gives a preliminary, non-binding read, and how it works walks through each stage from introduction to payment.

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. Rewards are payable only after the buyer pays and SourceX receives its fee, and no reward is guaranteed.

This is general information, not legal, tax or financial advice. Confirm recording-consent and privacy questions with your own counsel before acting.

Next step

If a company you know keeps years of recorded calls and meets the baseline, register as a partner and make the introduction. Owners who prefer to start the process themselves can use sourcex.si/apply.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Can automatic redaction tools handle call transcripts without human review?

Automated detection does most of the work, but it should not be the last step. Speech-to-text errors, numbers split across pauses and unusual names all produce misses that tools report with high confidence. A stratified manual sample, with reviewers both reading and listening, is how a company finds those misses and shows that the final archive meets the pass mark set in its redaction specification.

Should agent names be redacted as well as customer names?

Usually yes. Agents introduce themselves on most calls, and their names, logins and employee IDs also appear in metadata and QA scorecards. Replacing them with role tokens such as AGENT keeps the conversation readable while protecting staff. Some specifications keep a consistent agent token across calls so buyers can see how different agents handle similar problems without learning who they are.

Is a redacted recording anonymous if the voice is still audible?

Not necessarily. Masking spoken identifiers removes what was said, but a voice can still identify a person, and a caller can sometimes be recognized from the details of an unusual problem. That is why some datasets ship transcripts only, or derived records, and why voice-related laws in some states deserve their own review with counsel before audio is included.

What about card numbers captured before the company paused recording during payments?

Older recordings may contain full card numbers, expiry dates and security codes read aloud. Those segments need dedicated detection and audio masking, and some companies simply exclude the affected years or billing queues. Card data carries its own security obligations, so the redaction specification should state how such segments are found, masked and verified before those recordings are considered for licensing.

Can calls in languages other than English be included?

They can be considered, but each language needs its own transcription quality check, entity model and spoken-number rules, because people read out digits and dates differently. Many companies start with their English-language queues and add other languages only if a separate sample shows the redaction quality matches. Records that are primarily in English remain the strongest starting point.

Does redaction reduce what the recordings are worth to AI buyers?

Good redaction removes identity, not substance. Buyers training and evaluating agents care about the request, the steps taken, the tools used and the outcome, and typed tokens keep that structure intact. Over-redaction is the real risk: stripping product names, service dates or resolution codes makes a call harder to learn from, so the specification should say what must stay.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment