How to redact PII from call transcripts and audio before licensing them
To redact PII from call transcripts, agree a written redaction specification first, then remove identifiers from the text and the matching audio segments together, normalize spoken numbers so card, account and phone digits are caught, replace speaker names with role tokens, and hand-check a stratified sample of calls before anything leaves the company.
The short answer: redact the words and the audio together
Redacting PII from a call transcript means removing every identifier from the text and from the matching stretch of audio, then proving the result with a sample that people check by hand. A transcript that shows [DOB-1] while the recording still plays the caller saying "June fourth, seventy-nine" has not been redacted at all.
The aim is to strip identity while keeping the shape of the conversation: who spoke when, what the caller wanted, what the agent tried and how the call ended. That structure is what makes support, billing and sales calls useful to AI developers training and evaluating agents that handle real customer work. Over-redaction destroys it as surely as under-redaction leaks it.
With SourceX, the company and SourceX agree the de-identification and redaction requirements before any work begins, and recordings or transcripts move only after an executed agreement and the company's authorization. The steps below show what a sound specification covers, so owners and the partners who introduce them know what to expect.
What needs to be settled before anyone redacts a call?
Redaction cleans up recordings the company was entitled to make and is entitled to license. It cannot fix either problem after the fact.
- Lawful capture. The federal Wiretap Act permits a private party to record a call it takes part in, or one where a party gave prior consent, unless the purpose is criminal or tortious. Some states demand every party's consent; California's Penal Code section 632 is one example for confidential communications. Confirm what notice callers heard, year by year.
- Ownership of the calls. A contact center answering calls for clients may be bound by those clients' contracts; the guide to confidentiality clause use restrictions explains what to look for.
- A written redaction specification: entity list, replacement tokens, audio treatment, what stays in, and the pass mark for quality checks.
- An archive inventory: lines or queues, years covered, file formats, whether recordings are mono or two-channel, and which speech-to-text engine produced any stored text.
- A controlled workspace where recordings are processed inside the company's environment or one it has approved.
The overview of how company data is anonymized before AI licensing places call redaction inside the wider de-identification process.
Audio bleeping or transcript redaction: which does the dataset need?
It depends on what the buyer will train. A text-only deliverable needs transcript redaction; anything that ships audio needs both, with identical spans removed from each.
| Deliverable | What is removed | Main residual risk | Typical use |
|---|---|---|---|
| Redacted transcript only | Identifier spans replaced with typed tokens | Identifiers the speech engine misheard and wrote as ordinary words | Conversation flow, intent and resolution modeling |
| Audio with masked spans | The same time ranges replaced with silence or a tone of equal length | The voice itself, plus background speech | Speech recognition and voice agent training |
| Aligned transcript and audio | Matching spans in both, linked by word timestamps | Timestamp drift that leaves a syllable or digit audible | Multimodal agent training and evaluation |
| Derived records only | Raw speech dropped; intents, summaries, dispositions and handle times kept | Free-text summaries that quote the caller | Evaluation sets and workflow analysis |
Two details matter. Replace a span with silence or a tone of the same duration instead of cutting it out, so every later timestamp in the call stays valid. And masking words does not anonymize a voice; the guide to Illinois BIPA and call recordings covers when voice data can raise biometric questions.
How to redact a call transcript, step by step
- Define entity classes and tokens. Typical classes are person names, phone numbers, email addresses, street addresses, dates of birth, account, policy and order numbers, payment card data, government ID numbers, security answers and health details. Give each a typed token such as [CUSTOMER-1] or [ACCOUNT-1], numbered consistently within the call, and list what stays: product names, the company's own name, service dates and outcome codes.
- Start from the best transcript available. Re-transcribe with word-level timestamps and speaker diarization if the stored text lacks them. Two-channel recordings, with agent and caller on separate tracks, make speaker attribution far more reliable than diarizing a mixed mono file.
- Normalize speech before detection. Convert spoken numbers to digits ("four one one one" becomes 4111, "double five" becomes 55, "oh" becomes 0), join digit runs that a pause split across two segments, and collapse spelled-out words ("S as in Sam") into the word they spell.
- Detect in layers. Combine pattern rules for structured numbers, with a Luhn check for card numbers, a named-entity model for names and places, context triggers ("my date of birth is", "the last four are", "my mother's maiden name") and a lookup of the caller's own details from the CRM record linked to the call, which catches unusual names a model misses.
- Replace speaker labels with roles. Diarization output and contact-center exports often label speakers by agent name or login. Rename them AGENT and CUSTOMER, or CALLER-1 and CALLER-2 for transfers and conference calls, so no real name survives in the speaker field.
- Carry the spans into the audio. Map each redacted word range back to its timestamps, pad a short margin on both sides to absorb drift, and mask that range in the audio file. Then run speech-to-text again on the masked audio: any identifier it still recognizes is a miss.
- Clean metadata and side files. File names often contain the caller's phone number or an account ID; call detail records carry ANI, agent IDs and queue notes; QA scorecards and wrap-up notes hold free text. Each needs the same specification.
- Hand-check a stratified sample. Draw calls across every queue, year, call length and transcription-confidence band, including transfers and calls with long holds. Reviewers listen and read, log each miss by entity class, and the team fixes the rule and reprocesses the whole archive, not just the sampled calls.
- Keep a redaction log. Record the specification version, tool versions, sample design, miss counts and sign-off, so the company can show what was done if a buyer or its own counsel asks.
Where do numbers read aloud slip through?
Spoken numbers are the most frequent miss because detection rules are written for typed text. Test the sample for each of these patterns:
| How it is said | What the transcript often shows | Fix |
|---|---|---|
| Card or account number read in groups with pauses | Digits split across two segments, or "for" and "to" instead of 4 and 2 | Join adjacent segments and map homophones before matching |
| Repeated digits | "double five", "triple seven" | Expand repetitions during normalization |
| Agent reads the number back to confirm | The same digits in a second speaker turn, sometimes partial | Redact both turns and partial repeats such as "ending in" |
| Email address spelled out | "j dot ruiz at example dot com" | Rebuild the address from its spoken parts, then match |
| Date of birth in casual form | "third of March, eighty-two" | Parse spoken date forms, not just numeric dates |
| Keypad entry read back by the phone system | An account number inside an IVR prompt | Check IVR segments and keypad logs, not only live speech |
Card numbers deserve their own controls rather than one line in the entity list; PCI DSS and call recordings explains how card data is handled in recorded calls. Email archives fail in different places, such as quoted threads and signature blocks, which the guide on redacting PII from email archives covers.
Common mistakes and how to fix them
| Mistake | Why it hurts | Fix |
|---|---|---|
| Redacting the text and shipping the original audio | The recording still carries every identifier | Mask audio from the same spans, then re-transcribe to verify |
| One generic [REDACTED] token for everything | A model cannot tell a name from an account number, and the dialogue stops making sense | Typed, numbered tokens that stay consistent within each call |
| Removing outcomes along with identities | Dispositions, refund decisions and escalation notes are what buyers value | List in the specification what must stay |
| Trusting the vendor's confidence scores | Poor audio produces confident errors | Sample by confidence band and listen |
| Sampling only clean, short calls | Misses cluster in long, noisy and transferred calls | Stratify the sample across all of these |
| Forgetting file names and exports | Phone numbers and account IDs leak through metadata | Apply the specification to every field that ships |
Illustrative example: a home services dispatch line
Illustrative and fictional: a home services company with 160 full-time employees holds seven years of two-channel recordings from its dispatch and billing lines. Its specification covers names, phone numbers, street addresses, gate and alarm codes, card data and technician employee IDs, and keeps job types, appointment outcomes and callback reasons.
The first QA sample finds that gate codes read aloud pass straight through, because no rule expects a short code after the word "gate". The team adds context triggers for "gate", "alarm" and "lockbox", reprocesses the full archive and draws a fresh sample, which meets the pass mark. The redaction log records both rounds.
How partners should talk about call recordings
Partners introduce the company and share basic fit information; they never request, export or listen to recordings. A useful first conversation stays at the level of the archive, not its contents.
Fair questions for a partner: how many years of recordings exist, which platform stores them, whether callers hear a recording notice, and whether the calls are the company's own or handled for clients. To qualify, the company should be US-based, with 50+ full-time employees at peak (contractors excluded), a documented operating history going back several years, clear rights to the recordings and an owner or executive willing to sponsor the deal. The company fit checker gives a preliminary, non-binding read, and how it works walks through each stage from introduction to payment.
Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. Rewards are payable only after the buyer pays and SourceX receives its fee, and no reward is guaranteed.
This is general information, not legal, tax or financial advice. Confirm recording-consent and privacy questions with your own counsel before acting.
Next step
If a company you know keeps years of recorded calls and meets the baseline, register as a partner and make the introduction. Owners who prefer to start the process themselves can use sourcex.si/apply.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
Can automatic redaction tools handle call transcripts without human review?
Automated detection does most of the work, but it should not be the last step. Speech-to-text errors, numbers split across pauses and unusual names all produce misses that tools report with high confidence. A stratified manual sample, with reviewers both reading and listening, is how a company finds those misses and shows that the final archive meets the pass mark set in its redaction specification.
Should agent names be redacted as well as customer names?
Usually yes. Agents introduce themselves on most calls, and their names, logins and employee IDs also appear in metadata and QA scorecards. Replacing them with role tokens such as AGENT keeps the conversation readable while protecting staff. Some specifications keep a consistent agent token across calls so buyers can see how different agents handle similar problems without learning who they are.
Is a redacted recording anonymous if the voice is still audible?
Not necessarily. Masking spoken identifiers removes what was said, but a voice can still identify a person, and a caller can sometimes be recognized from the details of an unusual problem. That is why some datasets ship transcripts only, or derived records, and why voice-related laws in some states deserve their own review with counsel before audio is included.
What about card numbers captured before the company paused recording during payments?
Older recordings may contain full card numbers, expiry dates and security codes read aloud. Those segments need dedicated detection and audio masking, and some companies simply exclude the affected years or billing queues. Card data carries its own security obligations, so the redaction specification should state how such segments are found, masked and verified before those recordings are considered for licensing.
Can calls in languages other than English be included?
They can be considered, but each language needs its own transcription quality check, entity model and spoken-number rules, because people read out digits and dates differently. Many companies start with their English-language queues and add other languages only if a separate sample shows the redaction quality matches. Records that are primarily in English remain the strongest starting point.
Does redaction reduce what the recordings are worth to AI buyers?
Good redaction removes identity, not substance. Buyers training and evaluating agents care about the request, the steps taken, the tools used and the outcome, and typed tokens keep that structure intact. Over-redaction is the real risk: stripping product names, service dates or resolution codes makes a call harder to learn from, so the specification should say what must stay.
Related pages
- How confidentiality clause use restrictions decide what a company can license
- How is company data anonymized before AI licensing?
- Illinois BIPA and call recordings: when voice data becomes biometric
- PCI DSS and call recordings: how to remove card data before licensing
- How to redact PII from email archives at scale before licensing them
- Check Company Fit for Data Licensing
Free resources
- NPV calculator — Net present value with a discounted cash flow table.
- Time value of money calculator — Future and present value with optional regular payments.
- Business DSCR calculator — Debt service coverage from cash flow and loan terms.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment