Contact-center call recordings with transcripts

Recorded customer service calls from contact centers, with transcripts, speaker labels, disposition codes and call metadata. Labs use them to train and test speech recognition and voice agents on real phone audio, with accents, holds and transfers.

Last updated October 3, 2026

What a record contains

One call: the audio (often with agent and caller on separate channels), a time-aligned transcript with speaker turns, queue, disposition code, handle and hold time, and transfers.

FieldWhat it holds
call_idPseudonymous ID
audioTelephone-quality audio, mono or dual-channel
transcript[]Speaker, start, end and text for each turn
queue, dispositionWhere the call went and how it was coded
handle_seconds, hold_secondsCall timing
transfers[]Each transfer with its time and target queue
Illustrative record. The values are made up to show the shape of the data.
{
  "call_id": "K-33019",
  "audio": {"channels": 2, "format": "wav"},
  "queue": "claims_status",
  "disposition": "status_provided",
  "handle_seconds": 412,
  "transcript": [
    {"speaker": "agent", "start": 0.0, "end": 4.1, "text": "Thanks for calling, this is [AGENT]."},
    {"speaker": "caller", "start": 4.3, "end": 9.8, "text": "Hi, I'm checking on a claim from last week."}
  ]
}

How AI labs use it

Speech recognition
Fine-tune and test ASR on real phone audio and accents.
Voice agents
Learn turn-taking, holds and transfers from real calls.
Summaries and wrap-up
Disposition codes and agent notes label what each call was about.

Typical preparation requirements

Agreed with the supplier before any work begins. Typical requirements include:

  • Recording notices and consent checked, including whether they cover sharing recordings with a third party for AI training
  • Names, card numbers and account details redacted in transcripts and silenced in audio
  • Voice data reviewed against biometric privacy laws, such as Illinois’ BIPA, before licensing
  • Client ownership confirmed for outsourced contact centers

Every dataset has a documented owner and confirmed licensing rights. See data governance on sourcex.si.

What makes a strong package

  • Dual-channel audio with agent and caller separated
  • Human-corrected transcripts for a subset
  • Disposition codes used consistently

Compared with public datasets

Public sets such as CallCenterEN and Switchboard-1 Release 2 are useful references, but limited as enterprise training data. The calls and meetings category page compares them with licensed data.

Who typically holds it

  • Insurers
  • Utilities and telecom providers
  • Healthcare administration teams
  • Fintech and lending companies
  • Travel companies

Know a company like this?

Introduce the company to SourceX. If its data deal closes, you can earn up to $100,000 in referral fees, paid after the buyer accepts the data and SourceX receives payment.

Refer a company

Questions

Can I license audio, or only transcripts?

Both can be scoped. Audio raises extra consent and biometric questions, so some packages are transcripts only.

Which languages are available?

Specify the languages and accents you need in your request; availability depends on the companies that hold the data.

Need this data for a model?

Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.