Internal meeting recordings and transcripts

Recordings and transcripts of internal meetings, such as standups, planning sessions and design reviews, with the agendas, notes and follow-up tasks that came out of them. Labs use them for long-context summarization, action-item extraction and multi-speaker speech recognition.

Last updated October 3, 2026

What a record contains

One meeting: recording or transcript with speaker turns, the calendar invite, the shared agenda or document, and tasks or tickets created afterwards.

FieldWhat it holds
meeting_idPseudonymous ID
typeStandup, planning, review, all-hands and so on
speakers[]Pseudonymous IDs, consistent across meetings
transcript[]Speaker-labeled turns with times
agenda_ref, notes_refLinked documents
follow_ups[]Tasks or tickets created from the meeting
Illustrative record. The values are made up to show the shape of the data.
{
  "meeting_id": "M-1290",
  "type": "sprint_planning",
  "duration_minutes": 47,
  "speakers": ["eng_04", "eng_11", "pm_02"],
  "follow_ups": [{"ticket": "PAY-882", "owner": "eng_11", "due": "2025-02-14"}],
  "transcript": [
    {"speaker": "pm_02", "start": 12.4, "text": "Let's keep the retry work in this sprint and push the export."}
  ]
}

How AI labs use it

Long-context summarization
Summarize hour-long, multi-speaker conversations.
Action-item extraction
Linked follow-up tasks show which commitments were real.
Speaker diarization
Multi-speaker audio with consistent speaker IDs.

Typical preparation requirements

Agreed with the supplier before any work begins. Typical requirements include:

  • Employee notice and consent checked
  • HR, legal and board meetings excluded
  • Names pseudonymized consistently across meetings

Every dataset has a documented owner and confirmed licensing rights. See data governance on sourcex.si.

What makes a strong package

  • Meetings linked to the tickets or documents they produced
  • Recurring series from the same team over months
  • Native transcripts plus audio

Compared with public datasets

Public sets such as AMI Meeting Corpus and MeetingBank are useful references, but limited as enterprise training data. The calls and meetings category page compares them with licensed data.

Who typically holds it

  • Software companies
  • Consulting firms
  • Agencies
  • Product and engineering teams

Know a company like this?

Introduce the company to SourceX. If its data deal closes, you can earn up to $100,000 in referral fees, paid after the buyer accepts the data and SourceX receives payment.

Refer a company

Questions

Are confidential meetings included?

The agreement scopes which meeting types are included. Expect HR, legal and board meetings to be excluded.

Need this data for a model?

Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.