Slack and Microsoft Teams workspace archives

Multi-year exports of company Slack or Microsoft Teams workspaces: channels, threads, reactions, shared files and bot messages, with consistent pseudonymous user IDs. Labs use them to train agents that follow workplace context, answer questions over internal discussion and hand off work.

Last updated October 3, 2026

What a record contains

One channel: messages in order with thread structure, edits, reactions, file references and channel metadata such as purpose and membership over time.

FieldWhat it holds
channel_id, channel_typePublic or private, with purpose
message_id, thread_tsThread structure
userPseudonymous ID, consistent across systems if requested
text, tsMessage body and time
reactions[], files[]Reactions and file references
botWhether an integration posted it
Illustrative record. The values are made up to show the shape of the data.
{
  "channel": {"id": "CH-17", "type": "public", "purpose": "payments on-call"},
  "messages": [
    {"user": "eng_04", "ts": "2024-11-02T03:14:09Z", "text": "Settlement job failed again, same timeout as last week"},
    {"user": "eng_11", "ts": "2024-11-02T03:16:40Z", "thread_of": "2024-11-02T03:14:09Z",
     "text": "Rolling back the batch size change, see PAY-871", "reactions": [{"name": "+1", "count": 2}]}
  ]
}

How AI labs use it

Workplace agents
Train assistants that read threads, find decisions and answer follow-up questions.
Long-horizon context
Years of conversation around the same projects and people.
Retrieval evaluation
Questions whose answers are buried in real threads.

Typical preparation requirements

Agreed with the supplier before any work begins. Typical requirements include:

  • Private channels and direct messages included only where agreed and allowed by the company’s policies
  • HR, legal and personal channels excluded
  • Secrets and credentials scanned for and removed
  • Customer names and other personal data redacted or pseudonymized

Every dataset has a documented owner and confirmed licensing rights. See data governance on sourcex.si.

What makes a strong package

  • Linked to code, tickets or documents from the same company
  • Engineering and operations channels with long threads
  • A stable channel structure over years

Compared with public datasets

Public sets such as Ubuntu Dialogue Corpus and Software-related Slack Chats with Disentangled Conversations are useful references, but limited as enterprise training data. The chat and email category page compares them with licensed data.

Who typically holds it

  • Software companies
  • Remote-first teams
  • Agencies
  • Startups, including wound-down ones

Know a company like this?

Introduce the company to SourceX. If its data deal closes, you can earn up to $100,000 in referral fees, paid after the buyer accepts the data and SourceX receives payment.

Refer a company

Questions

Are direct messages included?

Only where the agreement scopes them in and the company’s policies and notices allow it.

Can Slack users be matched to Jira or GitHub users?

You can ask for one pseudonymous ID per person across systems, so threads, tickets and commits line up, where the agreed de-identification allows it.

Need this data for a model?

Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.