Workplace chat and email for AI training
Workplace chat and email show how teams coordinate, decide and hand off work over months and years. SourceX looks for Slack and Microsoft Teams workspaces, email archives and complete archives from wound-down companies, for training agents that work inside real organizations.
Listings
| Data | Typical sources | Modality | Availability |
|---|---|---|---|
| Slack and Microsoft Teams workspace archives | Slack, Microsoft Teams, Google Chat | Text and files | On request |
| Enterprise email archives | Gmail and Google Workspace, Outlook and Microsoft 365, Front | Text, documents and calendar data | On request |
| Complete archives from wound-down companies | Slack, Jira, Linear | Text, code and documents | On request |
What’s included
- Channels and threads with reactions, edits, files and bot messages
- Email threads with attachments, distribution lists and calendar events
- Consistent pseudonymous IDs so the same person can be followed across systems
- Complete multi-system archives from companies that have shut down
Public datasets and licensed data
Public chat data comes from open-source help channels, and the only sizable real corporate email sets are Enron (1998–2002) and Avocado, which is licensed for research only. Licensed archives add private, current workplace communication tied to real projects, tickets and code.
| Public dataset | Released | What it contains | Limits for enterprise use | License |
|---|---|---|---|---|
| Ubuntu Dialogue Corpus | McGill University, 2015 | About 930,000 two-person dialogues from Ubuntu technical-support chat logs (2004–2015). | Public open-source help channels; later work found many extracted conversations were inaccurate. | See source |
| Software-related Slack Chats with Disentangled Conversations | University of Delaware and others, 2020 | 38,955 conversations from public Python, Clojure and Elm Slack communities (2017–2019). | Public community workspaces, not private company Slack. | Open (Zenodo) |
| Enron Email Dataset | Made public by FERC; prepared at Carnegie Mellon | About 500,000 messages from about 150 mostly senior Enron employees, 1998–2002. | One company, more than 20 years old, released during a federal investigation. | Shared as a research resource |
| Avocado Research Email Collection | LDC, 2015 | Email and attachments from 279 accounts at a defunct IT company. | Research-only license; commercial use needs the owners’ written permission. | LDC research license |
Preparation and rights
Preparation requirements are agreed with each supplier before any work begins. For this kind of data they typically include:
- Checking employee notices and the company’s policies before any communications are scoped in
- Excluding HR, legal, personal and privileged conversations
- Detecting and removing secrets and credentials pasted into chat
Every dataset has a documented owner and confirmed licensing rights, and is delivered only after an executed agreement and the supplier’s authorization. See data governance on sourcex.si.
Use cases
More on sourcex.si
Questions
Can a company license its Slack history to AI developers?
Often, yes, if it controls the workspace data, its policies and notices allow it, and personal, HR and privileged content is removed. The company approves scope and permitted use before anything is shared.
Are direct messages included?
Only where the agreement scopes them in and the company’s policies allow it. A package can also be limited to public and selected private channels.
Need workplace chat and email?
Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.