Customer support data for AI training
Customer support datasets pair real customer problems with the steps agents took to resolve them, often with satisfaction and QA scores attached. AI labs license them to train and evaluate support agents, agent-assist tools and intent models. SourceX looks for multi-year archives from helpdesks such as Zendesk, Intercom, Freshdesk and Salesforce Service Cloud.
Listings
| Data | Typical sources | Modality | Availability |
|---|---|---|---|
| Helpdesk ticket threads with resolutions | Zendesk, Freshdesk, Salesforce Service Cloud | Text with structured metadata | On request |
| Live chat and messaging transcripts | Intercom, Zendesk Messaging, LivePerson | Text with structured metadata | On request |
| QA-scored support interactions | Zendesk QA, MaestroQA, Playvox | Text with scores | On request |
What’s included
- Ticket threads with every customer reply, agent response and internal note
- Live chat and in-app messaging, including bot-to-human handoffs
- Resolution codes, tags, priorities and status history
- CSAT ratings and QA rubric scores where the company collects them
- Macros and canned responses from the same period
Public datasets and licensed data
None of the widely used public sets are real company helpdesk tickets. ABCD and MultiWOZ are role-played, Bitext is partly synthetic, and the Twitter set covers public replies under a non-commercial license. Licensed archives add real customers, real company policies, internal notes and how each issue was resolved.
| Public dataset | Released | What it contains | Limits for enterprise use | License |
|---|---|---|---|---|
| ABCD (Action-Based Conversations Dataset) | ASAPP Research, 2021 | 10,042 human-to-human support dialogues across 55 intents, with agent actions that follow written company policies. | Agents and customers were crowdworkers in a role-play, not real staff and customers. | MIT |
| Customer Support on Twitter | Thought Vector on Kaggle, 2017 | Over 3 million tweets and replies between customers and large brands. | Public social replies rather than tickets, and a non-commercial license. | CC BY-NC-SA 4.0 |
| Bitext Customer Support LLM Chatbot Training Dataset | Bitext, 2024 | 26,872 question-and-answer pairs across 27 intents. | Partly synthetic: real seed texts were expanded with generation tools. | CDLA-Sharing-1.0 |
| MultiWOZ | University of Cambridge, 2018 | About 10,000 written task-oriented dialogues across travel and booking domains. | Crowdsourced role-play, with one person playing the system. | MIT |
Preparation and rights
Preparation requirements are agreed with each supplier before any work begins. For this kind of data they typically include:
- Redacting or replacing customer names, emails, phone numbers, order and account numbers with consistent placeholders
- Removing attachments that contain personal data
- Confirming the client company’s permission when support is outsourced, since the client usually owns the conversations
Every dataset has a documented owner and confirmed licensing rights, and is delivered only after an executed agreement and the supplier’s authorization. See data governance on sourcex.si.
Use cases
More on sourcex.si
Questions
What is a customer support dataset?
A collection of real support interactions, such as tickets, chats or calls, with the context around them: tags, status changes, the resolution and often a satisfaction or quality score. Licensed versions come from a company’s own helpdesk history.
Which helpdesk systems does SourceX source from?
Examples include Zendesk, Intercom, Freshdesk, Salesforce Service Cloud, ServiceNow and Jira Service Management. Name the systems you need in your request.
How far back do support archives go?
It depends on the company. Support archives commonly hold 2–10 years of tickets, and SourceX looks for companies with several years of documented operations.
Need customer support data?
Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.