Customer support data for AI training

Customer support datasets pair real customer problems with the steps agents took to resolve them, often with satisfaction and QA scores attached. AI labs license them to train and evaluate support agents, agent-assist tools and intent models. SourceX looks for multi-year archives from helpdesks such as Zendesk, Intercom, Freshdesk and Salesforce Service Cloud.

Last updated October 3, 2026

Listings

DataTypical sourcesModalityAvailability
Helpdesk ticket threads with resolutions
Multi-turn tickets with internal notes, tags, status history and the final resolution.
Zendesk, Freshdesk, Salesforce Service CloudText with structured metadataOn request
Live chat and messaging transcripts
Real-time chats and in-app messages, including bot-to-human handoffs and outcomes.
Intercom, Zendesk Messaging, LivePersonText with structured metadataOn request
QA-scored support interactions
Tickets, chats or calls graded by human reviewers against a scoring rubric.
Zendesk QA, MaestroQA, PlayvoxText with scoresOn request

What’s included

  • Ticket threads with every customer reply, agent response and internal note
  • Live chat and in-app messaging, including bot-to-human handoffs
  • Resolution codes, tags, priorities and status history
  • CSAT ratings and QA rubric scores where the company collects them
  • Macros and canned responses from the same period

Public datasets and licensed data

None of the widely used public sets are real company helpdesk tickets. ABCD and MultiWOZ are role-played, Bitext is partly synthetic, and the Twitter set covers public replies under a non-commercial license. Licensed archives add real customers, real company policies, internal notes and how each issue was resolved.

Public datasetReleasedWhat it containsLimits for enterprise useLicense
ABCD (Action-Based Conversations Dataset)ASAPP Research, 202110,042 human-to-human support dialogues across 55 intents, with agent actions that follow written company policies.Agents and customers were crowdworkers in a role-play, not real staff and customers.MIT
Customer Support on TwitterThought Vector on Kaggle, 2017Over 3 million tweets and replies between customers and large brands.Public social replies rather than tickets, and a non-commercial license.CC BY-NC-SA 4.0
Bitext Customer Support LLM Chatbot Training DatasetBitext, 202426,872 question-and-answer pairs across 27 intents.Partly synthetic: real seed texts were expanded with generation tools.CDLA-Sharing-1.0
MultiWOZUniversity of Cambridge, 2018About 10,000 written task-oriented dialogues across travel and booking domains.Crowdsourced role-play, with one person playing the system.MIT

Preparation and rights

Preparation requirements are agreed with each supplier before any work begins. For this kind of data they typically include:

  • Redacting or replacing customer names, emails, phone numbers, order and account numbers with consistent placeholders
  • Removing attachments that contain personal data
  • Confirming the client company’s permission when support is outsourced, since the client usually owns the conversations

Every dataset has a documented owner and confirmed licensing rights, and is delivered only after an executed agreement and the supplier’s authorization. See data governance on sourcex.si.

Use cases

More on sourcex.si

Questions

What is a customer support dataset?

A collection of real support interactions, such as tickets, chats or calls, with the context around them: tags, status changes, the resolution and often a satisfaction or quality score. Licensed versions come from a company’s own helpdesk history.

Which helpdesk systems does SourceX source from?

Examples include Zendesk, Intercom, Freshdesk, Salesforce Service Cloud, ServiceNow and Jira Service Management. Name the systems you need in your request.

How far back do support archives go?

It depends on the company. Support archives commonly hold 2–10 years of tickets, and SourceX looks for companies with several years of documented operations.

Need customer support data?

Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.