Helpdesk ticket threads with resolutions

Real customer support tickets from operating companies, sourced to order: the full thread between customer and agents, internal notes, tags, status changes and how the issue was resolved. AI labs use them to train and evaluate support agents on real policies and edge cases. Support archives commonly hold 2–10 years of tickets.

Last updated October 3, 2026

What a record contains

One ticket: every public reply and internal note in order, with channel, tags, priority, assignee group, status changes, timestamps and the closing resolution.

FieldWhat it holds
ticket_idPseudonymous ID, stable across the package
channelEmail, web form, chat or phone
messages[]Role (customer, agent, internal note), body and timestamp
tags, priority, groupAs the support team set them
status_history[]Each status change with its time
resolutionResolution code and closing summary
csatSatisfaction rating, where collected
Illustrative record. The values are made up to show the shape of the data.
{
  "ticket_id": "T-58213",
  "channel": "email",
  "tags": ["billing", "refund_request"],
  "priority": "normal",
  "messages": [
    {"role": "customer", "at": "2024-03-04T09:12Z", "body": "I was charged twice for March. Order [ORDER_ID]."},
    {"role": "internal_note", "at": "2024-03-04T09:40Z", "body": "Duplicate charge confirmed in billing. Refund one, policy 4.2."},
    {"role": "agent", "at": "2024-03-04T09:44Z", "body": "Sorry about that. I've refunded the duplicate charge..."}
  ],
  "resolution": {"code": "refund_issued", "solved_at": "2024-03-04T09:45Z"},
  "csat": 5
}

How AI labs use it

Fine-tuning support agents
Train on how experienced agents resolve real issues, including policy exceptions and escalations.
Outcome-based rewards
Resolution codes, reopen events and CSAT give reward signals for reinforcement learning.
Private evaluation
Build held-out evals from tickets that aren’t on the public web, in the domain you care about.
Triage and routing
Tags, priorities and assignee groups label intent and urgency.

Typical preparation requirements

Agreed with the supplier before any work begins. Typical requirements include:

  • Names, emails, phone numbers, order and account numbers redacted or replaced with consistent placeholders
  • Attachments containing personal data removed
  • Client ownership confirmed when support is outsourced
  • Scope, permitted use and de-identification requirements agreed before any work begins

Every dataset has a documented owner and confirmed licensing rights. See data governance on sourcex.si.

What makes a strong package

  • Several years of history with consistent tagging
  • Resolution codes or CSAT linked to each ticket
  • Internal notes kept, since they show reasoning customers never see
  • Macros and help-center articles from the same period

Compared with public datasets

Public sets such as ABCD and MultiWOZ are useful references, but limited as enterprise training data. The customer support category page compares them with licensed data.

Who typically holds it

  • Software and SaaS companies
  • E-commerce brands
  • Insurers
  • Fintech companies
  • Healthcare administration teams

Know a company like this?

Introduce the company to SourceX. If its data deal closes, you can earn up to $100,000 in referral fees, paid after the buyer accepts the data and SourceX receives payment.

Refer a company

Questions

How many tickets does a typical archive hold?

It depends on the company’s size and history. SourceX looks for companies with 20 or more full-time employees and several years of operations, and support archives commonly span 2–10 years. Put your minimum volume in the request.

Are internal notes included?

They can be. Internal notes often hold the most useful reasoning, and they go through the same de-identification steps agreed for the package.

Need this data for a model?

Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.