Helpdesk ticket threads with resolutions
Real customer support tickets from operating companies, sourced to order: the full thread between customer and agents, internal notes, tags, status changes and how the issue was resolved. AI labs use them to train and evaluate support agents on real policies and edge cases. Support archives commonly hold 2–10 years of tickets.
What a record contains
One ticket: every public reply and internal note in order, with channel, tags, priority, assignee group, status changes, timestamps and the closing resolution.
| Field | What it holds |
|---|---|
ticket_id | Pseudonymous ID, stable across the package |
channel | Email, web form, chat or phone |
messages[] | Role (customer, agent, internal note), body and timestamp |
tags, priority, group | As the support team set them |
status_history[] | Each status change with its time |
resolution | Resolution code and closing summary |
csat | Satisfaction rating, where collected |
{
"ticket_id": "T-58213",
"channel": "email",
"tags": ["billing", "refund_request"],
"priority": "normal",
"messages": [
{"role": "customer", "at": "2024-03-04T09:12Z", "body": "I was charged twice for March. Order [ORDER_ID]."},
{"role": "internal_note", "at": "2024-03-04T09:40Z", "body": "Duplicate charge confirmed in billing. Refund one, policy 4.2."},
{"role": "agent", "at": "2024-03-04T09:44Z", "body": "Sorry about that. I've refunded the duplicate charge..."}
],
"resolution": {"code": "refund_issued", "solved_at": "2024-03-04T09:45Z"},
"csat": 5
}How AI labs use it
- Fine-tuning support agents
- Train on how experienced agents resolve real issues, including policy exceptions and escalations.
- Outcome-based rewards
- Resolution codes, reopen events and CSAT give reward signals for reinforcement learning.
- Private evaluation
- Build held-out evals from tickets that aren’t on the public web, in the domain you care about.
- Triage and routing
- Tags, priorities and assignee groups label intent and urgency.
Typical preparation requirements
Agreed with the supplier before any work begins. Typical requirements include:
- Names, emails, phone numbers, order and account numbers redacted or replaced with consistent placeholders
- Attachments containing personal data removed
- Client ownership confirmed when support is outsourced
- Scope, permitted use and de-identification requirements agreed before any work begins
Every dataset has a documented owner and confirmed licensing rights. See data governance on sourcex.si.
What makes a strong package
- Several years of history with consistent tagging
- Resolution codes or CSAT linked to each ticket
- Internal notes kept, since they show reasoning customers never see
- Macros and help-center articles from the same period
Compared with public datasets
Public sets such as ABCD and MultiWOZ are useful references, but limited as enterprise training data. The customer support category page compares them with licensed data.
Who typically holds it
- Software and SaaS companies
- E-commerce brands
- Insurers
- Fintech companies
- Healthcare administration teams
Know a company like this?
Introduce the company to SourceX. If its data deal closes, you can earn up to $100,000 in referral fees, paid after the buyer accepts the data and SourceX receives payment.
Refer a companyQuestions
How many tickets does a typical archive hold?
It depends on the company’s size and history. SourceX looks for companies with 20 or more full-time employees and several years of operations, and support archives commonly span 2–10 years. Put your minimum volume in the request.
Are internal notes included?
They can be. Internal notes often hold the most useful reasoning, and they go through the same de-identification steps agreed for the package.
Need this data for a model?
Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.