Issue and sprint histories

Engineering issue trackers with the full lifecycle of each ticket: description, comments, estimates, status changes, sprint assignments and links to the code that resolved it. Labs use them to train planning and project agents and to build tasks grounded in real requirements.

Last updated October 3, 2026

What a record contains

One issue: fields and description, comment thread, change log, sprint and epic links, and the commits or pull requests that resolved it.

FieldWhat it holds
issue_key, typeBug, story, task and so on
summary, descriptionAs written by the team
comments[]Discussion thread
changelog[]Status, assignee and estimate changes
sprint, epic, story_pointsPlanning context
linked_prs[], resolutionHow it was closed
Illustrative record. The values are made up to show the shape of the data.
{
  "issue_key": "PAY-871",
  "type": "bug",
  "summary": "Nightly settlement job times out above 2k merchants",
  "story_points": 3,
  "changelog": [
    {"at": "2024-11-02T09:10Z", "field": "status", "from": "To Do", "to": "In Progress"},
    {"at": "2024-11-04T16:30Z", "field": "status", "from": "In Review", "to": "Done"}
  ],
  "linked_prs": ["PR-6120"],
  "resolution": "fixed"
}

How AI labs use it

Planning and estimation
Estimates compared with actual cycle times.
Requirements to code
Issues paired with the changes that closed them.
Project agents
Triage, prioritize and update tickets the way real teams do.

Typical preparation requirements

Agreed with the supplier before any work begins. Typical requirements include:

  • Customer names in bug reports pseudonymized
  • Attachments screened
  • Security tickets reviewed so undisclosed vulnerabilities aren’t exposed

Every dataset has a documented owner and confirmed licensing rights. See data governance on sourcex.si.

What makes a strong package

  • Consistent workflow states
  • Links to code
  • Years of sprints from the same team

Compared with public datasets

Public sets such as Public Jira Dataset and GH Archive are useful references, but limited as enterprise training data. The software engineering category page compares them with licensed data.

Who typically holds it

  • Software companies
  • IT departments
  • Agencies

Know a company like this?

Introduce the company to SourceX. If its data deal closes, you can earn up to $100,000 in referral fees, paid after the buyer accepts the data and SourceX receives payment.

Refer a company

Questions

How is this different from the public Jira dataset?

The public dataset covers 16 open-source Jira sites. Licensed histories come from private company projects, with links to private code and the business context around each ticket.

Need this data for a model?

Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.