Software engineering data for AI training

Software engineering data shows how professional teams write, review and ship code: private repositories with full history, pull requests with review threads, issue trackers and incident records. AI labs license it to train and evaluate coding agents on real codebases that aren’t public.

Last updated October 3, 2026

Listings

DataTypical sourcesModalityAvailability
Private codebases with full commit history
Proprietary repositories with every commit, branch and tag, plus build and test setup.
GitHub, GitLab, BitbucketCodeOn request
Pull requests with code review threads
Diffs, review comments, requested changes, approvals and CI results for each change.
GitHub, GitLab, BitbucketCode and textOn request
Issue and sprint histories
Tickets, sprints and status changes from Jira, Linear or GitHub Issues, linked to code.
Jira, Linear, GitHub IssuesText with structured metadataOn request
Incident, on-call and postmortem records
Alerts, incident timelines, responder chat and the postmortems that followed.
PagerDuty, Opsgenie, incident.ioText and logsOn request

What’s included

  • Private repositories with complete git history
  • Pull requests with diffs, review comments, approvals and CI results
  • Issue and sprint histories linked to the code that resolved them
  • Incident timelines and postmortems

Public datasets and licensed data

Public code data is open source by definition, so it misses the proprietary codebases, internal review norms and incident histories of real companies. Private repositories also reduce the benchmark contamination that comes with public GitHub, because they aren’t on the public web.

Public datasetReleasedWhat it containsLimits for enterprise useLicense
The Stack v2BigCode and Software Heritage, 20243.28 billion files from 104.2 million public GitHub repositories in 600+ languages.Public open-source code only; gated access and the original licenses still apply.Gated; original code licenses apply
SWE-benchPrinceton, 20232,294 tasks pairing real GitHub issues with the pull requests that fixed them, from 12 Python repositories.Public repositories that models may have seen during training.MIT
GH ArchiveIlya Grigorik, since 2011A record of public GitHub activity (commits, issues, comments and more), available as hourly files or in BigQuery.Public activity only; private company work never appears.Data may be subject to third-party rights
Public Jira DatasetMontgomery et al., MSR 202216 public Jira sites: 1,822 projects, 2.7 million issues and 9 million comments.Open-source projects, not private company trackers.CC BY 4.0

Preparation and rights

Preparation requirements are agreed with each supplier before any work begins. For this kind of data they typically include:

  • Confirming the company owns the code, including contractors’ work covered by IP assignments
  • Scanning for and removing secrets, keys and credentials from history
  • Identifying open-source components and listing their licenses
  • Removing customer data from test fixtures and logs

Every dataset has a documented owner and confirmed licensing rights, and is delivered only after an executed agreement and the supplier’s authorization. See data governance on sourcex.si.

Use cases

More on sourcex.si

Questions

Why do AI labs want private code?

Private code isn’t on the public web, so it adds examples public datasets lack and supports evaluations models are unlikely to have memorized. It also reflects how companies actually build software, with internal reviews and real incidents.

Which code hosts does SourceX source from?

GitHub, GitLab, Bitbucket and Azure DevOps for code and reviews, and Jira, Linear and GitHub Issues for issue histories.

Need software engineering data?

Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.