Software engineering data for AI training
Software engineering data shows how professional teams write, review and ship code: private repositories with full history, pull requests with review threads, issue trackers and incident records. AI labs license it to train and evaluate coding agents on real codebases that aren’t public.
Listings
| Data | Typical sources | Modality | Availability |
|---|---|---|---|
| Private codebases with full commit history | GitHub, GitLab, Bitbucket | Code | On request |
| Pull requests with code review threads | GitHub, GitLab, Bitbucket | Code and text | On request |
| Issue and sprint histories | Jira, Linear, GitHub Issues | Text with structured metadata | On request |
| Incident, on-call and postmortem records | PagerDuty, Opsgenie, incident.io | Text and logs | On request |
What’s included
- Private repositories with complete git history
- Pull requests with diffs, review comments, approvals and CI results
- Issue and sprint histories linked to the code that resolved them
- Incident timelines and postmortems
Public datasets and licensed data
Public code data is open source by definition, so it misses the proprietary codebases, internal review norms and incident histories of real companies. Private repositories also reduce the benchmark contamination that comes with public GitHub, because they aren’t on the public web.
| Public dataset | Released | What it contains | Limits for enterprise use | License |
|---|---|---|---|---|
| The Stack v2 | BigCode and Software Heritage, 2024 | 3.28 billion files from 104.2 million public GitHub repositories in 600+ languages. | Public open-source code only; gated access and the original licenses still apply. | Gated; original code licenses apply |
| SWE-bench | Princeton, 2023 | 2,294 tasks pairing real GitHub issues with the pull requests that fixed them, from 12 Python repositories. | Public repositories that models may have seen during training. | MIT |
| GH Archive | Ilya Grigorik, since 2011 | A record of public GitHub activity (commits, issues, comments and more), available as hourly files or in BigQuery. | Public activity only; private company work never appears. | Data may be subject to third-party rights |
| Public Jira Dataset | Montgomery et al., MSR 2022 | 16 public Jira sites: 1,822 projects, 2.7 million issues and 9 million comments. | Open-source projects, not private company trackers. | CC BY 4.0 |
Preparation and rights
Preparation requirements are agreed with each supplier before any work begins. For this kind of data they typically include:
- Confirming the company owns the code, including contractors’ work covered by IP assignments
- Scanning for and removing secrets, keys and credentials from history
- Identifying open-source components and listing their licenses
- Removing customer data from test fixtures and logs
Every dataset has a documented owner and confirmed licensing rights, and is delivered only after an executed agreement and the supplier’s authorization. See data governance on sourcex.si.
Use cases
More on sourcex.si
Questions
Why do AI labs want private code?
Private code isn’t on the public web, so it adds examples public datasets lack and supports evaluations models are unlikely to have memorized. It also reflects how companies actually build software, with internal reviews and real incidents.
Which code hosts does SourceX source from?
GitHub, GitLab, Bitbucket and Azure DevOps for code and reviews, and Jira, Linear and GitHub Issues for issue histories.
Need software engineering data?
Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.