Pull requests with code review threads
Pull requests from private repositories with the diff, every review comment, requested changes, approvals and CI results, linked to the issue behind each change. Labs use them to train code-review and coding agents on how experienced engineers critique and improve real changes.
What a record contains
One pull request: title and description, commits and diffs, inline review comments with replies, review states, CI runs and whether it merged.
| Field | What it holds |
|---|---|
pr_id, repo_id | Pseudonymous IDs |
linked_issue | The issue or ticket behind the change |
diff | Unified diff per commit |
review_comments[] | File, line, body, author, time |
reviews[] | Approved or changes requested |
ci_runs[], merged | Checks and outcome |
{
"pr_id": "PR-6120",
"linked_issue": "PAY-871",
"title": "Cap settlement batch size at 500",
"review_comments": [
{"path": "settle/batch.go", "line": 88, "author": "eng_02",
"body": "This hides the timeout instead of fixing it. Can we page the query instead?"}
],
"reviews": [{"author": "eng_02", "state": "changes_requested"}, {"author": "eng_02", "state": "approved"}],
"ci_runs": [{"status": "failed"}, {"status": "passed"}],
"merged": true
}How AI labs use it
- Code-review agents
- Real reviewer comments on real changes.
- Learning from revisions
- Before-and-after diffs show how feedback changed the code.
- Issue-to-fix tasks
- Issue and pull request pairs, like SWE-bench but from private code.
Typical preparation requirements
Agreed with the supplier before any work begins. Typical requirements include:
- The same ownership, secrets and open-source checks as the code itself
- Reviewer identities pseudonymized consistently
Every dataset has a documented owner and confirmed licensing rights. See data governance on sourcex.si.
What makes a strong package
- An active review culture, with several reviewers and substantive comments
- Pull requests linked to issues
- CI history retained
Compared with public datasets
Public sets such as SWE-bench and The Stack v2 are useful references, but limited as enterprise training data. The software engineering category page compares them with licensed data.
Who typically holds it
- Software companies
- Developer-tools companies
- Fintech companies
Know a company like this?
Introduce the company to SourceX. If its data deal closes, you can earn up to $100,000 in referral fees, paid after the buyer accepts the data and SourceX receives payment.
Refer a companyQuestions
Are review discussions part of a git clone?
No. Review threads live in the hosting platform, so they’re exported separately through its API and joined to the code.
Can you provide issue-to-PR pairs?
Yes, where the company linked issues and pull requests consistently.
Need this data for a model?
Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.