Private codebases with full commit history
Proprietary source code from operating or wound-down companies, with the full git history: every commit, branch and tag, plus build and test configuration. Labs license it to train coding models on production code that isn’t publicly available, and to build evaluation tasks that are unlikely to appear in training data.
What a record contains
One repository: a git mirror with complete history, plus metadata on languages, size, contributors (pseudonymous) and CI configuration.
| Field | What it holds |
|---|---|
repo_id | Pseudonymous ID |
languages{} | Lines of code by language |
commits | Count and date range |
ci_config | Build and test setup |
license_scan | Open-source components and their licenses |
secrets_scan | Report of credentials found and removed |
{
"repo_id": "R-208",
"languages": {"Go": 182400, "TypeScript": 96100, "SQL": 8700},
"commits": {"count": 21544, "first": "2017-05-02", "last": "2025-08-29"},
"ci_config": ".github/workflows",
"license_scan": {"components": 214, "copyleft": 3},
"secrets_scan": {"found": 41, "removed": 41}
}How AI labs use it
- Pretraining and code completion
- Production code from domains public GitHub covers thinly.
- Repository-level tasks
- Multi-file changes, refactors and migrations with real history.
- Held-out evaluation
- Code that isn’t on the public web, so it’s unlikely to be in training data.
Typical preparation requirements
Agreed with the supplier before any work begins. Typical requirements include:
- The company confirms it owns the code, including contractor work covered by IP assignments
- Secrets, keys and credentials scanned for and removed from history
- Open-source components identified and listed with their licenses
- Customer data in test fixtures removed
Every dataset has a documented owner and confirmed licensing rights. See data governance on sourcex.si.
What makes a strong package
- Years of history with meaningful commit messages
- Tests and CI that still run
- Linked issues and pull requests
- Less common languages or domains, such as embedded, scientific or financial code
Compared with public datasets
Public sets such as The Stack v2 and GH Archive are useful references, but limited as enterprise training data. The software engineering category page compares them with licensed data.
Who typically holds it
- SaaS companies
- Fintech companies
- Embedded and hardware companies
- Agencies that own their product code
- Shut-down startups
Know a company like this?
Introduce the company to SourceX. If its data deal closes, you can earn up to $100,000 in referral fees, paid after the buyer accepts the data and SourceX receives payment.
Refer a companyQuestions
Does licensing code let the buyer run the product?
No. The license covers the permitted use agreed, usually training and evaluation, not operating the product. Scope is set in the agreement.
What about open-source code inside the repositories?
It’s identified and listed with its license, and it stays under its original license terms.
Need this data for a model?
Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.