How to scan git history for secrets before licensing source code
To scan git history for secrets, mirror every repository, run two scanners across all branches, tags and pull requests, triage each finding, rotate every live credential, then exclude configuration and infrastructure files and re-scan the final set. Deleting a key from the latest commit does not remove it from history.
How do you scan git history for secrets before licensing source code?
Scan every commit on every branch and tag, not just the current files, then rotate anything that turns up, then decide what to exclude. A secret deleted in a later commit is still sitting in the history, so a clean working tree proves nothing. The sequence that works is: inventory the repositories, run a full-history scanner, triage the findings, revoke and reissue the credentials, and only then agree what goes into a licensed set.
This is a normal gate before engineering records leave a company, whatever the destination. Partners never touch repositories. The company's engineering lead runs the work, and SourceX agrees scope and redaction with the company before any work begins.
Why does history matter more than the latest commit?
Git keeps every version of every file. A developer who committed an API key on a Friday and removed it on Monday has left the key readable in the commit from Friday, in any clone, fork and backup made since. Pull requests, review comments, issue threads and CI logs can carry the same material. Licensing code, pull requests and ticket history together widens the surface, because buyers want the discussion around the change, not just the diff.
What do you need before you start?
- A list of every repository, including archived, forked and mirrored ones, with an owner for each.
- Admin access to the hosting platform so private, deleted-branch and pull-request refs can be reached.
- A secrets register: the cloud accounts, payment processors, email services, databases and third-party APIs the company uses, each with someone who can rotate credentials.
- A decision on scope: which repositories are candidates for the license and which are excluded outright.
- A place to record findings that is not a chat thread, because the findings themselves are sensitive.
Step by step: scan, rotate, exclude
- Inventory. Enumerate repositories and mark each as in scope, out of scope or unknown. Unknown means out until someone claims it.
- Mirror, do not trust the working copy. Work from a full mirror clone so all refs, branches and tags are included.
- Run two different scanners. Open-source tools such as Gitleaks and TruffleHog can scan history; check each tool's documentation for what it covers. Running more than one and merging findings reduces blind spots. Configure each with your own custom patterns for internal token formats.
- Include more than code. Scan wikis, issue and pull request text, build logs and exported chat attached to engineering work.
- Triage every hit. Mark each finding live, expired, test or false positive. Verify liveness safely by checking with the credential's owner rather than by trying the key yourself in production.
- Rotate first, clean later. Revoke and reissue every live credential, update the systems that use it, and record the date. Rewriting history is optional and secondary, because every clone, fork and backup would also need handling; a rotated key in history is a dead key.
- Exclude by file type. Leave out environment files, configuration and infrastructure-as-code, certificates and private keys, database dumps, deployment manifests and anything under a vendor's confidential terms.
- Re-scan the final set. Scan the exact files that will be delivered, after exclusions, and keep the report.
- Freeze and document. Record the tool versions, rules, scope and sign-off so the process can be repeated for later deliveries.
What should be excluded even when no secret is found?
| Category | Examples | Why leave it out |
|---|---|---|
| Credentials and keys | API tokens, private keys, certificates, password files | Direct access risk |
| Configuration | Environment files, secrets-manager exports, cloud settings | Map of internal systems |
| Infrastructure code | Terraform-style files, deployment scripts, network diagrams | Reveals architecture and accounts |
| Customer data in fixtures | Test files built from real customer records | Personal and confidential data |
| Third-party code under restrictive terms | Vendored libraries, licensed components, customer-specific code | The company may not hold the rights |
| Security reports | Penetration test results, vulnerability notes | Sensitive by nature |
| Personal files | Developers' dotfiles and local notes in repositories | Not company work product; see personal items in work systems |
Common mistakes
| Mistake | Why it hurts | Fix |
|---|---|---|
| Scanning only the default branch | Old branches and tags still hold secrets | Scan a full mirror with all refs |
| Deleting the file and calling it done | Commit history still contains it | Rotate the credential |
| Rewriting history before rotating | The old key remains valid and may exist in forks | Rotate first |
| Trusting one scanner | Each misses patterns | Run two and add custom rules |
| Ignoring false-positive review | Real findings get dismissed in the noise | Triage each hit with its owner |
| Forgetting archived repositories | Oldest code often has the worst habits | Include them in the inventory |
Illustrative example
Illustrative and fictional. A 90-person infrastructure software company wants to license six years of repositories, pull requests and tracker history. Its engineering lead mirrors 140 repositories, scans them with two tools, and finds a payment-processor key in a commit from three years ago, plus an old database password in a test fixture. Both are rotated within the week, the fixture directory is excluded, and the final delivery set is re-scanned and signed off by the lead and the CTO before the owner approves.
Why do buyers care about the scan?
Engineering records pair a change with the reasoning and the review, which is why buyers value them. They also carry a high leak risk, so the scan is a standard gate before delivery. Contract terms then limit how the code may be used, covered in field-of-use restrictions. Buyers also care that what they train on is clean, a point developed in FTC algorithmic disgorgement and licensed data.
What partners can say
You do not review code. You can ask whether the company has done a history scan, and that question alone shows the owner you understand the gate.
Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company; the reward is paid only after the buyer pays and SourceX receives its fee, and no reward is guaranteed. Check your own professional rules on referral fees and disclosure first.
When to wait
Pause if nobody owns the repositories, the code is mostly a customer's or outsourcer's work without consent, or no one can export the history. Companies can become ready later once ownership and exports are sorted.
Next step
If a company has 50+ full-time employees at peak (contractors excluded), years of engineering history and an engineering lead who can run the scan, register as a partner and make the introduction. The owner can start with the company fit checker, and the stages are laid out in how SourceX referrals work.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
Is deleting a secret from the latest commit enough?
No. Git retains earlier commits, so the secret stays readable in history, forks and backups. Treat any credential that ever appeared in a repository as exposed, rotate it, and then decide whether to clean history. Rotation is the control that removes the risk.
Should I use Gitleaks or TruffleHog?
Either can scan full history, and they differ in rules and coverage, so many teams run both and merge the findings. Add custom patterns for your own token formats. The best choice depends on your repositories and workflow, and your engineering lead should test them on a sample.
Which files should never go into a licensed code set?
Credentials and keys, environment and configuration files, infrastructure code, deployment manifests, database dumps, security reports, test fixtures built from real customer data and third-party code under restrictive terms. The company and SourceX agree the final exclusion list before delivery.
Do pull requests and tickets need scanning too?
Yes. Review comments, issue threads, wikis and CI logs often contain pasted tokens, customer details and internal URLs. If they will be part of a licensed set, scan and redact them with the same care as code before the company approves delivery.
Does a referral partner need technical skills?
No. The partner introduces the company and gives basic fit information. The company's engineering lead runs scanning and rotation, and SourceX agrees scope and redaction with the company. Partners never export, upload or describe confidential records.
Related pages
Free resources
- Business valuation calculator — Enterprise and equity value from EBITDA, your multiple, cash and debt.
- Portfolio data opportunity scanner — Screen several companies in one session.
- Working capital calculator — Net working capital, current ratio and quick ratio.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment