How to scan git history for secrets before licensing source code

To scan git history for secrets, mirror every repository, run two scanners across all branches, tags and pull requests, triage each finding, rotate every live credential, then exclude configuration and infrastructure files and re-scan the final set. Deleting a key from the latest commit does not remove it from history.

How do you scan git history for secrets before licensing source code?

Scan every commit on every branch and tag, not just the current files, then rotate anything that turns up, then decide what to exclude. A secret deleted in a later commit is still sitting in the history, so a clean working tree proves nothing. The sequence that works is: inventory the repositories, run a full-history scanner, triage the findings, revoke and reissue the credentials, and only then agree what goes into a licensed set.

This is a normal gate before engineering records leave a company, whatever the destination. Partners never touch repositories. The company's engineering lead runs the work, and SourceX agrees scope and redaction with the company before any work begins.

Why does history matter more than the latest commit?

Git keeps every version of every file. A developer who committed an API key on a Friday and removed it on Monday has left the key readable in the commit from Friday, in any clone, fork and backup made since. Pull requests, review comments, issue threads and CI logs can carry the same material. Licensing code, pull requests and ticket history together widens the surface, because buyers want the discussion around the change, not just the diff.

What do you need before you start?

  • A list of every repository, including archived, forked and mirrored ones, with an owner for each.
  • Admin access to the hosting platform so private, deleted-branch and pull-request refs can be reached.
  • A secrets register: the cloud accounts, payment processors, email services, databases and third-party APIs the company uses, each with someone who can rotate credentials.
  • A decision on scope: which repositories are candidates for the license and which are excluded outright.
  • A place to record findings that is not a chat thread, because the findings themselves are sensitive.

Step by step: scan, rotate, exclude

  1. Inventory. Enumerate repositories and mark each as in scope, out of scope or unknown. Unknown means out until someone claims it.
  2. Mirror, do not trust the working copy. Work from a full mirror clone so all refs, branches and tags are included.
  3. Run two different scanners. Open-source tools such as Gitleaks and TruffleHog can scan history; check each tool's documentation for what it covers. Running more than one and merging findings reduces blind spots. Configure each with your own custom patterns for internal token formats.
  4. Include more than code. Scan wikis, issue and pull request text, build logs and exported chat attached to engineering work.
  5. Triage every hit. Mark each finding live, expired, test or false positive. Verify liveness safely by checking with the credential's owner rather than by trying the key yourself in production.
  6. Rotate first, clean later. Revoke and reissue every live credential, update the systems that use it, and record the date. Rewriting history is optional and secondary, because every clone, fork and backup would also need handling; a rotated key in history is a dead key.
  7. Exclude by file type. Leave out environment files, configuration and infrastructure-as-code, certificates and private keys, database dumps, deployment manifests and anything under a vendor's confidential terms.
  8. Re-scan the final set. Scan the exact files that will be delivered, after exclusions, and keep the report.
  9. Freeze and document. Record the tool versions, rules, scope and sign-off so the process can be repeated for later deliveries.

What should be excluded even when no secret is found?

CategoryExamplesWhy leave it out
Credentials and keysAPI tokens, private keys, certificates, password filesDirect access risk
ConfigurationEnvironment files, secrets-manager exports, cloud settingsMap of internal systems
Infrastructure codeTerraform-style files, deployment scripts, network diagramsReveals architecture and accounts
Customer data in fixturesTest files built from real customer recordsPersonal and confidential data
Third-party code under restrictive termsVendored libraries, licensed components, customer-specific codeThe company may not hold the rights
Security reportsPenetration test results, vulnerability notesSensitive by nature
Personal filesDevelopers' dotfiles and local notes in repositoriesNot company work product; see personal items in work systems

Common mistakes

MistakeWhy it hurtsFix
Scanning only the default branchOld branches and tags still hold secretsScan a full mirror with all refs
Deleting the file and calling it doneCommit history still contains itRotate the credential
Rewriting history before rotatingThe old key remains valid and may exist in forksRotate first
Trusting one scannerEach misses patternsRun two and add custom rules
Ignoring false-positive reviewReal findings get dismissed in the noiseTriage each hit with its owner
Forgetting archived repositoriesOldest code often has the worst habitsInclude them in the inventory

Illustrative example

Illustrative and fictional. A 90-person infrastructure software company wants to license six years of repositories, pull requests and tracker history. Its engineering lead mirrors 140 repositories, scans them with two tools, and finds a payment-processor key in a commit from three years ago, plus an old database password in a test fixture. Both are rotated within the week, the fixture directory is excluded, and the final delivery set is re-scanned and signed off by the lead and the CTO before the owner approves.

Why do buyers care about the scan?

Engineering records pair a change with the reasoning and the review, which is why buyers value them. They also carry a high leak risk, so the scan is a standard gate before delivery. Contract terms then limit how the code may be used, covered in field-of-use restrictions. Buyers also care that what they train on is clean, a point developed in FTC algorithmic disgorgement and licensed data.

What partners can say

You do not review code. You can ask whether the company has done a history scan, and that question alone shows the owner you understand the gate.

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company; the reward is paid only after the buyer pays and SourceX receives its fee, and no reward is guaranteed. Check your own professional rules on referral fees and disclosure first.

When to wait

Pause if nobody owns the repositories, the code is mostly a customer's or outsourcer's work without consent, or no one can export the history. Companies can become ready later once ownership and exports are sorted.

Next step

If a company has 50+ full-time employees at peak (contractors excluded), years of engineering history and an engineering lead who can run the scan, register as a partner and make the introduction. The owner can start with the company fit checker, and the stages are laid out in how SourceX referrals work.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Is deleting a secret from the latest commit enough?

No. Git retains earlier commits, so the secret stays readable in history, forks and backups. Treat any credential that ever appeared in a repository as exposed, rotate it, and then decide whether to clean history. Rotation is the control that removes the risk.

Should I use Gitleaks or TruffleHog?

Either can scan full history, and they differ in rules and coverage, so many teams run both and merge the findings. Add custom patterns for your own token formats. The best choice depends on your repositories and workflow, and your engineering lead should test them on a sample.

Which files should never go into a licensed code set?

Credentials and keys, environment and configuration files, infrastructure code, deployment manifests, database dumps, security reports, test fixtures built from real customer data and third-party code under restrictive terms. The company and SourceX agree the final exclusion list before delivery.

Do pull requests and tickets need scanning too?

Yes. Review comments, issue threads, wikis and CI logs often contain pasted tokens, customer details and internal URLs. If they will be part of a licensed set, scan and redact them with the same care as code before the company approves delivery.

Does a referral partner need technical skills?

No. The partner introduces the company and gives basic fit information. The company's engineering lead runs scanning and rotation, and SourceX agrees scope and redaction with the company. Partners never export, upload or describe confidential records.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment