How are AI coding agents trained, and why do private engineering histories matter?

AI coding agents are trained in stages: pretraining on large volumes of code and text, fine-tuning on examples of requests solved well, and reinforcement learning in sandboxes where the agent edits a repository and automated tests score the result. The scarcest ingredient is private engineering history: tickets linked to pull requests, review threads and CI outcomes.

How are AI coding agents trained, stage by stage?

Most coding agents go through several rounds of training, and each round uses a different kind of data. By the end, a typical agent can be handed a real bug report and a snapshot of a repository, read the code, run commands, write a patch, and be scored on whether the project's own test suite passes and whether a reviewer would accept the change.

StageWhat the model learnsTypical dataWhat is scarce
PretrainingSyntax, libraries, common patternsLarge volumes of public code, documentation and technical Q&ALittle; public code is plentiful
Supervised fine-tuningTurning a request into a correct changePairs of instructions and accepted solutionsRealistic requests from commercial teams
Reinforcement learningActing step by step inside a repositoryTasks in sandboxed environments with automatic checks such as testsTasks with trustworthy pass or fail signals
EvaluationWhether the agent really improvedHeld-out tasks the model has never seenPrivate tasks that cannot have leaked into training

Developers combine these stages in different ways and rarely publish full recipes, so read this as the common shape rather than any one developer's method.

Why is public code not enough?

Public repositories show what code looks like now. They rarely show why it changed. In a commercial engineering team, a single change usually sits inside a chain: a customer complaint becomes a ticket, the ticket gets acceptance criteria, a developer opens a pull request, a reviewer asks for changes, CI fails on an edge case, the fix is merged, and occasionally it is rolled back after an incident.

That chain is the closest thing software engineering has to a recorded lesson, and most of it lives in private systems. Open-source projects also work differently from companies, with fewer customer deadlines, less legacy code, fewer compliance constraints and different review habits. The wider gap is covered in which kinds of work are missing from AI training data. Simulated repositories fill part of that gap, and the trade-offs between synthetic environments and real engineering logs are worth knowing before a client conversation.

Which engineering records are most useful?

The value sits in records that link together, so a task, its solution and its outcome can be followed end to end.

RecordWhere it usually livesWhy it helps
Issues and tickets with acceptance criteriaJira, Linear, Azure DevOps, GitHub IssuesDefines the task in the requester's words
Pull or merge requests with diffsGitHub, GitLab, BitbucketShows the accepted solution
Review threads and requested changesPull request commentsCaptures expert judgment on quality
CI and test resultsGitHub Actions, GitLab CI, JenkinsGives a verifiable pass or fail signal
Incident reviews and rollbacksPostmortem documents, on-call toolsRecords failures and how they were fixed
Design docs and decision recordsConfluence, Notion, shared drivesExplains why one approach won over another

The linked-history test for CTO and MSP partners

Fractional CTOs and managed service providers see engineering habits from the inside, which makes them good judges of whether a client's history deserves a conversation. Five questions do most of the work:

  • Do commit messages or branch names reference ticket IDs, so each change traces back to a request?
  • Is review required before merge, and are the review comments kept?
  • Has CI history been retained rather than purged after a few weeks?
  • Did past migrations, for example from Bitbucket to GitHub, carry full history instead of a fresh start?
  • Is the client a US business with 50+ full-time employees at peak (contractors excluded), several years of engineering history and an executive who can authorize a license?

B2B software companies, IT services firms and product teams inside larger operating businesses tend to score well. More role-specific ideas are on the fractional CTO partner page.

Which rights questions come first?

Code is licensable only if the company holds the rights. Under US copyright law, work an employee creates within the scope of employment is generally a work made for hire owned by the employer, while work from independent contractors may not be unless a qualifying signed agreement exists, as the Copyright Office's circular on works made for hire explains. Three situations deserve an early flag:

  • Software agencies and dev shops often build code that their contracts assign to clients, so the agency cannot license it without consent.
  • Repositories contain third-party open-source components that carry their own license terms.
  • Secrets, credentials and customer data sometimes sit in code, configuration or test fixtures and must be removed before any delivery.

Licensed data is also hard to pull back out of a trained model, as explained in whether licensed data can be removed from a model, so scope decisions belong at the start. This is general information, not legal, tax or financial advice. Confirm with your own counsel before acting.

What does a CTO or MSP partner actually do?

Your part is the introduction. You do not clone repositories, take screenshots or describe code to anyone.

  1. Raise the idea with the CEO or CTO and ask whether they want a screen.
  2. Register, then share your referral link or submit the company through the referral form.
  3. SourceX checks size, history, breadth of systems and rights with the company's sponsor.
  4. The company lists its own repositories, trackers and years of history in a data inventory, and agrees redaction rules before any work starts.
  5. Price and terms are agreed, buyers review, and if the company signs, the agreed scope is delivered and the company is paid.

A line that fits a quarterly technology review:

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company, and the reward becomes payable only after the buyer pays and SourceX receives its fee. It never reduces what the client receives. If you advise the client under a contract, check your engagement terms and tell the client about the referral relationship.

Next step

Pick one client with a long, reviewed engineering history and run it through the company fit checker. If it passes, register as a partner and make the introduction, or read how it works first.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Does using an AI coding assistant mean a company's code is already being used for training?

Not necessarily. Whether a coding tool's provider may use customer code depends on the provider's terms and the settings of the plan the company bought, so the answer sits in that contract. Licensing is a separate decision: the company negotiates scope, price and terms, and no repository content is delivered until it signs an agreement and authorizes delivery.

Is a repository with squashed or rewritten history still useful?

It can be, but less so. Squashed commits and fresh-start migrations remove the step-by-step trail that shows how a change evolved. Value often survives elsewhere, because tickets, pull request discussions and CI logs can still link requests to outcomes. The company's data inventory should record which systems kept full history and for which years, so buyers know what they are reviewing.

Can a software agency license code it wrote for clients?

Usually not without the client's consent, because development contracts commonly assign the code to the client. The agency may still own internal tools, templates, its own product code and its process records, and those can be worth reviewing. Check each master services agreement and statement of work early, and treat client-owned repositories as out of scope by default.

What happens to secrets and customer data found in repositories?

They are dealt with before anything is delivered. De-identification and redaction requirements are agreed with the company at the start, and credentials, keys, personal data and client identifiers are removed or masked under those rules. A company that cannot produce a clean export today is not ruled out forever; it may need preparation work first, which the inventory step helps scope.

Does a fractional CTO need to show anyone the code to make an introduction?

No. The introduction is a conversation plus a referral link or form submission with basic fit information: company name, rough size, years of history and the systems in use. The company's own team handles the inventory, rights review and any later delivery directly with SourceX under a signed agreement. Partners never export, upload or describe confidential code or records.

Can a company license only some of its repositories?

Yes. The company decides the scope, including which repositories, trackers, years and teams are in and which are out, such as client-owned code or sensitive products. A narrower scope can make sense when parts of a codebase carry third-party obligations. Scope, exclusivity for AI training and price are settled before signing, and nothing is binding until the company agrees and signs.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment