How are AI coding agents trained, and why do private engineering histories matter?
AI coding agents are trained in stages: pretraining on large volumes of code and text, fine-tuning on examples of requests solved well, and reinforcement learning in sandboxes where the agent edits a repository and automated tests score the result. The scarcest ingredient is private engineering history: tickets linked to pull requests, review threads and CI outcomes.
How are AI coding agents trained, stage by stage?
Most coding agents go through several rounds of training, and each round uses a different kind of data. By the end, a typical agent can be handed a real bug report and a snapshot of a repository, read the code, run commands, write a patch, and be scored on whether the project's own test suite passes and whether a reviewer would accept the change.
| Stage | What the model learns | Typical data | What is scarce |
|---|---|---|---|
| Pretraining | Syntax, libraries, common patterns | Large volumes of public code, documentation and technical Q&A | Little; public code is plentiful |
| Supervised fine-tuning | Turning a request into a correct change | Pairs of instructions and accepted solutions | Realistic requests from commercial teams |
| Reinforcement learning | Acting step by step inside a repository | Tasks in sandboxed environments with automatic checks such as tests | Tasks with trustworthy pass or fail signals |
| Evaluation | Whether the agent really improved | Held-out tasks the model has never seen | Private tasks that cannot have leaked into training |
Developers combine these stages in different ways and rarely publish full recipes, so read this as the common shape rather than any one developer's method.
Why is public code not enough?
Public repositories show what code looks like now. They rarely show why it changed. In a commercial engineering team, a single change usually sits inside a chain: a customer complaint becomes a ticket, the ticket gets acceptance criteria, a developer opens a pull request, a reviewer asks for changes, CI fails on an edge case, the fix is merged, and occasionally it is rolled back after an incident.
That chain is the closest thing software engineering has to a recorded lesson, and most of it lives in private systems. Open-source projects also work differently from companies, with fewer customer deadlines, less legacy code, fewer compliance constraints and different review habits. The wider gap is covered in which kinds of work are missing from AI training data. Simulated repositories fill part of that gap, and the trade-offs between synthetic environments and real engineering logs are worth knowing before a client conversation.
Which engineering records are most useful?
The value sits in records that link together, so a task, its solution and its outcome can be followed end to end.
| Record | Where it usually lives | Why it helps |
|---|---|---|
| Issues and tickets with acceptance criteria | Jira, Linear, Azure DevOps, GitHub Issues | Defines the task in the requester's words |
| Pull or merge requests with diffs | GitHub, GitLab, Bitbucket | Shows the accepted solution |
| Review threads and requested changes | Pull request comments | Captures expert judgment on quality |
| CI and test results | GitHub Actions, GitLab CI, Jenkins | Gives a verifiable pass or fail signal |
| Incident reviews and rollbacks | Postmortem documents, on-call tools | Records failures and how they were fixed |
| Design docs and decision records | Confluence, Notion, shared drives | Explains why one approach won over another |
The linked-history test for CTO and MSP partners
Fractional CTOs and managed service providers see engineering habits from the inside, which makes them good judges of whether a client's history deserves a conversation. Five questions do most of the work:
- Do commit messages or branch names reference ticket IDs, so each change traces back to a request?
- Is review required before merge, and are the review comments kept?
- Has CI history been retained rather than purged after a few weeks?
- Did past migrations, for example from Bitbucket to GitHub, carry full history instead of a fresh start?
- Is the client a US business with 50+ full-time employees at peak (contractors excluded), several years of engineering history and an executive who can authorize a license?
B2B software companies, IT services firms and product teams inside larger operating businesses tend to score well. More role-specific ideas are on the fractional CTO partner page.
Which rights questions come first?
Code is licensable only if the company holds the rights. Under US copyright law, work an employee creates within the scope of employment is generally a work made for hire owned by the employer, while work from independent contractors may not be unless a qualifying signed agreement exists, as the Copyright Office's circular on works made for hire explains. Three situations deserve an early flag:
- Software agencies and dev shops often build code that their contracts assign to clients, so the agency cannot license it without consent.
- Repositories contain third-party open-source components that carry their own license terms.
- Secrets, credentials and customer data sometimes sit in code, configuration or test fixtures and must be removed before any delivery.
Licensed data is also hard to pull back out of a trained model, as explained in whether licensed data can be removed from a model, so scope decisions belong at the start. This is general information, not legal, tax or financial advice. Confirm with your own counsel before acting.
What does a CTO or MSP partner actually do?
Your part is the introduction. You do not clone repositories, take screenshots or describe code to anyone.
- Raise the idea with the CEO or CTO and ask whether they want a screen.
- Register, then share your referral link or submit the company through the referral form.
- SourceX checks size, history, breadth of systems and rights with the company's sponsor.
- The company lists its own repositories, trackers and years of history in a data inventory, and agrees redaction rules before any work starts.
- Price and terms are agreed, buyers review, and if the company signs, the agreed scope is delivered and the company is paid.
A line that fits a quarterly technology review:
Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company, and the reward becomes payable only after the buyer pays and SourceX receives its fee. It never reduces what the client receives. If you advise the client under a contract, check your engagement terms and tell the client about the referral relationship.
Next step
Pick one client with a long, reviewed engineering history and run it through the company fit checker. If it passes, register as a partner and make the introduction, or read how it works first.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
Does using an AI coding assistant mean a company's code is already being used for training?
Not necessarily. Whether a coding tool's provider may use customer code depends on the provider's terms and the settings of the plan the company bought, so the answer sits in that contract. Licensing is a separate decision: the company negotiates scope, price and terms, and no repository content is delivered until it signs an agreement and authorizes delivery.
Is a repository with squashed or rewritten history still useful?
It can be, but less so. Squashed commits and fresh-start migrations remove the step-by-step trail that shows how a change evolved. Value often survives elsewhere, because tickets, pull request discussions and CI logs can still link requests to outcomes. The company's data inventory should record which systems kept full history and for which years, so buyers know what they are reviewing.
Can a software agency license code it wrote for clients?
Usually not without the client's consent, because development contracts commonly assign the code to the client. The agency may still own internal tools, templates, its own product code and its process records, and those can be worth reviewing. Check each master services agreement and statement of work early, and treat client-owned repositories as out of scope by default.
What happens to secrets and customer data found in repositories?
They are dealt with before anything is delivered. De-identification and redaction requirements are agreed with the company at the start, and credentials, keys, personal data and client identifiers are removed or masked under those rules. A company that cannot produce a clean export today is not ruled out forever; it may need preparation work first, which the inventory step helps scope.
Does a fractional CTO need to show anyone the code to make an introduction?
No. The introduction is a conversation plus a referral link or form submission with basic fit information: company name, rough size, years of history and the systems in use. The company's own team handles the inventory, rights review and any later delivery directly with SourceX under a signed agreement. Partners never export, upload or describe confidential code or records.
Can a company license only some of its repositories?
Yes. The company decides the scope, including which repositories, trackers, years and teams are in and which are out, such as client-owned code or sensitive products. A narrower scope can make sense when parts of a codebase carry third-party obligations. Scope, exclusivity for AI training and price are settled before signing, and nothing is binding until the company agrees and signs.
Related pages
- Which kinds of work are missing from AI training data?
- Synthetic environments vs real business logs: what AI agents learn from each
- Referral opportunities for fractional CTOs
- Can licensed data be removed from a trained AI model?
- Check Company Fit for Data Licensing
- How SourceX US company data referrals work
Free resources
- Due diligence checklist generator — A tailored document request list by deal type.
- Cash flow calculator — A 12-month cash forecast with shortfalls highlighted.
- Referral earnings calculator — Hypothetical partner earnings with the per-company cap.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment