GitLab export project with history: what merge requests, issues and CI records exist
A GitLab project export can package issues, merge requests and related data, but coverage varies by route, version and artifact expiry. Merge requests with review threads, linked issues and pipeline outcomes are the engineering task records AI buyers ask about. Client-owned code must be excluded, and SourceX handles rights review.
What does a GitLab project export contain, and why do AI buyers care?
A GitLab instance holds the full engineering task record: merge requests with their review threads, the issues they close, commit history and the CI pipeline results that said whether a change was acceptable. Buyers building coding agents want exactly that chain, because it shows a task, the attempts to solve it, the feedback and the verdict.
For a fractional CTO, GitLab is often the system you already know best from the inside. This page covers what a project export does and does not carry, how self-managed archives differ from hosted projects, and how to flag a company without touching its code.
What engineering records live in GitLab
| Record | What it shows | Why it matters for agent training and evaluation |
|---|---|---|
| Merge requests | Description, diff context, reviewer comments, approvals, revisions | Human feedback on real changes, with the final decision |
| Issues and boards | Problem statement, labels, milestones, assignees, state changes | The task as the business described it |
| Linked issue and merge request pairs | Which change closed which problem | Task-to-solution pairs |
| Pipelines and jobs | Build, test and deploy outcomes per commit | Pass or fail labels on each attempt |
| Wikis and snippets | Runbooks, design notes, how-tos | Context that explains decisions |
| Groups and permissions | Team structure across projects | Helps scope what belongs to the company |
The valuable part is the linkage. A repository alone is source code. Code plus the threads and outcomes around it is a record of how engineers worked.
Export routes and what each one covers
GitLab offers several routes, and their coverage differs. Treat the following as questions to check against the current GitLab documentation for the version in use, not as fixed guarantees.
| Route | General idea | Question to ask |
|---|---|---|
| Project export | A single project packaged with its issues, merge requests and related data | Does the file include review comments and attachments, and does it import cleanly? |
| Group export or direct transfer | Moves groups and projects between instances | Which relations are included for this version and plan? |
| Self-managed backup | A full instance backup taken by the administrator | Who restores it, how old is the oldest backup and is it tested? |
| API extraction | Programmatic pull of issues, merge requests and pipelines | Is anyone able to run it, and is there a rate or retention limit? |
Pipeline logs and job artifacts are the part most often shorter than the rest, because many teams set expiry on artifacts and logs to save storage. Ask how long they are kept before assuming a pipeline history exists.
Self-managed instances versus hosted projects
Companies that run GitLab on their own servers often have the longest histories, because nothing forces an upgrade or a cleanup. They also have the risk that a server is retired without a full backup being kept.
- A self-managed instance with a recent, tested backup is a strong signal. Ask the infrastructure owner, not the developers, who holds it.
- Hosted projects depend on the plan and settings the company chose, including artifact expiry.
- A company that moved from one host to another may have older history left behind. The Azure DevOps work item histories page covers the same pattern for Microsoft's platform.
Which code and records must be excluded?
This is where engineering records most often go wrong, and where your judgment as a technical adviser helps most.
- Client-owned code. A software agency or consultancy building for customers usually does not own the repositories. Without client consent, those projects are out of scope.
- Open source contributions. Upstream code carries its own license terms, and the company may not be free to include it as its own.
- Third-party and vendored code. Libraries copied into a repository belong to their authors.
- Secrets and credentials. Keys, tokens and passwords committed by mistake must be dealt with before anything leaves the company.
- Personal data in issues. Support-linked issues can include customer names and contact details.
The rights review and any redaction rules are agreed between SourceX and the company before work begins. Your role is only to spot whether these exclusions will be large.
The five-question engineering depth check
Ask the CTO or head of engineering, in a short call:
- How many years of merge requests exist in the instance you use today?
- Do reviews happen inside merge requests, or in chat and pull-request-free workflows?
- Are issues linked to the merge requests that closed them?
- How long are pipeline logs and artifacts kept?
- Who could run a full export, and has anyone tested it?
Answers of several years, reviews in merge requests, linked issues and a named exporter point toward a strong record. The PagerDuty incident timelines page explains what an operations layer adds on top of this.
What to say to a company CTO
For the broader role, read the fractional CTO overview and the Procore project data page for a non-software example of the same logic.
How the introduction and rewards work
You introduce the company and share basic fit information, nothing more. SourceX qualifies the company, the company completes its inventory, price and terms are agreed, buyers review and the deal closes before delivery. The data inventory builder helps list systems without describing content, and the who qualifies page sets the baseline of 50+ full-time employees at peak (contractors excluded).
Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. Payment happens only after the buyer pays and SourceX receives its fee, and no reward is guaranteed. If you advise the company under a services contract, check that contract and your own policies on referral fees first.
When GitLab is not the signal
Pass when the instance mostly hosts client repositories, when the history is under a year, when the company is below the size baseline, or when the owner will not consider an exclusive license.
Next step
If a company you advise has years of reviewed merge requests, register as a partner and introduce the sponsor, or send them to sourcex.si/apply with your referral link.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
Does a GitLab project export include merge request comments?
Project exports are designed to carry issues and merge requests with related data, but what is included depends on the GitLab version and settings. Run a trial export on a small project and check that review comments, labels and attachments appear before you rely on it for a larger archive.
Is a GitLab group export enough for a full archive?
Not always. Group export or direct transfer covers groups and their projects, but pipeline logs, artifacts and instance-level settings may sit outside it. For a self-managed instance, a tested administrator backup is a separate safeguard worth confirming with the person who runs the servers.
How do I keep history when migrating from GitLab to GitHub?
Migration tools vary in what they preserve, and issues, review comments and pipeline results often map imperfectly. Before cutting over, keep a native GitLab export or backup of the old instance, so history that does not translate is still available later.
Can an agency license its GitLab history?
Only for projects it owns. Repositories built for clients usually belong to those clients or are covered by contract terms, so they stay out without consent. Internal tooling, the agency's own products and its internal issue tracking may qualify if the company holds the rights.
Does the partner need to look at any code?
No. A partner never exports, uploads or describes confidential records. You only note basic fit facts, such as how long the team has used GitLab and who owns exports, and the company works with SourceX on inventory and rights.
Related pages
- Azure DevOps export of work items, repos and pipeline history: what a deep record looks like
- PagerDuty export incidents: what timelines and postmortems exist and how deep they go
- Referral opportunities for fractional CTOs
- Procore export of project data: RFIs, submittals and what survives closeout
- Build a metadata-only business data inventory
- Which US businesses are a fit for a SourceX data licensing introduction
Free resources
- MCP ROI calculator — Estimate hours saved, implied savings and first-year ROI from MCP.
- Business exit readiness assessment — A preliminary exit readiness score and checklist for advisors.
- SDE vs EBITDA calculator — Seller's discretionary earnings next to market-rate EBITDA.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment