What is pretraining data?
Pretraining data is the very large body of text, code and other content a language model learns from first, before it is adapted for specific tasks. It gives the model general knowledge of language and the world. Later stages, such as fine-tuning and evaluation, use smaller and more targeted data.
Think of pretraining as general education and the later stages as job training. This page explains the terms so you can follow what AI buyers say they are missing.
What is pretraining data made of?
Developers have described pretraining mixes that draw on broad public sources. The usual categories are below. Mixes differ by developer and are not always disclosed.
| Source type | Examples | Typical limitation |
|---|---|---|
| Public web pages | Articles, forums, reference sites | Quality varies; little of how businesses actually operate |
| Books and long-form writing | Published books, papers | Rights questions; not records of workflows |
| Code repositories | Open-source projects | Public code only; lacks company context |
| Reference and academic text | Encyclopedias, journals | Facts, not decisions with outcomes |
| Licensed collections | Content under agreement | Scope set by contract |
What these have in common is that they are public or deliberately shared. Private business records, such as internal email, tickets, CRM histories and approvals, were never part of that pool.
Why were private business records never included?
Because they were never published. A company's internal records sit behind logins, in tools like email, Slack or Teams, CRM, finance and support systems. Access requires the company's permission, and the content is confidential, so it could not be gathered by crawling the public web.
That is the gap. Models can read about business in general, but they have limited examples of real work done step by step inside organizations: the request, the exceptions, the approvals and the results.
What is the difference between pretraining and the stages after it?
| Stage | Purpose | Data it uses | Where company records fit |
|---|---|---|---|
| Pretraining | Learn general language and knowledge | Huge, broad, mostly public corpora | Rarely; private records are outside it |
| Fine-tuning | Adapt a model to a task or style | Smaller, targeted examples | Examples of tasks done well |
| Preference or feedback training | Shape behavior using judgments | Ranked or rated outputs | Expert decisions with outcomes |
| Evaluation | Measure performance | Held-out test sets | Cases with known results; see golden datasets |
The licensing demand SourceX serves sits mostly after pretraining: material for building and testing agents that perform tasks. Where a particular buyer uses a licensed set is the buyer's decision and is set out in the agreement, not something a partner can promise.
Is public text running out?
A research paper from Epoch AI estimated the effective stock of public human-written text at roughly 300 trillion tokens and projected that, if trends continue, language models could fully use that stock between 2026 and 2032. It is a forecast with wide uncertainty, and the authors also discuss synthetic data and data efficiency as ways to cope. The takeaway for partners is limited: researchers expect public text to become a constraint, which adds interest in non-public, permissioned data.
Why does this matter for referral partners?
It explains why a US company's old systems can have value. Years of tickets, deal histories and engineering reviews describe real work in a way public pages do not. Whether any company's records suit a buyer is assessed during qualification and inventory, never by the partner. Partners make introductions and give basic fit information only.
Two related questions arise early. First, rights: buyers want rights-cleared data with a clear basis to license. Second, documentation: the comparison of data lineage vs data provenance explains how origin is recorded. Buyers may also use controlled access, covered in compute-to-data, and prepare records through data curation. For the commercial side, see what data monetization means and the guide to enterprise AI data licensing deals.
What should a partner say?
How do rewards work?
Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. The reward is paid only after the buyer pays and SourceX receives its fee, and no reward is guaranteed.
Next step
If you know a company with years of records and an authorized sponsor, register as a partner and introduce it, or start with the company fit checker. The steps are in how it works.