What is pretraining data, and why were private business records never part of it?

Short answer

Pretraining data is the very large body of mostly public text, code and other content a language model learns from first, before fine-tuning for tasks. Private business records were never included because they are confidential and sit behind company systems, which is why buyers license them.

What is pretraining data, and why were private business records never part of it?: overview of What is pretraining data?, What is pretraining data made of?, Why were private business records never included?, What is the difference between pretraining and the stages after it?, Is public text running out?
Covered on this page: What is pretraining data? · What is pretraining data made of? · Why were private business records never included? · What is the difference between pretraining and the stages after it? · Is public text running out?

What is pretraining data?

Pretraining data is the very large body of text, code and other content a language model learns from first, before it is adapted for specific tasks. It gives the model general knowledge of language and the world. Later stages, such as fine-tuning and evaluation, use smaller and more targeted data.

Think of pretraining as general education and the later stages as job training. This page explains the terms so you can follow what AI buyers say they are missing.

What is pretraining data made of?

Developers have described pretraining mixes that draw on broad public sources. The usual categories are below. Mixes differ by developer and are not always disclosed.

Source typeExamplesTypical limitation
Public web pagesArticles, forums, reference sitesQuality varies; little of how businesses actually operate
Books and long-form writingPublished books, papersRights questions; not records of workflows
Code repositoriesOpen-source projectsPublic code only; lacks company context
Reference and academic textEncyclopedias, journalsFacts, not decisions with outcomes
Licensed collectionsContent under agreementScope set by contract

What these have in common is that they are public or deliberately shared. Private business records, such as internal email, tickets, CRM histories and approvals, were never part of that pool.

Why were private business records never included?

Because they were never published. A company's internal records sit behind logins, in tools like email, Slack or Teams, CRM, finance and support systems. Access requires the company's permission, and the content is confidential, so it could not be gathered by crawling the public web.

That is the gap. Models can read about business in general, but they have limited examples of real work done step by step inside organizations: the request, the exceptions, the approvals and the results.

What is the difference between pretraining and the stages after it?

StagePurposeData it usesWhere company records fit
PretrainingLearn general language and knowledgeHuge, broad, mostly public corporaRarely; private records are outside it
Fine-tuningAdapt a model to a task or styleSmaller, targeted examplesExamples of tasks done well
Preference or feedback trainingShape behavior using judgmentsRanked or rated outputsExpert decisions with outcomes
EvaluationMeasure performanceHeld-out test setsCases with known results; see golden datasets

The licensing demand SourceX serves sits mostly after pretraining: material for building and testing agents that perform tasks. Where a particular buyer uses a licensed set is the buyer's decision and is set out in the agreement, not something a partner can promise.

Is public text running out?

A research paper from Epoch AI estimated the effective stock of public human-written text at roughly 300 trillion tokens and projected that, if trends continue, language models could fully use that stock between 2026 and 2032. It is a forecast with wide uncertainty, and the authors also discuss synthetic data and data efficiency as ways to cope. The takeaway for partners is limited: researchers expect public text to become a constraint, which adds interest in non-public, permissioned data.

Why does this matter for referral partners?

It explains why a US company's old systems can have value. Years of tickets, deal histories and engineering reviews describe real work in a way public pages do not. Whether any company's records suit a buyer is assessed during qualification and inventory, never by the partner. Partners make introductions and give basic fit information only.

Two related questions arise early. First, rights: buyers want rights-cleared data with a clear basis to license. Second, documentation: the comparison of data lineage vs data provenance explains how origin is recorded. Buyers may also use controlled access, covered in compute-to-data, and prepare records through data curation. For the commercial side, see what data monetization means and the guide to enterprise AI data licensing deals.

What should a partner say?

How do rewards work?

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. The reward is paid only after the buyer pays and SourceX receives its fee, and no reward is guaranteed.

Next step

If you know a company with years of records and an authorized sponsor, register as a partner and introduce it, or start with the company fit checker. The steps are in how it works.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

What is the difference between pretraining and fine-tuning?

Pretraining teaches a model general language and knowledge from very large, broad datasets. Fine-tuning adapts the pretrained model to a particular task or style using smaller, targeted examples. Fine-tuning data is usually much more specific, which is where examples of real work tasks become relevant.

Were company emails and tickets used to pretrain models?

Private business records are not part of public web collections because they sit behind company logins and are confidential. Use of any private data requires the owner's permission. Companies should not assume a past use occurred either way; any license is an explicit, signed agreement with defined scope.

Why would a buyer want business records if it already has the web?

The public web has limited examples of real work done step by step: requests, decisions, exceptions and results inside organizations. Buyers training agents to perform tasks look for those records. Each purchase depends on the data, rights and terms, and nothing is binding until the company signs.

Does licensing data to a buyer mean it becomes pretraining data?

Not necessarily. A license defines permitted use, such as AI training or evaluation, and the buyer chooses how to apply it within those limits. Companies should read the permitted-use clause carefully and ask counsel how it treats training stages and models trained during the term.

Is the Epoch AI projection a certainty?

No. It is a forecast with wide uncertainty and depends on assumptions about training trends and data use. The authors also discuss synthetic data and efficiency as ways to reduce reliance on new public text. It is context for market interest, not a prediction about any specific company.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment