Where does AI training data come from? The five main sources

AI training data comes from five main sources: text and media crawled from the public web, content licensed from publishers and platforms, work contracted from people who label and write examples, synthetic data generated by models, and records licensed from operating companies. The last, permissioned records of real business work, is the hardest to obtain.

The short answer

AI training data comes from five main sources. Developers crawl the public web, license content from publishers and platforms, contract people to label, rank and write examples, generate synthetic data with existing models, and license records from operating companies. Each fills a different need, and developers often do not publish their full mix.

The fifth source is the smallest and the hardest to obtain: permissioned records of real business work, such as tickets carried to resolution, deal histories and engineering reviews. It matters more as developers build agents that perform tasks rather than only answer questions.

What are the five sources?

SourceWhat it isWhat it is good forWhere it runs short
Public web crawlPages, forums, code and documents collected by crawlersBroad language and general knowledgeHigh-quality fresh text is finite, and rights questions are contested
Licensed contentPublisher archives and platform content under agreementEdited, reliable and current material with permissionLimited to what publishers and platforms hold
Contracted human workLabels, preference rankings, expert-written examples and evaluations made to orderTeaching specific behaviors and judging qualityEach example is commissioned and written for the task, not drawn from real work
Synthetic dataExamples generated by existing modelsVolume and variety for targeted skillsNeeds real data to check against; records nothing that actually happened
Licensed enterprise recordsOperating companies' own records under a licenseReal multi-step workflows with decisions and outcomesSpread across private systems; needs rights review and redaction

For a plain definition of the term itself, see what is AI training data.

How is training data used at each stage?

The sources feed different stages of model building:

  1. Pretraining. A model learns language and general knowledge from a very large text collection, drawn largely from the web and licensed archives.
  2. Post-training. Developers refine it with curated examples, many written or rated by people, so it follows instructions and behaves as intended.
  3. Agent training. Models that take actions learn from demonstrations of multi-step tasks, which is where records of real work earn their place.
  4. Evaluation. Held-out tests check whether a model or agent can do the job; realistic tasks with known outcomes make better tests.

Why is web data the most contested source?

Crawled data is enormous in volume, but it raises copyright questions. The US Copyright Office's AI initiative released Part 3 of its report, on generative AI training, as a pre-publication version in May 2025. It examines where copying for training may implicate copyright, how fair use may apply and how practical licensing approaches are, and it notes that model performance depends heavily on data quality. It is a report, not law. For where the US court cases stand, see is AI training fair use.

This is general information, not legal, tax or financial advice.

How does licensed content reach developers?

Through agreements with the owners of archives and platforms. In December 2023, OpenAI and Axel Springer announced a partnership under which Axel Springer's content is used to advance the training of OpenAI's models and ChatGPT shows summaries of selected articles with attribution and links; financial terms were not disclosed. It is a public market event, and neither company is a SourceX buyer or partner. The supplier side is mapped in who sells AI training data, and disclosed deal values are compared in enterprise AI data licensing deals.

Where do human-made and synthetic examples fit?

Human-data work fills gaps the web cannot: experts write model answers, raters compare outputs, and annotators label examples to a developer's specification. Synthetic data fills gaps at volume, with models generating new examples for skills a developer wants to strengthen.

Both are produced for training. Neither records what actually happened inside a business, which is why records generated with AI in order to sell them are a red flag at SourceX rather than a source of supply.

Why are licensed company records the scarcest category?

Much of the economy's real work happens inside private companies and never reaches the public web. The SBA Office of Advocacy's 2026 small business FAQ reports that small businesses, meaning independent firms with fewer than 500 employees, employ 45.9% of private-sector employees, about 62.3 million workers. The email, tickets, CRM histories and engineering reviews those people produce sit behind logins.

Three things keep this category scarce:

  • Fragmentation. Records live in many systems; most strong companies run 10-15 or more.
  • Rights work. Each company has to confirm it can license its records and agree redaction rules before any work begins.
  • No owner. Nobody inside a typical company is responsible for licensing records, so the opportunity goes unnoticed.

That is the gap SourceX works in. How developer demand for these records is shifting is tracked in AI data licensing trends for 2026.

What this means for referral partners

You do not need to understand pretraining to make a useful introduction. You need to recognize a company whose records belong in the fifth category: a US business with 50+ full-time employees at peak (contractors excluded), several years of documented operations across many systems, outcomes that can be traced, the right to license its records, and an owner or executive who can sponsor a license.

The partner's part is the introduction, made with a referral link or the referral form. From there SourceX screens the company, the company maps its systems in a data inventory, an all-in price is set before any buyer looks, and a signed license leads to delivery and payment. Partners never handle the records. Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company, payable only after the buyer pays and SourceX receives its fee; no reward is guaranteed. How it works walks through each step.

Limits and open questions

  • Developers often do not disclose their full training mix, so the share of each source is not public.
  • The Copyright Office report is a pre-publication version, and this page does not summarize court decisions.
  • The SBA figure covers all small businesses, many of them well below SourceX's baseline.

Next step

Think of one company you know that fits the fifth category and test it with the company fit checker. If it fits, register as a partner and introduce it, or point the owner to sourcex.si/apply with your referral link.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Do AI companies pay for all of their training data?

No. Much of it is collected from the public web without payment, which is where the copyright debate centers. Developers do pay for some sources: licensed archives and platform content, contracted human work and licensed company records. Synthetic data is generated in-house. The mix varies by developer and is often not disclosed, so we know of no reliable public breakdown of paid versus unpaid data.

What makes business records different from web text for training?

Web text mostly shows finished, public-facing writing. Business records show work in progress: a request, the steps taken across several tools, the people involved, the decision and the result. That structure is what agents need to learn from and what evaluators need to test against. It also comes with obligations, since rights must be confirmed and personal details redacted before delivery.

How do developers know licensed data is genuine?

Through the licensing process rather than the data alone. A company's rights are reviewed, its systems are inventoried, and redaction rules are agreed before any work begins, so the dataset arrives with a documented origin. That provenance is one reason SourceX rejects records generated with AI in order to sell them: they would undermine the very thing a buyer is paying for.

Is synthetic data replacing human-created data?

Not replacing it. Synthetic data adds volume and variety, and researchers discuss it as one response to limits on public text, but it is generated from what existing models already learned. Developers still need real examples to train on and to check synthetic output against, and real records of business work are among the hardest of those to obtain.

Can individuals sell their own data for AI training through SourceX?

No. SourceX works only with US companies licensing their own business records, and datasets made up mainly of consumer personal data with no licensing basis are a red flag. Individuals can take part in AI data work as contracted writers, labelers or raters for human-data firms, or join SourceX as referral partners who introduce qualifying companies.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment