How large language models are trained, stage by stage, in plain English

Large language models are trained in stages: pretraining on huge volumes of text to predict the next word, supervised fine-tuning on example prompts and good responses, preference or reinforcement training that rewards better answers and actions, and evaluation on held-out tests. Licensed business records fit mostly in the later stages, where realistic examples of real work are scarce.

The short answer: four stages, each fed by different data

Large language models are trained in sequence. First comes pretraining, where the model reads an enormous volume of text and learns to predict the next word. Post-training then turns that raw predictor into a useful assistant: supervised fine-tuning on worked examples, followed by preference and reinforcement training that rewards better responses. Evaluation runs throughout, testing the model on material it has never seen.

Each stage needs a different kind of data, and that is where licensed business records come in. Pretraining wants volume. The later stages want quality: realistic examples of tasks done well, clear outcomes, and tests the model could not have memorized. Records of real company work are scarce on the public web, so they tend to matter most in those later stages.

The training pipeline at a glance

StageWhat the model learnsTypical dataWhere business records can fit
Data preparationNothing yet; this step decides what the model will seeWeb pages, books, code and licensed collections, filtered and de-duplicatedLicensed, rights-cleared records enter here with documented provenance
PretrainingGrammar, facts, reasoning patterns and writing stylesVery large volumes of general text, split into tokensA small but distinctive share: professional writing from real workplaces
Supervised fine-tuningHow to follow instructions and complete tasksExample prompts paired with high-quality responsesWorked examples: a request, the steps taken and the result
Preference and reinforcement trainingWhich responses and actions are betterHuman rankings, automated graders and practice tasksRecords that show outcomes, such as approved or rejected, won or lost
EvaluationNothing; this step measuresHeld-out test sets the model has not seenNever-published records with known right answers

Stage 1: pretraining teaches language from text at scale

Pretraining is the expensive part. Text is split into tokens, which are words or word fragments, and the model repeatedly guesses the next token in a passage, then nudges billions of internal numbers, called weights, to guess better next time. Repeat that across a vast corpus and the model absorbs vocabulary, grammar, facts and patterns of reasoning without anyone labeling anything.

The constraint is supply. Researchers at Epoch AI estimated the effective stock of human-generated public text at roughly 300 trillion tokens and projected that, if current trends continue, language models will have fully used it at some point between 2026 and 2032 (Epoch AI). That forecast carries wide uncertainty, but it explains why developers look beyond the open web.

Stage 2: supervised fine-tuning teaches the model to follow instructions

A pretrained model can continue a passage, but it does not reliably answer questions or carry out requests. Supervised fine-tuning fixes that by training on curated examples: a prompt, followed by the response a skilled person would give. These datasets are far smaller than pretraining corpora, and quality counts for more than quantity.

Fine-tuning is also how developers specialize a model, for example toward customer support, finance operations or software engineering, by adding examples from that field.

Stage 3: preference and reinforcement training shape behavior

Next, the model learns which of its own outputs are better. In preference training, people compare two responses and pick the stronger one, and those choices train a reward signal that steers the model toward helpful, accurate answers. In reinforcement learning, the model attempts tasks, a grader checks the result, and successful attempts are reinforced.

Reinforcement learning grows in importance as models become agents that carry out multi-step work inside software. Developers build practice workspaces for this, described in the explainer on RL environments.

Stage 4: evaluation tests the model on work it has not seen

Evaluation is less a final step than a constant check. Developers run models against held-out test sets to measure accuracy, safety and task completion, and compare versions before release. A test only works if the model did not see its questions during training.

There is a quieter risk as well. As more web text is itself generated by AI, models trained heavily on it can degrade over generations, the problem covered in what model collapse is. Fresh human-created records help on both fronts: they are unseen, and they reflect how people actually work.

What makes a business record usable for training

Useful is not the same as permitted. A record only helps a developer if the company can license it and the developer can document where it came from. Three checks come first:

  • Ownership: the US Copyright Office's circular on works made for hire explains that material employees create within the scope of their jobs is generally owned by the employer, while content from outside contractors may not be unless rights were assigned in writing.
  • Third-party content: documents that belong to clients, or that were licensed from vendors, usually need the other party's agreement or must be excluded.
  • Personal data: names, contact details and sensitive information are redacted or de-identified under rules the company agrees before any work begins.

This is general information, not legal, tax or financial advice. Confirm with your own counsel, tax adviser or professional body before acting.

The explainer on ethical sourcing of AI training data sets out why provenance and consent weigh heavily with developers, and enterprise AI data licensing deals beyond the media headlines shows how operating companies take part.

What this means if you advise companies

You do not need to understand the mathematics to spot a client whose records suit the later training stages. Look for work recorded end to end, with results attached:

  • The company has 50+ full-time employees at peak (contractors excluded) and several years of documented operations.
  • Work runs through many systems, such as email, chat, CRM, finance, support and engineering tools, and retired systems were archived rather than deleted.
  • Records show complete tasks with outcomes, not just final documents.
  • The company created the records, and its contracts and notices allow licensing.
  • An owner, CEO, CFO or other authorized representative would sponsor the conversation.

The company fit checker runs a preliminary, non-binding version of these checks. Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company; the reward becomes payable only after the buyer pays and SourceX receives its fee, and no reward is guaranteed. The rewards policy explainer sets out the details in plain language.

Next step

If a client's records read like worked examples with outcomes, register as a partner and send the introduction. SourceX handles qualification, the data inventory, terms and delivery, and the partner never touches the data.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

How long does it take to train a large language model?

It varies widely with model size and budget. Pretraining a frontier model typically runs for weeks to months on large clusters of specialized chips, and post-training continues in rounds as developers fine-tune, gather feedback and test new versions. Fine-tuning an existing model for one field can take days or less. Developers rarely publish exact timelines for their largest models.

Is a chatbot like ChatGPT trained the same way?

Broadly, yes. Major chat assistants are built on a pretrained language model that is then fine-tuned on example conversations and refined with preference or reinforcement training, with evaluation before each release. Developers differ in the data, scale and techniques they use, and each developer's own published model documentation is the authority on the specifics of its product.

Does a model memorize the documents it is trained on?

Mostly it learns patterns rather than storing documents, but models can sometimes reproduce passages they saw many times or that are highly distinctive. That is why licenses for business records include redaction and de-identification rules agreed before delivery, and why agreements limit verbatim reproduction in outputs. Removing sensitive details before training is more reliable than trying to remove them afterwards.

Can a model be trained only on one company's records?

Not from scratch in any practical sense, because one company's records are far too small to teach general language. What a company's records can do is improve an already pretrained model through fine-tuning, preference training or evaluation, adding examples of real work the public web lacks. Developers usually combine records from many sources, each licensed separately.

Why don't developers rely only on synthetic data?

Synthetic data, meaning text generated by other models, is useful for some tasks and widely used. But it mostly reflects what existing models already know, and heavy reliance on model-generated text can degrade quality over successive generations. Human-created records add information models do not yet have: real decisions, real exceptions and real outcomes from everyday work.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment