What are AI scaling laws, and why does the data term matter?

Short answer

AI scaling laws are empirical patterns showing that a neural network's error falls predictably as model size, training compute and training data increase. Because data is one of those inputs and public text is finite, scaling laws help explain sustained demand for new, high-quality, rights-cleared data such as licensed business records.

What are AI scaling laws, and why does the data term matter?: overview of What are AI scaling laws?, How do scaling laws work?, Scaling laws at a glance, Why does the data term create sustained demand?, What scaling laws do not say
Covered on this page: What are AI scaling laws? · How do scaling laws work? · Scaling laws at a glance · Why does the data term create sustained demand? · What scaling laws do not say

What are AI scaling laws?

Scaling laws are empirical relationships showing that a neural network's error falls in a predictable way as you increase three inputs: model size, training compute and the amount of training data. They are observations from experiments, not laws of nature, and they let labs forecast how much a model will improve before they spend on training it.

The term matters to referral partners because of the third input. If data is one of the terms that determines results, then the supply of high-quality data influences how far a lab can push a model, which is the root of demand for licensed business records.

How do scaling laws work?

The usual form is a curve. Plot a model's loss on a held-out test against compute, parameters or tokens, and the points tend to follow a smooth line over many orders of magnitude. Two practical ideas follow.

  1. Balance matters. For a fixed compute budget there is a better mix of model size and training data than simply making the model bigger. Research in the early 2020s, often summarized by the name "Chinchilla," argued that many large models had been trained on too little data for their size.
  2. Returns diminish. Each doubling of an input gives a smaller improvement, so progress needs steady increases in all three.

Exact exponents and ratios differ by paper, model type and measurement, so treat any single number you see quoted with caution and check it against the original research.

Scaling laws at a glance

TermPlain meaningWhat it implies for data
Model size (parameters)How many adjustable weights the network hasLarger models can use more data productively
ComputeTotal processing used in trainingSpending more compute without more data hits diminishing returns
Data (tokens)Amount of text or other examples trained onQuality and novelty matter, not just volume
LossA measure of prediction errorFalls smoothly as the other three increase
Compute-optimalThe best split of a fixed budgetOften calls for more data than earlier practice used

Why does the data term create sustained demand?

Because the data a lab can use is finite and not all of it is equal. Researchers at Epoch AI estimate the stock of public human-written text at roughly 300 trillion tokens and project that, if current trends continue, language models could use most of it between 2026 and 2032. That is a forecast with wide uncertainty, and the same work discusses synthetic data and other ways to extend supply.

Three consequences matter for the licensing market:

  • Public web text is being used heavily, so additional gains lean on data that is not public.
  • Training agents that carry out tasks needs records of real work, such as decisions and outcomes, that are thin on the open web.
  • Labs also need fresh evaluation data, as covered in AI agent benchmarks and private tasks.

This is the logic behind the market described in enterprise AI data licensing deals. Whether any one company's records are worth licensing depends on rights, structure and buyer demand, not on the scaling argument alone.

What scaling laws do not say

  • They do not say more data always wins. Low-quality or duplicated data can add little.
  • They do not set a price for any dataset or guarantee any deal.
  • They do not prove that new architectures or synthetic data will not change the picture. Whether distillation reduces the need for data is a live debate.
  • They describe averages across experiments, so a specific model can behave differently.

How should a referral partner use this?

You do not need to explain the mathematics. You need an honest one-line reason why buyers exist, plus the discipline to leave pricing and technical claims to SourceX. SourceX manages sourcing, rights review, contracting, delivery and payment between companies and AI labs and data buyers, as described in what is an AI data intermediary, and it does not train models. For common misconceptions, see AI data licensing myths versus facts.

Be clear about what is licensed and what is not. Licensing differs from simply sharing data, as the comparison of data licensing versus data sharing explains, and chain of title questions are covered in chain of title for training data. To understand who is on the other side of the deal, read what is a data buyer. This is general information, not legal, tax or financial advice.

Next step

Pick one US company you know with years of records and 50+ full-time employees at peak (contractors excluded), run it through the company fit checker and review how it works. Then register as a partner. Partners earn 25% of the eligible platform fees SourceX actually collects, capped at $100,000 cumulative per referred company, only after the buyer pays and SourceX receives its fee. No reward is guaranteed.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Who discovered neural scaling laws?

Several research groups documented them in the early 2020s by measuring how language model error changed with size, compute and data. The details and ratios differ between studies, so cite the specific paper you rely on instead of a general claim, and treat any single headline number as approximate.

What does the Chinchilla result mean in simple terms?

It argued that for a given compute budget, many large models had been trained on too little data and that size and data should grow together. The takeaway for the data market is that data volume and quality are an input to performance, not an afterthought.

Do scaling laws mean AI will always need more data?

Not necessarily. They describe past trends, and techniques such as better data selection, synthetic data and distillation may change how much new data is needed. The safer statement is that data remains one of the main inputs, and that high-quality, rights-cleared data is scarce.

Do scaling laws set prices for licensed datasets?

No. They describe how model quality changes with inputs, not what a buyer will pay. Price depends on the dataset's uniqueness, structure, rights and buyer demand, and is agreed between the company and SourceX before buyers review. Nothing is binding until the company signs.

Why does this matter if I only refer companies?

It helps you explain why the market exists without overstating it. You can say buyers look for real records because public text is limited, then leave technical and pricing questions to SourceX. Credibility with owners comes from plain, accurate statements and no promises about deals.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment