What are AI scaling laws?
Scaling laws are empirical relationships showing that a neural network's error falls in a predictable way as you increase three inputs: model size, training compute and the amount of training data. They are observations from experiments, not laws of nature, and they let labs forecast how much a model will improve before they spend on training it.
The term matters to referral partners because of the third input. If data is one of the terms that determines results, then the supply of high-quality data influences how far a lab can push a model, which is the root of demand for licensed business records.
How do scaling laws work?
The usual form is a curve. Plot a model's loss on a held-out test against compute, parameters or tokens, and the points tend to follow a smooth line over many orders of magnitude. Two practical ideas follow.
- Balance matters. For a fixed compute budget there is a better mix of model size and training data than simply making the model bigger. Research in the early 2020s, often summarized by the name "Chinchilla," argued that many large models had been trained on too little data for their size.
- Returns diminish. Each doubling of an input gives a smaller improvement, so progress needs steady increases in all three.
Exact exponents and ratios differ by paper, model type and measurement, so treat any single number you see quoted with caution and check it against the original research.
Scaling laws at a glance
| Term | Plain meaning | What it implies for data |
|---|---|---|
| Model size (parameters) | How many adjustable weights the network has | Larger models can use more data productively |
| Compute | Total processing used in training | Spending more compute without more data hits diminishing returns |
| Data (tokens) | Amount of text or other examples trained on | Quality and novelty matter, not just volume |
| Loss | A measure of prediction error | Falls smoothly as the other three increase |
| Compute-optimal | The best split of a fixed budget | Often calls for more data than earlier practice used |
Why does the data term create sustained demand?
Because the data a lab can use is finite and not all of it is equal. Researchers at Epoch AI estimate the stock of public human-written text at roughly 300 trillion tokens and project that, if current trends continue, language models could use most of it between 2026 and 2032. That is a forecast with wide uncertainty, and the same work discusses synthetic data and other ways to extend supply.
Three consequences matter for the licensing market:
- Public web text is being used heavily, so additional gains lean on data that is not public.
- Training agents that carry out tasks needs records of real work, such as decisions and outcomes, that are thin on the open web.
- Labs also need fresh evaluation data, as covered in AI agent benchmarks and private tasks.
This is the logic behind the market described in enterprise AI data licensing deals. Whether any one company's records are worth licensing depends on rights, structure and buyer demand, not on the scaling argument alone.
What scaling laws do not say
- They do not say more data always wins. Low-quality or duplicated data can add little.
- They do not set a price for any dataset or guarantee any deal.
- They do not prove that new architectures or synthetic data will not change the picture. Whether distillation reduces the need for data is a live debate.
- They describe averages across experiments, so a specific model can behave differently.
How should a referral partner use this?
You do not need to explain the mathematics. You need an honest one-line reason why buyers exist, plus the discipline to leave pricing and technical claims to SourceX. SourceX manages sourcing, rights review, contracting, delivery and payment between companies and AI labs and data buyers, as described in what is an AI data intermediary, and it does not train models. For common misconceptions, see AI data licensing myths versus facts.
Be clear about what is licensed and what is not. Licensing differs from simply sharing data, as the comparison of data licensing versus data sharing explains, and chain of title questions are covered in chain of title for training data. To understand who is on the other side of the deal, read what is a data buyer. This is general information, not legal, tax or financial advice.
Next step
Pick one US company you know with years of records and 50+ full-time employees at peak (contractors excluded), run it through the company fit checker and review how it works. Then register as a partner. Partners earn 25% of the eligible platform fees SourceX actually collects, capped at $100,000 cumulative per referred company, only after the buyer pays and SourceX receives its fee. No reward is guaranteed.