How much data are frontier AI models trained on, and does size matter?
Frontier language models are trained on trillions of tokens. Epoch AI estimates the public human-written text supply at roughly 300 trillion tokens and projects it could be fully used between 2026 and 2032. A company archive is far smaller, so relevance, not size, drives its value.
How much data are AI models trained on?
Frontier language models are trained on trillions of tokens, a token being a small chunk of text, roughly a short word or word fragment. Exact figures for specific models are often not disclosed, so rely on research estimates of the available supply. Epoch AI estimates the effective stock of human-generated public text at roughly 300 trillion tokens (90% confidence interval 100 trillion to 1,000 trillion) and projects that language models could fully use it between 2026 and 2032 if current trends continue, according to its published research.
That is a forecast with wide uncertainty, not a measurement of any one model. Use it to explain scale and scarcity, and avoid quoting it as the size of a particular system.
What is a token, in plain terms?
Models read text as tokens, not words. In English, a token is often a bit under one word on average, so a million tokens is somewhere around three-quarters of a million words. Treat that as a rule of thumb; the ratio varies by language, formatting and code.
| Unit | Rough meaning | Why it matters here |
|---|---|---|
| Token | Chunk of text a model reads | Training size is counted in tokens |
| Million tokens | Several novels' worth of text | Small by model standards |
| Billion tokens | A large corporate archive of text | Still small next to public corpora |
| Trillion tokens | Scale of large pretraining sets | Needs the public web plus other sources |
How does a company archive compare with that scale?
Honestly, it is small. Illustrative example: a company with 200 employees and eight years of records might hold a few billion tokens of usable text across email, chat, tickets and documents. Against a public stock estimated in the hundreds of trillions, that is a tiny fraction, on the order of one hundred-thousandth of the estimate.
If size drove value, company archives would not matter. They matter because of what they contain.
Why does relevance beat size?
Buyers are not trying to add more generic text. They want material that their data does not already have, such as:
- Multi-step work histories with the outcome attached, like a ticket that was resolved or escalated.
- Records from domains that are thin on the public web, such as procurement, finance operations and engineering review.
- Material with clear provenance, so the buyer knows the company created it and can license it.
The explainer on why AI needs so much data separates pretraining from post-training and evaluation, which is where small, high-quality sets are most useful. Negotiation records are a good example, covered in negotiation threads as training data.
What does this mean for a referral partner?
Do not sell a company on its size in tokens. Sell on depth and connection: years of records, many systems, outcomes recorded, rights in place. The SourceX baseline is 50+ full-time employees at peak (contractors excluded), several years of documented operations, rights to license the data and an authorized sponsor.
Strong companies typically keep records across 10-15+ systems, and long histories help. The page on what AI training data is defines the term simply, and what an AI lab is explains who the buyers are.
What are the limits of the numbers?
- The 300 trillion figure covers public human-generated text, not private company data.
- Its range spans an order of magnitude; treat it as an estimate.
- The 2026-2032 window depends on training trends and could move earlier if models are heavily overtrained.
- Synthetic data and efficiency gains could change demand.
How do I explain this to a client?
Compare this with chain of title for training data, which explains why rights matter as much as content.
How are partners rewarded?
Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. The reward is payable only after the buyer pays and SourceX receives its fee, and no reward is guaranteed. It is a share of SourceX's fee and is never deducted from what the company receives.
Next step
Skip the size debate and screen for depth with the company fit checker, then see how it works and enterprise AI data licensing deals. When a company fits, register as a partner.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
How many tokens was GPT trained on?
Developers often do not disclose exact figures for specific models, so be cautious with numbers you see online. A safer statement is that frontier models are trained on trillions of tokens, and that research estimates of the whole public text supply are in the hundreds of trillions.
What does Epoch AI estimate about public text?
Epoch AI estimates the effective stock of human-generated public text at roughly 300 trillion tokens, with a 90% confidence interval of 100 trillion to 1,000 trillion, and projects language models could fully use it between 2026 and 2032. It is a forecast with wide uncertainty.
Is a company archive too small to matter to AI developers?
Small does not mean unimportant. A company archive is tiny next to the public web, but buyers look for relevance: connected workflows, recorded outcomes and clear rights. Size in tokens matters less than whether the records show real work that public text lacks.
How many tokens does a typical company have?
There is no typical number, and it depends on headcount, years of operation and how many systems were used. Partners should not estimate token counts. The company's data inventory with SourceX lists systems, years of history and what can be exported.
Does more data always make a model better?
Not always. Quality, relevance and coverage matter, and repeated or low-quality text adds little. That is why developers look for new kinds of records, and why a smaller, well-documented dataset can be more useful than a large generic one.
Related pages
- Why does AI need so much data, and why new records matter most
- Negotiation threads as AI agent training data
- What is AI training data?
- What is an AI lab? Frontier labs and model developers explained
- What is chain of title for AI training data?
- Enterprise AI data licensing deals: what advisors should know beyond the headlines
Free resources
- EBITDA calculator — Reported and adjusted EBITDA from net income.
- MOIC calculator — Multiple on invested capital from realized and unrealized value.
- PDF bank statement to CSV converter — Turn Chase, Bank of America or Wells Fargo PDF statements into CSV, privately in your browser.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment