Why pre-AI business archives are becoming more valuable to AI developers

Pre-AI data has value because records a company created before generative AI tools went mainstream at the end of 2022 are reliably human-made, with real decisions and outcomes attached. As more new text is machine-assisted, AI developers prize verifiable human records, so a company's 2010-2022 email, tickets and project files are worth reviewing before they are lost.

Why do pre-AI business records carry a premium?

Records a company created before generative AI tools went mainstream at the end of 2022 can be shown to be human-made, and that matters to the people building AI. As more new email, code and documents are drafted with AI assistance, developers increasingly want material that clearly reflects human judgment, with the decisions and consequences that followed.

Picture an engineering firm's 2014 log of requests for information. Each entry holds a contractor's question, the engineer's reply, the drawing revision that resulted and the change order that came later. No model wrote any of it. A company with ten or more years of records like that holds something a newer business cannot manufacture.

What is the low-background steel analogy?

Steel made before the first nuclear weapon tests carries less radioactive contamination than steel made since, so it is prized for instruments that must measure very low levels of radiation. People who work on AI data borrow the idea for text produced before AI writing tools spread: material free of machine-generated content.

The analogy has limits. Not every pre-AI record is valuable, and recent records are far from worthless. But it explains why provenance now carries weight. Developers are wary of training on large volumes of machine-generated text, which is one reason data with a clearly human origin is attractive. This page cites no study on that point.

How scarce is human-made text?

Public human writing is finite. Researchers at Epoch AI estimated the effective stock of human-generated public text at roughly 300 trillion tokens and projected that, if current trends continue, language models will fully use it between 2026 and 2032 (Epoch AI analysis). It is a forecast with wide uncertainty, but it explains why attention has turned to non-public sources, company archives among them.

Company records add something the public web lacks: they show work being done, not just writing about work. The wider gap is described in the work that rarely reaches AI training sets, and the limits of manufactured substitutes in how synthetic environments compare with real business logs.

Which archives count as pre-AI data?

Almost any system that has run for years can hold pre-AI records. The question is whether the history survived.

Record typeWhere it usually livesWhat the original dates proveWhat often goes wrong
Email threadsExchange or Google Workspace, archive files, journaling servicesNegotiations, escalations and decisions in their original sequenceOlder mailboxes left behind in a platform migration
Support and service ticketsHelpdesk or PSA systemsHow problems were diagnosed and resolved, step by stepHistory purged when the platform changed
Proposals, SOWs and RFP responsesShared drives, document managementHow the firm scoped and priced real workOnly final versions kept; drafts deleted
SOPs and policy versionsWikis, intranets, document controlHow procedures changed over timeOld versions overwritten
Code and review historyGit hosting, issue trackersHow engineers reasoned about changesCommits squashed during a migration
Chat historySlack or TeamsDay-to-day coordination and problem solvingEarly years removed by retention settings
CRM activitySales and account management systemsDeals pursued, won and lost, with the reasonsActivity logs dropped when switching vendors

Email is often the deepest of these; why business email archives are valuable for AI explains what makes it useful.

The archive provenance check

Work through this list with whoever runs your IT, whether that is an internal team or an outside provider.

  • We know how far back each main system goes, and someone can confirm it.
  • Past migrations carried history across, or the old system's export was kept.
  • Retired systems, former employees' mailboxes and backups still exist and can be restored.
  • Original timestamps and metadata survive, not just re-saved copies.
  • Our own people created these records in our own work; we are not holding them for clients or bought them from others.
  • Nothing in the archive was generated with AI to pad it out.
  • We are a US company with 50+ full-time employees at peak (contractors excluded), several years of documented operations, and an owner or executive who can authorize a license.

If most boxes are ticked, answer the short questions in the company fit checker for an early, non-binding indication; it does not ask who you are.

What raises or lowers the value of an old archive?

Age helps only when the records are intact, connected and yours.

Raises valueLowers value
Continuous history from well before 2022 to todayGaps of several years after a migration
Many connected systems, typically 10-15+ in strong companiesOne system in isolation
Cases that end in a recorded outcomeFragments with no result attached
Records the company created and ownsRecords that belong to clients or partners
Original metadata and timestampsRe-exported copies with dates lost
Never licensed for AI training beforeAlready licensed for AI training

Which rights questions do older archives raise?

Old records were collected under old promises, and those promises still count.

  • Privacy commitments. FTC staff warned in February 2024 that it may be unfair or deceptive for a company to adopt more permissive data practices, such as using data for AI training, and tell people only through a quiet, retroactive change to its terms or privacy policy (FTC staff post). Check what customers were told when the records were collected.
  • Client contracts. Confidentiality clauses signed years ago may limit what can be licensed today.
  • Personal content. Old mailboxes and chat channels contain personal messages, so the company sets masking and redaction rules for that content before anyone touches the archive.
  • Acquired businesses. Archives that came with an acquisition depend on what the purchase documents transferred.

This is general information, not legal, tax or financial advice. Confirm with your own counsel before acting.

Why does timing matter for owners?

Archives tend to disappear at transitions: a sale, a system consolidation, a wind-down or a retirement. McKinsey reports that more than half of US small-business owners are over 55 and that about six million US small and medium-size businesses will face ownership transitions by 2035 as baby boomers retire (McKinsey, The great ownership transfer). For an owner thinking about the next five years, the practical rule is simple: before any system is retired, keep a complete export.

Companies that are still operating, have been acquired or have wound down can all qualify if the data still exists. Some owners worry about what models trained on their records might do; whether AI trained on a company's records could replace it is answered directly elsewhere.

How does an owner take this forward?

The process is built so the company stays in control at every step.

  1. Run the fit checker, then apply directly at sourcex.si/apply, or ask the adviser who raised the idea to introduce you with their referral link.
  2. SourceX confirms headcount, years of history, systems and rights with you.
  3. Your team lists each system and the years it covers in a data inventory.
  4. You agree one all-in price and the terms, with SourceX's fee included and no separate charges; nothing is binding until you sign.
  5. Buyers, meaning AI labs and other data buyers, look at the opportunity and typically respond within about two weeks once the company is deal-ready.
  6. After signing, the agreed records are prepared under the agreed redaction rules and delivered, and you receive a one-time payment, which typically arrives within about 60 days of invoicing after the buyer has chosen the data.

You keep ownership throughout; the data is licensed, not sold. More detail is in how it works.

Owners and advisers who know other established companies can also refer them. Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company, and the reward becomes payable only after the buyer pays and SourceX receives its fee.

Next step

If your archives reach back well before 2022, apply at sourcex.si/apply. If you advise owners with long histories, register as a partner and make the introduction.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Does data created after 2022 have no value for AI?

It still has value. Recent records reflect current tools, regulations and practices, and buyers weigh freshness alongside depth. The point about pre-AI archives is provenance: older records are easier to accept as human-made. A company with continuous history running from well before 2022 to today can tell the strongest story, because its records show both depth and recency.

Is a 15-year-old archive too old to be useful?

Not necessarily. Older records can still show how decisions were made, escalated and resolved, even if the software of the day has changed. Long histories of five to ten years or more help, especially when archived systems connect to current ones. Value drops if records are fragmentary, unreadable or impossible to export, so condition matters as much as age.

We changed email platforms twice. Is the old history gone?

Not always. Migrations sometimes moved only recent mail, but old exports, archive files, backup tapes, journaling services or a retired server image may still exist. Ask whoever ran each migration what was moved and what was kept. List what you find in a data inventory with the years each source covers, and preserve it before any decommissioning or sale.

Can a company that has been sold or closed still license its pre-AI archives?

Yes, if the data still exists and someone has authority to license it. Operating, acquired and wound-down companies can all qualify. After a sale, the current owner of the assets decides; in an insolvency or wind-down, a court, trustee or assignee may control the records and must be involved before any licensing discussion goes further.

How can anyone tell old records were not generated with AI?

Provenance comes from the systems themselves: original timestamps, message headers, ticket histories and version logs that predate AI writing tools. That is why intact metadata matters more than tidy copies. Records created with AI in order to sell them are a red flag that rules a company out, and an owner should expect to stand behind what it says about where its data came from.

Will licensing old archives expose former employees' messages?

Not without agreed safeguards. De-identification and redaction requirements are set with the company before any work begins, and the company decides which mailboxes, channels or years are in scope. Personal messages, health details and other sensitive content can be excluded or masked. Nothing is delivered until there is an executed agreement and the company has authorized delivery.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment