What public AI data partnership programs actually ask companies for

Public AI data partnership programs have mostly asked large content owners for three things: deep archives, fresh or real-time access, and rights to use the material for training and in products, usually over multiyear terms. Announced deals with news publishers and Reddit show the pattern; deals for private operating-company records are rarely announced.

What do AI data partnership programs ask for?

The public record points to three asks: depth, freshness and usable rights. Depth means long archives. When the Associated Press agreed in July 2023 to license part of its text archive to OpenAI, that archive reached back to 1985, AP reported. Freshness means continuing access rather than a single file transfer. Usable rights means clear permission to train on the material and, in some deals, to show it inside AI products with attribution.

"Data partnership program" is a loose label. In practice it covers bilateral agreements between an AI developer and a content owner, usually announced in a short press release with few commercial details. Reading several of those announcements side by side shows a consistent pattern in what developers wanted.

None of the companies named on this page is a SourceX buyer, client or partner. Their deals are cited only as dated, public evidence of what AI developers have asked content owners for.

Which public deals show the pattern?

The table lists announced or reported agreements and what each one disclosed. Dates are when the deal was announced or reported.

DatePartiesAccess the AI developer receivedOther consideration disclosedMoney disclosed
July 2023AP and OpenAIPart of AP's text archive, back to 1985AP gets access to OpenAI technology and product expertiseNot disclosed
December 2023Axel Springer and OpenAI (announcement)Content used to advance model training; summaries of selected articles, including paywalled ones, shown in ChatGPTAttribution and links back to the publicationsNot disclosed
January 2024, disclosed February 2024Reddit and unnamed licensees (Form S-1)Continuous access to Reddit's data API plus quarterly data transfers, for terms of two to three yearsNone statedAn aggregate contract value of $203.0 million across the arrangements and their full terms, not per year
February 2024Reddit and Google, as reported by ReutersReddit content made available to train Google's AI modelsNot reportedReuters cited one unnamed source at about $60 million a year; both companies declined to comment
May 2024Reddit and OpenAI (announcement)Access to Reddit's Data API: real-time, structured contentReddit builds AI features on OpenAI models; OpenAI becomes a Reddit advertising partnerNot disclosed
May 2024News Corp and OpenAI (report)Current and archived content from titles including The Wall Street Journal, New York Post, The Times, Barron's and MarketWatch, under a multiyear agreementNot disclosed by the companiesThe Wall Street Journal reported more than $250 million over five years, in cash and credits for OpenAI technology

The six asks behind public partnerships

Read together, the deals reduce to six asks: depth, freshness, structure, rights, display and term. Each has an equivalent for an operating company that holds internal records rather than published content.

AskWhat developers soughtPublic exampleEquivalent for an operating company
DepthYears of history, not a sampleAP's archive back to 19855-10+ years of tickets, email, CRM and project files, including archived systems
FreshnessContinuing access to new materialReddit's real-time Data APIUsually a defined snapshot, licensed for an agreed term
StructureMachine-readable content with metadataReddit's structured API content and quarterly transfersExports with timestamps, status fields and links between systems
RightsClear permission to trainAxel Springer content used to advance trainingThe company created the records, and its contracts and notices allow licensing
DisplayUse inside products, with attributionChatGPT summaries linking to Axel Springer titlesRarely relevant: internal records are licensed for AI training, not shown to the public
TermMultiyear accessTwo-to-three-year terms in Reddit's filing; News Corp's multiyear agreementA license, typically exclusive for AI training, for an agreed term

The middle four asks transfer most directly. A company whose records are deep, structured, rights-cleared and well documented is answering the same questions a publisher answered, in a different format.

How do operating companies differ from the publishers in the headlines?

Publishers license content made to be read. Operating companies hold records made to get work done: a support ticket escalated and resolved, a quote revised twice, an approval refused with a reason. AI developers building agents need exactly that kind of multi-step, outcome-labeled work, and it is thin on the public web because it was never published.

Those records also carry heavier obligations. Client confidentiality, employee and customer personal data, and contracts that limit reuse all have to be checked before anything moves. That is why a mid-size company rarely fits the press-release model of partnership. It needs a managed process: an inventory of systems, a rights review, agreed redaction rules, one price and delivery only after signing. The guide to enterprise AI data licensing deals covers how those private deals differ from the headlines, and the collected public evidence that AI companies pay for data lists the documented examples.

How can a mid-size company partner with AI labs on data?

Large content owners negotiate directly. A US company with 50+ full-time employees at peak (contractors excluded) usually has no buyer contacts, no pricing reference and no delivery pipeline, so the practical route is a process that supplies all three. With SourceX it runs like this:

  1. The company applies at sourcex.si/apply, often through a partner's referral link, or a partner submits it with the referral form.
  2. SourceX checks the fundamentals: headcount at peak, years of documented operations, how many systems hold records, whether the company owns the rights and who will sponsor the decision.
  3. The company lists each system, how far back it goes and what can be exported.
  4. SourceX and the company settle one all-in price and the license terms before any buyer sees the opportunity.
  5. AI labs and data buyers review it; once a company is deal-ready, they typically respond within about two weeks.
  6. After the agreement is signed, the records are prepared under the agreed de-identification rules and delivered, and the company is paid, typically within about 60 days of invoicing once the buyer selects the data.

The full sequence is on how it works.

What does this mean if you refer companies?

You do not need to explain the whole market to an owner. Three points from the public deals carry weight in a first conversation: buyers value depth, they insist on clean rights, and terms stay confidential, so nobody should quote a price from a headline.

Before you raise it, check the basics:

  • The company is US-based with 50+ full-time employees at peak (contractors excluded).
  • It has kept records across many systems for several years, including archives of retired tools.
  • It created those records itself rather than holding them for clients.
  • You can reach the owner, CEO, CFO or another authorized representative.
  • The same data has never been licensed for AI training.

Run the company fit checker first for a preliminary, non-binding read.

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. The reward becomes payable only after the buyer pays and SourceX receives its fee, and rewards are not guaranteed. Data and analytics consultants, who often see these archives first, have their own guide: referral opportunities for data and analytics consultants.

Limits and open questions

  • Announcements are marketing documents. They describe access and attribution, and rarely mention price, exclusivity or excluded data.
  • Several deals mixed cash with other consideration, such as technology access, an advertising partnership or product credits, so a reported value is not a price per record.
  • The reported figures measure different things: an aggregate contract value over multi-year terms, one anonymous source's annual estimate and a newspaper's five-year total. Do not compare them as like for like.
  • The public deals cluster in news and online platforms and say little directly about demand for internal business records. Whether that demand lasts is debated; see is AI training data demand a bubble?
  • Buyers increasingly ask how material was gathered and whether people were told; ethically sourced AI training data explains what that means in practice.

This is general information, not legal, tax or financial advice. Figures are as disclosed or reported on the dates shown.

Next step

If you know an owner whose company has years of operational records, register as a partner and send your referral link. The company can also apply directly at sourcex.si/apply.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Can a mid-size company apply to an AI lab's data partnership program directly?

It can try, but the announced deals so far involve large publishers and platforms with legal teams and API infrastructure. A mid-size company usually lacks a buyer contact, a pricing benchmark and a delivery pipeline. Going through SourceX gives it an inventory, a rights review, one all-in price and access to several AI labs and data buyers, with nothing binding until it signs.

Do AI data partnerships pay once or as recurring revenue?

The public examples were mostly multiyear arrangements, and Reddit's filing described terms of two to three years. Licenses of internal business records through SourceX work differently: the company receives one all-in price as a one-time payment, typically within about 60 days of invoicing once the buyer selects the data, for a license that is usually exclusive for AI training over an agreed term.

Does licensing data to an AI developer mean giving up ownership?

No. In the public deals the content owners kept their archives and licensed access to them. The same holds for companies working with SourceX: data is licensed, not sold, the company keeps ownership, and it approves the scope, the price and the redaction rules before anything is delivered to a buyer.

Why don't the announcements say what data was excluded?

Press releases are written for customers and investors, not as contract summaries. Exclusions such as personal data, client-owned material or privileged documents are normally handled inside the agreement. With SourceX, those exclusion and redaction rules are written down with the company before preparation begins, so nothing outside the agreed scope is ever delivered.

Are the companies named in these deals SourceX buyers?

No. They are named only because their deals were publicly announced or reported, with dates and sources. SourceX does not name its buyers, and nothing on this page suggests a relationship between SourceX and any company mentioned. The deals are useful only as evidence of what AI developers have asked content owners to provide.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment