Training license vs RAG license: how the two AI data licenses differ
A training license lets an AI developer use a dataset to change a model's weights, usually through a one-time delivery, while a RAG (retrieval) license lets the developer index content and quote it in answers through ongoing access. Training buys learning; retrieval buys current content and attribution. SourceX deals typically center on AI training use.
The short verdict: training licenses buy learning, retrieval licenses buy access
A training license lets an AI developer use a dataset to adjust a model's weights, so the model learns patterns from the records while the records themselves are never shown to end users. A retrieval license, often called a RAG or grounding license, lets the developer store content in a search index and pull passages into answers at the moment a user asks a question.
That single difference drives everything else. Training licenses usually cover a fixed snapshot delivered once and paid once, because the value is in what the model learns. Retrieval licenses usually need ongoing, current access and recurring payment, because the value is in fresh content shown with attribution.
For most established companies the choice is simple. Internal operational records, such as support tickets, project files, CRM histories and approval chains, are confidential and should never be quoted to the public, so a training license, sometimes alongside an evaluation license, is the relevant structure. Retrieval and display licenses suit publishers and platforms whose content is already public-facing.
Training license vs RAG license, side by side
| Dimension | Training license | Retrieval (RAG) license |
|---|---|---|
| What the buyer does | Uses the data to change model weights during training or fine-tuning | Indexes the content and retrieves passages to ground answers at query time |
| When the data is used | Before the model ships, during training runs | Every time a user asks a relevant question |
| Delivery | Usually a one-time transfer of an agreed snapshot | Usually continuous access, such as an API or regular feeds |
| Freshness | Historical depth matters more than recency | Recency is often the main source of value |
| What end users see | Model behavior shaped by the data, not the records themselves | Excerpts, summaries or citations drawn from the content |
| Attribution | Not usually expected | Often required, with links back to the source |
| Payment shape | Typically one price for the dataset and term | Typically recurring fees tied to access or usage |
| Exclusivity | Common for the agreed training use and term | Less common, since publishers want wide distribution |
| End-of-term question | What happens to models already trained on the data | Whether indexed content must be deleted from the system |
| Typical licensor | Companies with deep internal records | Publishers, platforms and reference sources |
The end-of-term row is the one contracts spend the most time on. A model cannot easily forget what it learned, so a training license should say plainly whether models trained during the term can keep being used after it ends.
Where evaluation and display licenses fit
Training and retrieval are not the only rights an AI developer may ask for. Two others come up often, and one agreement can combine several of them.
| License type | What it permits | Restriction to expect |
|---|---|---|
| Evaluation | Using records as a held-out test set to measure model quality | Usually a ban on training with the same records, so the test stays clean |
| Display | Showing excerpts or summaries to end users | Limits on length, attribution rules and linking requirements |
| Training | Learning from the records during training or fine-tuning | Limits on reproducing records verbatim in outputs |
| Retrieval | Pulling passages into answers at query time | Freshness obligations and deletion when the license ends |
Evaluation rights matter because a test set loses its value once a model has seen it; the guide to benchmark contamination explains why never-published records make clean tests.
How public deals have combined these rights
Published announcements show the categories often travel together. These examples involve publishers and platforms, not SourceX deals, and are listed with their dates.
| Date | What was announced | Rights described | Source |
|---|---|---|---|
| December 13, 2023 | Axel Springer and OpenAI partnership | Summaries of selected content shown in ChatGPT with attribution and links, plus use of content to advance model training | OpenAI announcement |
| January 2024, disclosed in an IPO filing | Reddit data licensing arrangements with unnamed licensees | Aggregate contract value of $203.0 million over two-to-three-year terms, delivered through continuous API access plus quarterly data transfers | Reddit registration statement |
| May 16, 2024 | OpenAI and Reddit partnership | Access to Reddit's Data API for real-time, structured content; no financial terms disclosed | OpenAI announcement |
Two lessons carry over to any company. The Reddit figure is a multi-year contract total across several agreements, not annual revenue, so read headline numbers with care. And deals built on continuous access look very different from a one-time training delivery; the overview of enterprise data licensing deals covers the quieter side of the market, where operating companies rather than media brands license records.
Why the rights can be licensed separately
US copyright law treats ownership as divisible. Under 17 U.S.C. 201, ownership may be transferred in whole or in part, and any of the exclusive rights may be transferred and owned separately. That is the legal footing for licensing training use while keeping every other right, or for granting retrieval rights to one party and training rights to another.
Business records also carry contract and confidentiality obligations that copyright does not cover, so scope is set in the license agreement itself. Structure can affect accounting too: Deloitte's guidance on ASC 606 licenses of intellectual property distinguishes a right to use IP as it exists when granted from a right to access IP over the license period, which can change when the licensor recognizes revenue. A company should ask its auditors how a specific license will be treated.
This is general information, not legal, tax or financial advice. Confirm with your own counsel, tax adviser or professional body before acting.
When a training license is the better fit
Choose a training structure when most of these are true:
- The records are internal and confidential: tickets, email threads, CRM stages, approvals.
- The value lies in patterns across years of work, not in last week's content.
- The company prefers a one-time payment to an ongoing technical integration.
- Nobody wants the records quoted, summarized or linked in someone else's product.
- The company can deliver an agreed snapshot once, under redaction rules it sets.
When a retrieval or display license is the better fit
Retrieval and display rights make sense when content is meant to be read by the public and loses value as it ages: news, reference material, product documentation, reviews or forums. The owner usually wants attribution and traffic, can support a live feed or API, and accepts revenue that rises and falls with usage.
Internal business records almost never fit that description. A company selling retrieval access to confidential records would be agreeing to have them surfaced in another firm's product, which few owners, or their clients, would accept.
How SourceX deals are structured
SourceX deals typically center on AI training. They are usually exclusive for AI training for an agreed term, priced as one all-in figure with SourceX's fee included, and paid as a one-time payment, typically within about 60 days of invoicing once the buyer selects the data. The company keeps ownership because the data is licensed, not sold, as the page on licensing versus selling data explains. Nothing is binding until the company agrees price and terms and signs.
Before signing, the company and its counsel should settle these scope questions:
- Which uses are permitted: training, evaluation, retrieval or display.
- How long exclusivity lasts and exactly which use it covers.
- Whether models trained during the term may be used after it ends.
- What limits apply to reproducing records verbatim in model outputs.
- Which redaction and de-identification rules apply before delivery.
- Whether privacy notices and client contracts permit the use; the page on ethically sourced AI training data covers consent and provenance.
Referral partners do not negotiate any of this. A partner makes the introduction; SourceX and the company handle qualification, the data inventory, terms, buyer review and delivery, as laid out in how it works. Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company; the reward becomes payable only after the buyer pays and SourceX receives its fee, and no reward is guaranteed.
Next step
If a client holds years of internal records and asks how an AI license would work, start with the scope questions above. Run a quick, non-binding screen with the company fit checker, then register as a partner to make the introduction. A company can also apply directly at sourcex.si/apply.
Common questions
Can one agreement grant both training and retrieval rights?
Yes. Public announcements show deals that combine training use with display or retrieval of summaries. Each right should be named separately in the agreement, with its own scope, term and restrictions, so neither side assumes a right it was not granted. For confidential business records, most companies grant training and possibly evaluation rights only, and exclude anything that would surface the records to end users.
Does a training license let the buyer quote my documents to its users?
Not by default. A training license covers learning from the records, and the agreement should restrict reproducing them verbatim in model outputs. Redaction and de-identification rules, agreed before any work begins, remove names and sensitive details before delivery. If a buyer wants to show excerpts or summaries to end users, that is a separate display or retrieval right the company can decline.
What happens to a trained model when the training license ends?
That depends on the contract, which is why it should be settled before signing. Removing what a model learned from specific records is difficult, so the agreement needs to say whether models trained during the term may continue to be used, whether new training must stop, and what happens to any copies of the delivered data. The company's counsel should confirm the wording.
Is a RAG license ever a good idea for internal company records?
Rarely, when the licensee is an outside AI developer. Retrieval means passages from the records appear in answers to other people's questions, which conflicts with client confidentiality and most privacy commitments. Many companies build retrieval tools over their own records for internal use, but that is an internal system, not a license. Licensing to developers usually works better as a training or evaluation grant.
Why are AI training licenses often exclusive?
A buyer paying for a dataset wants confidence that competitors are not training on the same records during the same period. SourceX deals are typically exclusive for AI training for an agreed term. That exclusivity is limited to the stated use and term; the company keeps ownership of its data and keeps running its business on it. The exact scope is set in the signed agreement.
Related pages
- Benchmark contamination: why never-published business data matters for AI evaluation
- Enterprise AI data licensing deals: what advisors should know beyond the headlines
- Licensing vs selling data: what is the difference?
- What does 'ethically sourced AI training data' mean?
- How SourceX US company data referrals work
- Check Company Fit for Data Licensing
Free resources
- Time value of money calculator — Future and present value with optional regular payments.
- Business DSCR calculator — Debt service coverage from cash flow and loan terms.
- MCP ROI calculator — Estimate hours saved, implied savings and first-year ROI from MCP.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment