Is AI training fair use, and what does it mean for licensed business data?

There is no blanket answer: in the US, whether AI training is fair use is decided case by case, weighing four factors against the facts of each lawsuit, and the litigation is still moving. Licensed business records largely avoid the question, because the buyer trains under express permission in a signed license rather than relying on a fair-use defense.

The short answer

It depends, and no single ruling settles it. Under US copyright law, fair use is a defense that a court weighs for a particular use of particular works, so whether training an AI model is fair use gets decided case by case, on the facts of each lawsuit. As of October 2026, AI developers, authors, publishers and other rights holders are still arguing those facts in court.

For business records licensed under a signed agreement, the question mostly falls away. A buyer that licenses a company's support tickets and project files for an agreed term relies on permission, not on a defense. What remains is a narrower question: does the company granting the license actually hold the rights it is granting?

What does fair use weigh?

Fair use allows some unlicensed uses of copyrighted works. Courts consider four factors together, and no single factor decides the outcome. The US Copyright Office's Copyright and Artificial Intelligence initiative is where the agency publishes its analysis of how those factors may apply to training.

FactorWhat a court asksHow it surfaces in AI training disputes
Purpose and character of the useIs the use transformative, and is it commercial?Whether training serves a different purpose from the original works
Nature of the workIs the work creative or factual, published or unpublished?Novels, music and art sit differently from reference and factual material
Amount usedHow much of the work was copied, and was it the heart of it?Training pipelines often copy entire works
Market effectDoes the use harm the market for the work, including licensing markets?Whether a working market for licensed training data exists and is being displaced

The fourth factor is where licensing enters the legal debate. Whether rights holders and developers already license training material to each other is part of the market-effect question.

What has the US Copyright Office said?

The Office is publishing a multi-part report on copyright and AI. Part 3, which covers generative AI training, appeared in May 2025 as a pre-publication version. It examines which steps in building a model may involve copying that implicates copyright, how fair use may apply, and whether licensing approaches are practical. It also observes that a model's performance depends heavily on the quality of its data.

Two cautions before quoting it. The report is the agency's analysis, not law, and courts are not bound by it; judges decide fair use. And because it was released in pre-publication form, check the Office's site for the current version. This page deliberately does not summarize the report's conclusions; read them in the original or ask counsel.

Where do the US court cases stand?

This page does not report individual rulings, because we have not verified them against the court opinions and the picture keeps shifting as cases move through appeals, settlements and new filings. Ask counsel for the current status of any case, and read any ruling through five questions rather than relying on a headline.

The ruling reader:

  1. Which court decided it? A federal trial court decision does not bind other courts. An appeals court decision binds trial courts in its circuit, and only the Supreme Court sets a nationwide rule.
  2. At what stage? A ruling on a motion to dismiss, a summary judgment decision and a jury verdict carry different weight, and a settlement sets no precedent at all.
  3. What was copied, and how was it obtained? Where the training copies came from can matter as much as what the training did with them.
  4. What does the model produce? Outputs that reproduce or substitute for the original works raise different questions from outputs that do not.
  5. Is there a licensing market for that material? Evidence of willing licensors and paying licensees feeds straight into the market-effect factor.

For the commercial side of that last question, AI data licensing trends for 2026 tracks how the market is developing.

Why do licensed business records sidestep the question?

Because the buyer has written permission for an agreed scope, and because the records were never public to begin with.

QuestionPublic web content gathered without a licenseBusiness records licensed under a signed agreement
How the buyer gets accessCrawling or bulk downloadsAn export by the company after it authorizes delivery
Legal basis for trainingA fair-use defense, if challengedExpress permission in the license
Who confirms the rightsNobody, in advanceThe licensing company, through a rights review before delivery
Privacy and redactionWhatever happened to be publishedDe-identification and redaction rules agreed before work begins
ExclusivityNone; anyone can gather the same pagesTypically exclusive for AI training for an agreed term

Three points explain the gap:

  • Fair use does not open doors. Even a ruling broadly favorable to training would give nobody access to a company's email, Slack or Teams threads, CRM history or ticket queue. Those records sit behind company systems, and the practical route to them runs through the company.
  • Copyright is only one right in play. Many operational records are largely factual, and copyright protects original expression, so confidentiality duties, contracts and privacy law often matter as much as copyright for this material.
  • Buyers value a paper trail. A signed license documents what the buyer may do, for how long and under which de-identification rules, which matters to developers that need a clear record of where their training data came from.

Developers have signed licenses even while the fair-use question stays open. The Associated Press reported in July 2023 that it would license part of its text archive, dating back to 1985, to a major AI developer. Business records add something news archives lack: the exceptions, escalations and decisions behind real work, which is why AI models need rare examples. For the full map of sources, see where AI training data comes from.

Does the company actually hold the rights it licenses?

This is where the real work sits, because a license is only as good as the licensor's rights.

  • Staff-created material. The Copyright Office's circular on works made for hire explains that the employer, not the employee, is treated as the author of material prepared as part of the job. Freelance material is treated that way only in listed categories with a signed written agreement, so contractor work needs checking.
  • Partial rights. Copyright can be divided: 17 U.S.C. section 201 allows any exclusive right to be transferred and owned separately. That is why a company can grant AI-training rights for a set term while keeping ownership of its records.
  • Third-party material. Purchased research, vendor manuals, licensed images and files supplied by clients arrive with their own terms and are normally carved out unless the owner agrees.

How does this play out in common partner situations?

Situation a partner hearsWhat to checkTypical outcome to confirm with counsel
The owner says AI developers will take the data anywayWhether the records are public at allInternal systems are not public; a license is how buyers reach them
The owner worries a license concedes that training infringesWhether the license needs to take any position on the lawsuitsA license is a commercial permission, not a legal admission
Marketing and engineering work was done by freelancersWritten assignments or work-made-for-hire termsCovered material stays in scope; the rest is assigned or excluded
Shared drives hold purchased reports and vendor manualsThe terms those materials came withUsually carved out of the dataset
The company runs a public blog, help center or forumWhether that content differs from the private records on offerPublic pages are a separate question; licenses focus on non-public records
The owner wants to keep ownershipHow the license defines scope, term and exclusivityAI-training rights for an agreed term, with ownership staying with the company

Disclosure and consent good practice

Partners make introductions and share basic fit information. They do not give legal opinions or handle records.

  • Do not offer a view on whether a client's records would be fair game without a license; point to the company's counsel and the rights review.
  • Never forward records, screenshots or samples to show what a company holds.
  • Tell the company you may earn a referral reward; licensed professionals should check their own rules on referral fees and disclosure.
  • Expect the company to approve scope, de-identification and redaction before any work starts; nothing is delivered without an executed agreement and the company's authorization.

If a client raises the lawsuits, a short answer works:

Questions to ask counsel

The company's counsel, or yours if you advise the company:

  1. Which of our records include third-party copyrighted material, and how should it be excluded?
  2. Do our contractor, agency and freelancer agreements assign work product to us in writing?
  3. Do customer contracts, employee notices or our privacy policy limit using records for AI training?
  4. What rights and warranties would we give in a license, and how is liability limited?
  5. Does any recent AI-copyright decision change our analysis, and which cases should we watch?
  6. For partners who hold a professional license: does my professional body restrict referral fees, and what must I disclose?

This is general information, not legal, tax or financial advice. Confirm with your own counsel, tax adviser or professional body before acting.

Where does this leave a referral partner?

You do not need a view on fair use to make a useful introduction. The better question for a client is practical: did the company create its records, and does it control the systems they sit in? For the other legal layers, read whether it is legal to license business data to AI developers; for how the public deals were structured, see enterprise AI data licensing deals.

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, up to $100,000 per referred company, payable only after the buyer pays and SourceX receives its fee. No reward is guaranteed, and the reward is never deducted from what the company receives.

Next step

If a client has years of records it created itself, run it through the company fit checker and read how it works. Then register as a partner to make the introduction, or have the company apply directly at sourcex.si/apply.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

If courts decide AI training is fair use, will companies still be paid for their data?

The case for licensing private records does not depend on that outcome. Fair use is a defense for copying material a developer can already reach, and it gives no access to a company's email, tickets, CRM history or engineering records, which sit behind company systems. Buyers pay for that access, for complete workflows with their outcomes and for documented rights. Demand for any specific dataset is never assured, so each company is assessed on its own records.

Does licensing our records mean we are taking a side in the AI copyright lawsuits?

No. A license is a commercial agreement that sets what a buyer may do with specific records, for how long and under which de-identification rules. It does not require the company to take a position on whether training on other people's works is fair use. Companies worried about public perception can ask counsel how the agreement describes the use and what confidentiality terms apply to the deal itself.

Can a company license records it bought or collected from others?

Only to the extent it holds the rights. Purchased research, vendor manuals, licensed images and files supplied by clients usually come with terms that limit reuse, so they are normally carved out unless the original owner agrees. Records generated with AI in order to sell them are a red flag. The rights review before any delivery is where such items are found and either excluded or cleared.

Are employee emails and chat messages copyrighted, and who owns them?

Some may be, where they contain original expression, and material employees write as part of their jobs is generally treated as the employer's under the work-made-for-hire rules. Ownership is only part of the picture. Employee notices, confidentiality duties, privacy law and customer contracts also shape what can be licensed, which is why de-identification and redaction rules are agreed with the company before any work begins.

What happens to a license if a court ruling changes the law during its term?

The signed agreement governs the relationship, so the answer sits in its terms. Before signing, the company's counsel can ask how the agreement handles changes in law, what each side warrants about rights and how disputes are resolved. A license that rests on permission is less exposed to shifts in fair-use case law than unlicensed training, but counsel should confirm how that applies to the specific agreement.

Can a partner tell a client that licensing is legally safe?

No. Partners make introductions and share basic fit information; they should not give legal opinions. Point the client to its own counsel and to the rights review that happens before any delivery, which covers ownership, contracts, privacy and de-identification. It is reasonable to explain that a license gives the buyer permission for the licensed records; promising that a dataset is free of rights issues is not.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment