The US Copyright Office report on generative AI training, explained

The US Copyright Office's Part 3 report on generative AI training, released in May 2025 as a pre-publication version, discusses where AI training may implicate copyright, how fair use may apply and how practical licensing is. It is regulator analysis, not law, and it does not decide any one company's rights.

What does the Copyright Office report on AI training cover?

Part 3 of the US Copyright Office's "Copyright and Artificial Intelligence" series covers generative AI training. According to the Office's AI initiative page, it examines where copying during training may implicate copyright, how fair use may apply, and how practical licensing approaches are. It was released in May 2025 as a pre-publication version, which means it is regulator analysis and not law.

For partners the useful takeaway is narrow: a federal regulator treated licensing as a real, discussable part of the AI training picture, and it noted that model performance depends heavily on data quality. This is general information, not legal, tax or financial advice. Read the report itself, and ask your own counsel, before relying on any summary.

What is in the series, and where does Part 3 fit?

The Office has published its work in parts. Knowing which is which keeps you from quoting the wrong one.

PartSubjectWhy a partner might care
Part 1Digital replicasMostly about voice and likeness; rarely relevant to business records
Part 2Copyrightability of AI outputsMatters if a company's own content was generated by AI
Part 3Generative AI trainingThe one that discusses copying in training, fair use and licensing

If an owner asks "what does the government say," Part 3 is the document, and the pre-publication label is the first thing to mention.

How do the fair use factors enter the discussion?

Fair use is decided case by case, using a set of statutory factors that look at the purpose of the use, the nature of the work, how much was used and the effect on the market for the original. The report discusses how those factors may apply to training. We do not summarize its conclusions here, because the version is pre-publication and the Office's final position, and the courts' views, may differ.

The practical point for a company thinking about a license is the market-effect factor. When a working market for licensed training data exists, that fact can matter in the analysis, a thread explored in the guide on why a functioning licensing market matters in AI fair use cases. A single court ruling on a legal research tool is covered in Thomson Reuters v. Ross, explained.

How would a company use the report in a licensing decision?

A company does not need to resolve fair use to license. It needs to answer five simpler questions, in this order:

  1. Did we create these records, or do clients and third parties hold rights?
  2. Do our contracts, privacy notices and employee policies allow a license?
  3. Which systems hold the material, and can someone export it?
  4. Are we comfortable with an exclusive AI-training license for an agreed term?
  5. Has any of this data already been licensed for AI training?

These map to the title questions in the explainer on chain of title for training data. The report is background; the contract and the rights review do the real work.

What does the report suggest about licensing as a path?

The Office page lists the practicality of licensing approaches among the topics Part 3 addresses. Beyond that, treat any claim that the report "endorses" or "requires" licensing as unverified. Sellers should hear the honest version: licensing is one lawful route buyers use to get permissioned data, and the terms are set by agreement.

The market context is visible elsewhere. Large public deals are covered in enterprise AI data licensing deals, and a separate guide explains what a large investment in a data vendor signaled. Neither tells a mid-sized operating company what its own records are worth, which is set by the buyers who review its inventory.

Which situations call for extra care?

SituationWhat to checkTypical outcome to confirm
Company content was drafted with generative AIPart 2 of the series on copyrightability of AI outputsRights in AI-assisted material may be narrower
Records include client deliverablesClient contracts and consentClient rights may block a license
Material is mostly public web contentWhether the company owns it at allLittle to license
Customer personal data is mixed inPrivacy notices and redaction plansDe-identification agreed before work begins
Another party controls the assetsCourt, trustee or assignee approvalTheir sign-off is needed first

What should a company read before it talks to a buyer?

Keep the reading list short and ordered, so the owner is not buried in law-firm alerts.

  • The Copyright Office AI page, to see which part of the series is which and whether the training report has been updated.
  • The company's own customer contracts, to see who owns deliverables and whether confidentiality clauses cover internal records.
  • Employee handbooks and notices, which show what was disclosed about monitoring and recording.
  • Privacy notices, which describe what the company promised customers about their personal information.

Only after that does the report help. It frames the market, while the four items above decide what can be licensed. A company that has them in hand shortens the inventory and rights review considerably, because the documents buyers ask for first are already identified.

It also helps to separate three kinds of statements when you hear them. A regulator's report states analysis. A court ruling states a result on specific facts. A contract states what the parties actually agreed. Owners get confused when advisors blur the three, and partners who keep them apart sound credible without giving legal advice. For deeper context on what buyers actually look for, see the explainer on how much data frontier models are trained on, which shows why relevance, not volume, drives interest in business records, and the page on negotiation threads as training data for one example of a record type with real demand.

What should a partner say about the report?

Avoid predicting outcomes, quoting conclusions you have not read, or telling anyone that licensing is required.

How do partner rewards work?

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. The reward is payable only after the buyer pays and SourceX receives its fee, and no reward is guaranteed. The reward is a share of SourceX's fee and is never deducted from what the company receives. If you are a lawyer, accountant or other licensed professional, check your own rules on referral fees and disclosure first, and confirm with your own counsel, tax adviser or professional body before acting.

Next step

Read the report's own page, then test one candidate with the company fit checker and walk through how the process works. If the company fits, register as a partner and introduce the owner.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Is the Copyright Office AI training report binding law?

No. Part 3 was released in May 2025 as a pre-publication version of a Copyright Office analysis. It informs the debate but does not bind courts or companies. Fair use is decided case by case, so confirm any question about your specific situation with your own counsel.

What is Part 3 of the Copyright and AI report about?

According to the Copyright Office, Part 3 addresses generative AI training: where copying during training may implicate copyright, how fair use may apply, and how practical licensing approaches are. Part 1 covers digital replicas and Part 2 covers whether AI outputs can be copyrighted.

Does the report say companies must license their data?

Nothing we can verify says that. The report discusses licensing among its topics but does not turn licensing into a requirement for sellers. A company chooses whether to license, and the terms, scope and price are set by agreement between the company and the buyer.

Why would a business with internal records care about this report?

Most business records are not public, so a buyer can use them only with permission. The report matters mainly as evidence that regulators treat data quality and licensing as part of the AI picture. The company's own rights review still decides whether it can license.

Has the Copyright Office issued a final version of Part 3?

Our source describes the May 2025 release as pre-publication. Check the Copyright Office AI page for the current version before quoting it, because later revisions could change wording. Tell contacts which version you read and when.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment