Proposed US federal bills on AI training data transparency: what partners should know

Proposed federal AI training data transparency bills would require AI developers to disclose, or let copyright owners discover, what material trained their models. They are proposals until enacted, so check congress.gov for current status. Whatever passes, buyers prefer suppliers that can document where licensed data came from and on what rights basis.

Is there a federal AI training data transparency law?

Treat the federal picture as proposals, not settled law. Bills introduced in Congress would make AI developers disclose what went into their training data, or give copyright owners a way to find out, but a bill binds no one until it is enacted. Before relying on any summary, including this one, look up each bill on congress.gov, where an enacted bill shows a public law number.

For a partner, the practical point is narrower. The proposals are aimed at the developers that build and release models and at the training datasets they use; how far a given bill reaches beyond developers depends on its definitions. Either way, a company that licenses records to AI labs and data buyers should expect to be asked for clear documentation of what the data is and where it came from. That expectation exists today, whatever Congress does. If the vocabulary is new, start with what is AI training data.

What would the proposed bills require?

Policy proposals on this topic broadly fall into three designs. The table describes these generic patterns, not any single bill; the exact duties sit in each bill's current text.

Proposal designWhat it would do, in general termsWho would actWhat a licensing company would notice
Disclosure filingRequire a notice summarizing copyrighted works in a training dataset to be filed with a federal office in connection with a model's releaseThe developer, and possibly whoever assembles or substantially changes the datasetBuyers asking for a precise, written description of the dataset's contents
Discovery for rights holdersGive a copyright owner who believes its works were used a procedure to obtain training records from a developerThe developer, on requestBuyers keeping detailed records of every licensed source
Documentation standardsDirect a federal agency to set standards for the information developers publish about training dataThe developerStandard fields that any dataset description has to fill

Searchers often look for the Generative AI Copyright Disclosure Act and the TRAIN Act by name. Look up each on congress.gov to see its current text and whether it resembles the first two designs. Bills are often reintroduced with new numbers and edits in a later Congress, so search the name on congress.gov and read the latest version and its stage: introduced, reported by committee, passed one chamber, or enacted.

What already applies while bills are pending?

Several existing sources already shape how training data is described and used. None is a federal training data disclosure law, but each matters to a company thinking about licensing.

  • The Copyright Office's study. The Office's Copyright and Artificial Intelligence initiative includes a Part 3 report on generative AI training, released in pre-publication form in May 2025, which examines where copying for training may implicate copyright and how practical licensing approaches are. It is a report, not law; the Copyright Office training report explainer covers it in depth.
  • FTC staff guidance on quiet policy changes. FTC staff wrote in February 2024 that adopting more permissive data practices, such as using consumers' data for AI training, and disclosing that only through a surreptitious, retroactive amendment to terms of service or a privacy policy may be unfair or deceptive. It is staff guidance issued under the prior FTC leadership, not a rule.
  • California's privacy statute. Under the California Consumer Privacy Act, Civil Code section 1798.100 and following, covered businesses must give notice at collection of the categories of personal information, the purposes and whether it is sold or shared, and need a written agreement limiting use when they sell, share or disclose personal information to a service provider or contractor.
  • California's training data documentation law. AB 2013 is a separate state measure aimed at developers; see California's AB 2013 explained.

How do these rules affect the companies partners introduce?

SituationWhat to checkWhat to confirm with counsel
The CEO worries a disclosure law would publish the company's name or recordsWhether the license lets the buyer describe the dataset in regulatory filings, and in what termsHow the company would be identified, if at all, and what confidentiality the license keeps
The records include material the company did not create, such as vendor manuals, purchased research or client documentsWhether the data inventory flags third-party materialWhether to exclude it or obtain permission before licensing
The company recently edited its privacy policy to mention AIWhat the policy said when the data was collectedWhether the planned license fits the promises made at that time
The records include personal information of California residentsNotices given at collection and existing service provider contractsWhich obligations apply and what the license agreement must say
A federal bill passes during the license termWhether the agreement has a change-in-law or cooperation clauseWho bears the cost of any new documentation duty

Why do documentation-ready suppliers win whatever Congress does?

Every design above ends in the same request from a buyer: show what this dataset is, where it came from and what rights stand behind it. A company that can answer quickly is easier to license today and better placed if a law passes later. The same documentation already drives the enterprise data licensing deals that rarely make headlines.

The company prepares this documentation pack with SourceX; the partner does not handle it.

  1. Data inventory: each system, the years it covers, the record types and how they can be exported. The data inventory builder helps a company list systems and records.
  2. Rights basis note: why the company owns or controls each source, including staff-created material and the terms of client contracts.
  3. Exclusion list: material left out, such as third-party content, privileged files or regulated personal data.
  4. Redaction and de-identification rules: agreed with the company before any preparation begins.
  5. Executed license: scope, term, AI-training exclusivity, confidentiality and how the buyer may describe the data in any filing.
  6. Delivery record: what was delivered, when and in what form, produced only after the agreement is signed and the company authorizes delivery.

The record-level fields that make all of this traceable, such as source system, author, timestamp and status, are covered in why metadata raises the value of business data.

Disclosure and consent good practice

  • Read what the company promised staff, customers and clients before licensing anything; privacy policies, employee handbooks and client contracts set the outer limits.
  • Keep the dataset description in the license precise and agreed by both sides, so any future public summary matches what each expects.
  • Do not quietly rewrite a privacy policy to permit AI training after the fact; get advice on notice and consent instead.
  • As a partner, never forward, upload or summarize confidential records. An introduction carries basic fit information and nothing more.

Questions to ask your counsel

  • If a pending federal bill were enacted, would its definitions treat our client as a dataset creator or supplier?
  • What would a required disclosure reveal about a licensor, and can the license limit or shape it?
  • Does the license allocate the cost of new compliance duties between buyer and licensor?
  • Which state laws apply to these records today, and what notices were given when they were collected?
  • Do any client or vendor contracts restrict reuse of the material?

This is general information, not legal, tax or financial advice. Confirm with your own counsel, tax adviser or professional body before acting.

Next step

Legislative uncertainty is a reason to introduce companies that keep good records, not a reason to wait. A strong candidate is a US company with 50+ full-time employees at peak (contractors excluded), several years of documented operations, clear rights to its records and an owner or senior executive able to sponsor the process. After your introduction, SourceX qualifies the company, the company completes its inventory, price and terms are agreed, buyers review, and the company is paid when a deal closes; how it works has the detail.

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company, and rewards become payable only after the buyer pays and SourceX receives its fee. The reward comes out of SourceX's fee, never out of the company's proceeds.

Register as a partner to make introductions, or point an owner to sourcex.si/apply.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

How can I tell whether an AI training data bill has become law?

Look the bill up on congress.gov and check its latest action. A bill becomes federal law only after the House and Senate pass identical text and it clears the presidential step, usually a signature; enacted bills receive a public law number. Introduced or committee-stage bills create no obligations, although they can signal what buyers may start asking for.

Would a transparency law reveal which company licensed the data?

It depends on the text of any bill that passes. The proposals discussed here focus on describing copyrighted works and training sources, and a disclosure could in principle reach the origin of licensed datasets. A company can negotiate how it is described in the license itself, and counsel should review confidentiality and filing language before signing.

Should a company wait for Congress to act before licensing its records?

That is the company's decision with its advisers, but uncertainty is not by itself a reason to wait. Agreements can address future rules through change-in-law and cooperation clauses, and good documentation is useful under any outcome. Nothing is binding until the company agrees price and terms and signs, so exploring a license early commits it to nothing.

Do these proposals create obligations for referral partners?

The proposals discussed here are aimed at developers and training datasets, not at people who make business introductions. A partner's own obligations come from elsewhere: the program terms, any professional rules on referral fees that apply to them, and disclosure when publicly recommending a service they are paid for. Partners never handle records, which keeps them out of the data chain.

How does a federal bill differ from California's AB 2013?

California's AB 2013 is a state law with its own scope and effective dates, while a federal bill imposes nothing unless Congress enacts it. A developer can therefore face state documentation duties even while federal proposals stall. Companies licensing records should ask counsel which state and federal rules apply on the date they sign, because both levels keep changing.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment