California's AI training data transparency law (AB 2013), explained

AB 2013 is California's Generative AI Training Data Transparency Act, reported to take effect January 1, 2026. It requires developers of generative AI systems offered in California to post high-level documentation about their training datasets. Suppliers that license records are not the posting party, but buyers have reason to favor data with clear origin and rights.

What is California's AB 2013, in one paragraph?

AB 2013 is California's Generative Artificial Intelligence Training Data Transparency Act. In short, it requires developers of generative AI systems made available to Californians to post documentation on their websites describing the datasets used to train those systems. It is reported to have taken effect on January 1, 2026; confirm the date and any later amendments in the enacted text. Details such as definitions, exemptions and the exact list of required items sit in the statute, so read the enacted text on the California Legislature's website before relying on any summary, including this one.

This is general information, not legal, tax or financial advice. Confirm with your own counsel before acting.

Who has to post documentation, and who does not?

The duty falls on AI developers, not on the companies that supply them with data. A mid-sized US business that licenses its records is a supplier, not a developer, so the statute does not make the supplier post anything. That distinction is the point most partners need.

PartyRole under the lawWhat changes for them
AI developer offering a generative system to CaliforniansCovered by the posting requirementMust publish dataset documentation on its website
Company that licenses records to that developerSource of a datasetWill be asked questions about origin, rights and content
Referral partnerMakes the introduction onlyNo documentation duty; never handles records
SourceXManages licensing between sellers and buyersKeeps rights review and inventory paperwork organized

What does the documentation cover?

Summaries of the bill describe the documentation as a high-level overview rather than a copy of the data. It is described as addressing themes such as where datasets came from, what kinds of data points they contain, whether they include copyrighted or personal information, whether they were purchased or licensed, and when they were collected. Check the statute itself for the precise list and for any carve-outs; this page does not reproduce it.

The practical reading for a supplier: a developer that must describe its datasets in public will prefer data it can describe cleanly. Records with a clear origin, a documented owner and a written license are easier to summarize than scraped material of unknown provenance.

How should a company prepare so its data is easy to document?

Treat it as a paperwork exercise done before any buyer asks. A company that can answer these questions quickly is a better supplier.

  • Origin: which systems produced the records, and over which years?
  • Ownership: did the company create them, and which client or vendor contracts touch them?
  • Personal information: is any present, and how will it be removed or de-identified?
  • Notices: do employee and customer notices, privacy policies and terms permit this use? FTC staff have said in a 2024 blog post that quietly and retroactively changing privacy terms to allow AI use can be unfair or deceptive, so notice history matters.
  • Authority: who can sign a license for the company?
  • Inventory: is there a written list of systems, record types and date ranges?

The data inventory builder helps a company list systems and records, and the company fit checker gives a preliminary screen with no contact details required.

What does AB 2013 mean for a referral partner?

It changes the conversation, not your role. You are not asked to assess compliance, draft anything for a buyer or describe confidential records. You introduce a company that fits the baseline: US, 50+ full-time employees at peak (contractors excluded), several years of documented operations, rights to license and an authorized sponsor.

Where the law helps is the opening line. A company owner who worries that "AI buyers do not care where data comes from" can hear that buyers that must describe their datasets publicly have a reason to favor documented, permissioned sources. That is an inference from the law's purpose, not a measured trend. The broader argument is in the guide on what ethically sourced AI training data means.

Common situations and what to check

SituationWhat to checkTypical outcome to confirm with counsel
Company asks whether it must post documentationWhether it develops and offers a generative AI system itselfUsually no if it only supplies records, but confirm
Company has ageing archives with unclear originContracts, employee notices, system exportsMay need rights review before any license
Records include customer personal informationPrivacy policy, state privacy law, de-identification planOften excluded or redacted before delivery
Buyer asks for a dataset descriptionWhat the company is willing to say about origin and scopeAgreed in the license, not improvised
Company serves clients in regulated sectorsClient contracts and confidentiality termsClient consent may be needed

What to say to a cautious owner

Do not promise that any company will be accepted, priced or paid. Nothing is binding until the company agrees price and terms and signs.

How rewards work

Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company. The reward is paid only after the buyer pays and SourceX receives its fee; an introduction, meeting or signed agreement alone does not trigger payment, and no reward is guaranteed. Licensed professionals should check their own rules on referral fees and disclosure; see the program terms.

Limits of this page

  • It summarizes purpose and structure, not every definition, exemption or deadline in the statute.
  • Amendments, regulations or court challenges may change how the law applies; verify current text.
  • Other states and countries have their own AI and privacy rules, and the enterprise data licensing deals guide covers how deals are structured in practice.

Questions to bring to counsel

Partners are not expected to answer these, but a company owner will want them answered before signing anything.

  1. Does the company itself build or offer any generative AI feature that Californians can use?
  2. Which contracts with clients, vendors or employees restrict use of the records for AI training?
  3. What will the license say about the buyer describing the dataset publicly, and in how much detail?
  4. Which records contain personal information, and what redaction standard applies?
  5. Is exclusivity for AI training acceptable for the agreed term, and what happens to the company's own use of its records?

Companies keep ownership of their data; it is licensed, not sold, and the owner approves scope and price.

Next step

If you know a US company with years of operational records and an owner who wants clean paperwork, register as a partner and make the introduction, or point the owner to sourcex.si/apply. See how it works for the full process.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Does AB 2013 apply to a company that only licenses its records to an AI developer?

The posting duty is aimed at developers of generative AI systems, not at suppliers of data. A company that only licenses records is generally a source of a dataset rather than the party that posts documentation. Whether a specific company is covered depends on the statute's definitions, so it should confirm with its own counsel.

Why would AB 2013 make AI buyers more interested in licensed business records?

A developer that must describe its datasets publicly has a reason to prefer sources it can describe confidently: known origin, documented ownership and a written license. Business records with a clean paper trail fit that need better than material of unclear provenance. This is a reasoned inference, not a promise of demand for any given company.

Do referral partners need to do anything to comply with AB 2013?

No. Partners make introductions and share basic fit information only. They never export, upload or describe confidential records, and they do not prepare dataset documentation. Partners who are licensed professionals should still check their own rules on referral fees and disclosure.

What should a supplier be able to say about its data?

At minimum, which systems produced the records, the years covered, who owns them, whether personal information is present and how it will be handled, and who is authorized to sign. A written inventory makes those answers consistent. SourceX agrees redaction and de-identification requirements with the company before any work begins.

Where can I read the law itself?

Read the enacted text of Assembly Bill 2013 on the California Legislature's official website, and check for later amendments. Summaries, including this page, can lag behind changes. A California-licensed attorney can advise on how the statute applies to a specific company or product.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment