Is a company liable for what an AI model trained on its data does?

Generally the developer, not the data owner, answers for how a model is trained, deployed and used, because it controls those choices, and a well-drafted license says so. The owner remains responsible for its own promises: that it had the rights and permissions to license the records and that it described them accurately.

The short answer

In most training licenses, the developer carries responsibility for what its model does, because it chooses the training method, tests the model, decides how it is deployed and sets the terms for its own users. A well-drafted license puts that in writing. The data owner keeps responsibility for its own promises: that it had the rights and permissions to license the records, that it described them accurately, and that it kept the commitments it made to its customers and staff.

So the practical question is less who is liable for AI and more which promises your company makes, both in the license and to everyone whose information sits in the records. Both are within the owner's control before anything is signed.

Which risks sit with the developer and which with the owner?

Responsibility tends to track who makes the decision. Use this table as a map for reading a draft, not as a statement of what any particular agreement says.

RiskWho makes the decisionWhere a license commonly places it
Training method, model design and testingDeveloperDeveloper
Model outputs, including errors and harmful contentDeveloperDeveloper, often backed by an indemnity
Products built on the model and their end usersDeveloperDeveloper
Use of the data outside the agreed field of useDeveloperDeveloper, as a breach of the license
Security of the data after deliveryDeveloperDeveloper
Rights and permissions to license the recordsOwnerOwner, through warranties
Accuracy of the dataset descriptionOwnerOwner
Promises in the owner's privacy notices and contractsOwnerOwner
Applying the agreed redaction and de-identificationSet in advanceWhichever party the agreement assigns

Two rows deserve emphasis. If a model gives bad advice in a field your records covered, that points to the developer's training and safety choices. If a customer complains that you promised never to share their messages, that points to your promise, whoever trained the model. How the contract converts these lines into indemnities and caps is covered in indemnification in AI training data licenses.

Where does an owner's own exposure come from?

Mostly from commitments the company has already made to other people.

Privacy notices and terms of service. FTC staff have written that companies' promises not to use customer data for undisclosed purposes, such as training or updating models, are enforceable, whether the promise sits in a privacy policy, terms of service or promotional material (FTC staff post, January 2024). The following month, FTC staff warned that adopting more permissive practices, such as sharing data with third parties or using it for AI training, and telling consumers only through a quiet, retroactive change to terms could be unfair or deceptive (FTC staff post, February 2024). Both are staff guidance rather than rules, but they show where scrutiny lands: on what the company promised.

Customer and vendor contracts. Confidentiality clauses in customer agreements can bar reuse of what customers sent you, quite apart from privacy law.

Health information. Records built around protected health information are a red flag unless they are authorized or de-identified. HHS guidance describes two de-identification methods, Expert Determination and Safe Harbor, and health information de-identified under either is no longer protected health information under the HIPAA Privacy Rule (HHS de-identification guidance).

The description itself. If the schedule says eight years of resolved support tickets and the delivery holds three, the accuracy warranty is where the problem lands. Gaps, duplicates and missing outcomes also affect price, as what lowers the value of company data explains.

How does this play out in common situations?

ScenarioCheck firstPoint to confirm with your lawyer
A model gives wrong answers in your industryWhether the license assigns output risk to the developerThe developer handles its model's behavior
A model surfaces a person's name from your ticketsWhether the agreed redaction was applied, and by whomDepends on which party the agreement made responsible for that step
A customer objects that its data was licensedYour privacy notice and that customer's contractExposure depends on what you promised that customer
The buyer says the delivery differs from the descriptionThe dataset schedule against your inventoryOwner exposure sits with the accuracy warranty
The developer uses records beyond training and evaluationThe field-of-use clause and its remediesA breach by the developer, with remedies set in the agreement

What to settle before you sign

  • Read your privacy notices, terms of service and largest customer contracts for promises about reuse.
  • Remove or de-identify personal information as agreed, and keep protected health information out unless authorized or de-identified.
  • Match the dataset schedule to your inventory: systems, years, record types and exclusions.
  • Confirm the field of use is narrow and that resale or onward licensing is addressed.
  • Look for a developer commitment not to attempt re-identification of individuals.
  • Confirm the developer takes responsibility for its models, outputs and products in writing.
  • Check how the indemnities and the liability cap divide the risks in the table above.

The data inventory builder helps an owner list systems and record types with metadata only, which makes the schedule easier to get right.

Questions to ask your counsel

  1. Which promises in our privacy notices and contracts limit what we can license?
  2. Does the draft confine our warranties to the dataset as described?
  3. Does it state plainly that training, outputs and deployment are the developer's responsibility?
  4. What remedies do we have if the developer breaches the field-of-use or re-identification terms?
  5. Do any state or foreign privacy laws apply to people who appear in our records?

This is general information, not legal, tax or financial advice. Confirm with your own counsel before acting.

What should a referral partner say about liability?

Nothing that sounds like advice. Partners make the introduction and share basic fit information; they never handle records and should not offer liability assurances.

Owners who want a wider view can start with the pros and cons of licensing company data to AI developers or with whether licensing company data to AI is ethical.

Next step

Register as a partner to introduce a company whose owner is weighing these questions, or point the owner to sourcex.si/apply. The full sequence, from qualification to delivery, is on how it works.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Can a model trained on licensed records reveal confidential details later?

Models can sometimes reproduce fragments of their training data, which is why redaction and de-identification requirements are agreed before any work begins and why owners look for a commitment not to re-identify individuals. Removing names, account numbers and other identifiers before delivery reduces what could surface. Ask counsel how the agreement handles an incident if one occurs.

Is the data owner responsible if the model turns out to be biased?

Design, training and testing choices sit with the developer, so bias in a model's behavior is generally its responsibility. The owner's part is to describe the dataset honestly, including what it covers and what it leaves out, such as the years, regions or customer types included. Accurate description is a warranty the owner fully controls.

Does de-identification remove all liability?

No. It reduces privacy exposure, and health information de-identified under the HHS methods is no longer protected health information under the Privacy Rule. Other obligations can remain, including confidentiality promises in customer contracts and rights in content the company did not create. De-identification is one control among several, not a release from every claim.

What if the developer breaks the license's use restrictions?

That is a breach by the developer, and the agreement sets the remedies, which can include termination, deletion obligations and damages. The owner's own exposure to customers may still depend on what it promised them, which is why the field of use, re-identification terms and remedies deserve as much attention as the price.

Is the company exposed when a developer is sued over its other training data?

Claims about a developer's other sources, such as content collected from the public web, concern the developer's own choices. An owner's exposure generally relates to its own dataset and its own warranties about rights and accuracy. Check that any indemnity you give is limited to claims caused by the records you delivered, as described in the agreement.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment