Bartz v. Anthropic explained: the authors settlement and what it means for licensed data

Bartz v. Anthropic is a copyright case brought by book authors in 2024. In June 2025 the court found training on lawfully bought books was fair use but sent claims over pirated library copies toward trial, and the parties then agreed a class settlement. The lesson for licensed data: how material was acquired matters, so documented provenance carries weight.

The short answer

How training material was obtained matters, and Bartz v. Anthropic is the clearest public example. In June 2025 the federal judge hearing the case in the Northern District of California ruled on fair use: using lawfully purchased books to train a language model was fair use, and so was scanning purchased print books into digital copies, but downloading and keeping copies from pirate library sites was not excused, and those claims were headed for trial. In September 2025 the parties announced a class settlement of the pirated-copies claims, which then went through the court's approval process.

For anyone licensing data, the lesson is narrow but real. A favorable view of training as fair use did not protect material acquired without permission. Buyers now ask harder questions about where data came from and who had the right to provide it, and records licensed by the company that created them answer those questions directly.

Anthropic is named here only because it was a party to a public court case. Nothing on this page identifies it, or any other party, as a SourceX buyer, client or partner.

What happened in the case?

WhenStepWhy it matters for licensed data
2024Book authors sued over the use of their books to build AI modelsPublished books were the training material at issue
June 2025Ruling on fair use: training on lawfully acquired books and scanning purchased print copies were fair use; pirated library copies were not excusedThe method of acquisition separated lawful from unlawful copying
July 2025A class of rights holders in the pirated books was certifiedThe remaining exposure turned on how copies were obtained
September 2025Class settlement announced, then submitted for court approvalThe settled claims concerned the pirated copies
Since thenApproval, claims and payment stepsCheck the case docket for current status

The dates and holdings above summarize the court's public orders and press coverage; verify each against the case docket before relying on it. For amounts, eligibility rules and deadlines, rely on the official settlement notice and the court's orders rather than press summaries; some deadlines may already have passed. This was one district court ruling on its own facts, and courts in other AI copyright cases can take different approaches; New York Times v. OpenAI is another case to follow.

What does the law actually say?

The case applied fair use, which courts decide case by case. The rules that matter most for companies licensing their own records concern ownership and promises:

  • Ownership starts with the author, and rights can be licensed separately. Under 17 U.S.C. § 201, copyright vests initially in the author; for a work made for hire, the employer is considered the author; and any of the exclusive rights can be transferred and owned separately. That is what lets a company license specific uses of material it owns while keeping everything else.
  • Employee work usually belongs to the employer. 17 U.S.C. § 101 defines a work made for hire as one prepared by an employee within the scope of employment, or certain commissioned works in listed categories where both parties sign a written agreement.
  • The Copyright Office has studied AI training. Its copyright and AI initiative released Part 3 of its report, on generative AI training, as a pre-publication version in May 2025. It addresses where copying in training may implicate copyright, how fair use may apply and how practical licensing is. It is a report, not law.
  • Privacy promises are enforceable. FTC staff wrote in January 2024 that companies' promises not to use customer data for undisclosed purposes, such as training models, can be enforced whether they appear in privacy policies, terms of service or marketing (FTC staff guidance, January 2024). That is staff guidance, not a rule.

Why does provenance now carry weight with buyers?

Provenance is the documented answer to three questions: where the data came from, who had the right to provide it, and what the people in it were told. What data provenance is, and why AI buyers pay for it covers the concept in depth.

A license from an operating company builds that record as it goes: an authorized sponsor signs, the scope is written down, de-identification and redaction rules are fixed before preparation starts, and nothing is delivered without an executed agreement and the company's authorization. Compared with scraped or downloaded material, that trail is the point. Buyers' wider expectations are set out in what ethically sourced AI training data means, and the kind of license matters too: training versus retrieval licenses explains the difference.

How does this apply in common partner situations?

SituationWhat to checkTypical outcome to confirm with counsel
Employees wrote the records: tickets, email, SOPs, codeScope of employment, policies, client contractsUsually the company's to license, subject to privacy and contract limits
Shared drives hold purchased books, industry reports, standards or vendor manualsPurchase or subscription termsUsually excluded: owning a copy is not the right to license it
The company holds documents for its clients, as an agency or outsourcerClient contracts and consentExcluded unless the clients agree
Freelancers or contractors produced some contentWritten assignment or work-made-for-hire termsIncluded only where rights were assigned
The privacy policy or customer terms limited data useThe exact promises and their datesAffected data excluded or consent obtained; terms not changed quietly
The data was already licensed for AI trainingExclusivity in the earlier agreementA red flag that may block a new license
Records were scraped from other websitesSource and site termsExcluded

What does good disclosure and consent practice look like?

  • Keep a written inventory: each system, its date range, who created the content and what is excluded.
  • Settle exclusions before anyone prepares data, not after a buyer asks.
  • Do not change privacy terms in order to license data without counsel's advice and proper notice.
  • Partners never export, upload or describe records; they make the introduction and give basic fit information.
  • If you recommend SourceX publicly while earning a referral reward, say so plainly.

Questions to ask your counsel

  1. Did the company create these records, or acquire them from someone else, and on what terms?
  2. Do client contracts, NDAs or master service agreements limit reuse of documents we hold?
  3. What did our privacy policy and customer terms say over the period the records cover?
  4. Do our contractor and freelancer agreements assign rights to us?
  5. Which folders hold third-party published works that should be excluded?
  6. Has any of this data been licensed before, and with what exclusivity?
  7. What de-identification do employee and customer details need?

This is general information, not legal, tax or financial advice. Confirm with your own counsel before acting.

What should referral partners take from the case?

The case does not change what a good referral looks like; it sharpens it. Companies whose records were created by their own people, under clean contracts and clear privacy terms, are the strongest candidates. When a founder asks why licensing is different from what they have read about, keep it simple:

The guide on how to explain company data licensing to a founder has more ways to frame it. Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company, payable only after the buyer pays and SourceX receives its fee; rewards are not guaranteed.

Next step

Run a candidate through the company fit checker. If it looks promising, register as a partner and introduce the company; how it works walks through qualification, inventory, pricing and payment.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Did the Bartz ruling decide that all AI training is fair use?

No. It was one district court's ruling on the facts of one case, and it distinguished lawfully acquired books from pirated copies. Courts in other cases can reach different conclusions, and new cases continue. Treat it as an important data point about acquisition and provenance, not as a general rule covering all training data.

Where can I find the settlement amount and claim details?

In the official settlement notice and the court's orders on the case docket, which are the authoritative sources for amounts, eligibility, deadlines and approval status. Press summaries vary, and some deadlines may already have passed. This page focuses on what the case means for companies licensing their own records rather than on claims administration.

Does the case affect companies licensing internal business records?

Not directly, because it concerned published books. Its relevance is the emphasis on how material was obtained. A company licensing records its own employees created, under a signed agreement with documented scope, sits at the opposite end from copies downloaded without permission, which is why buyers increasingly look for that kind of documented provenance.

Can a company license third-party books or reports it bought for its staff?

Generally not as part of a data license, because buying a copy does not usually carry the right to license it to others. Those materials are typically excluded during the rights review. The company and its counsel should check the purchase or subscription terms; SourceX focuses on records the company created itself.

Does SourceX work with the companies involved in this case?

This page names the case only as a public legal event. SourceX does not identify its buyers, and nothing here suggests that any party to the case is a SourceX buyer, client or partner. Companies licensing through SourceX deal with AI labs and data buyers under signed agreements that the company itself approves.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment