AI copyright is the question hanging over the entire generative AI industry, and in 2026 it still has no simple answer. Can a company feed millions of copyrighted books into a large language model without permission from the authors who wrote them? TechCrunch put that exact question to intellectual property lawyers this week in a widely shared explainer, and the honest professional answer was: it’s complicated. Courts have blessed some training runs, punished others, and produced the largest copyright settlement in American history along the way.

This article walks through where the AI copyright fight actually stands: the $1.5 billion Anthropic settlement and what it did and did not decide, the fair use doctrine every case turns on, the scoreboard of rulings so far, the very different position in the UK, and what all of it means for a business that builds on or simply uses AI models. The short version is that how the books were obtained now matters as much as what the model did with them — and, as the long-running debate over machine learning and intellectual property shows, that distinction is worth real money.

ai copyright training models on copyrighted books b courthouse columns

Strip away the technology and the AI copyright question is old-fashioned: when does using someone else’s creative work without permission become infringement? US copyright law has not been substantially rewritten since 1976 — decades before anyone imagined a model ingesting half a million novels — so judges are stretching pre-internet doctrine over a technology its drafters never saw.

What copyright actually protects

Copyright protects original expression: the sentences an author wrote, not the facts or ideas behind them. Infringement traditionally requires copying that expression. Training complicates this because a model does copy — wholesale, at ingestion — yet what it stores is statistical patterns, and what it outputs is usually something new. As IP lawyer Cathy Gellis told TechCrunch, “copyright law hinges on copying, but it doesn’t hinge on using the work.” Reading a book, learning from it, and writing something informed by it has never required a licence.

Why training sits in a grey zone

That is precisely the analogy AI companies lean on: training is reading at industrial scale. Rights holders answer that a machine ingesting 500,000 books in a week is nothing like a person reading them, and that the resulting systems compete for the same readers’ attention and money. Both arguments have now won in court, which is why every AI copyright headline seems to contradict the last one.

ai copyright training models on copyrighted books c open book blank pages

The defining AI copyright case so far is Bartz v. Anthropic, brought in 2024 by thriller novelist Andrea Bartz and two other authors over the books used to train Claude. In June 2025, Judge William Alsup handed down a split decision that has framed the debate ever since.

Training was fair use — piracy was not

Alsup ruled that training itself was lawful. The models, he wrote, trained on works “not to race ahead and replicate or supplant them — but to turn a hard corner and create something different.” That reading treats training as transformative fair use — closer to studying a library than photocopying one.

But Anthropic had also downloaded millions of books from pirate “shadow libraries,” and Alsup refused to excuse that acquisition. Facing a December 2025 trial on the piracy claims, with statutory damages of up to $150,000 per willful infringement on the table, Anthropic settled in late August 2025.

The largest copyright settlement in history

The numbers were unprecedented: $1.5 billion, roughly $3,000 per work across an estimated 500,000 books, later refined to a final list of about 482,000. Judge Alsup granted preliminary approval before retiring; Judge Araceli Martínez-Olguín gave final approval in July 2026. By then about 91% of the 482,000 covered books had been claimed by authors or publishers — roughly 439,000 works with cheques on the way, and about 43,000 still unclaimed.

Anthropic settlement: share of the 482,000 covered books claimed by mid-2026
Claimed by authors or publishers 91%
Still unclaimed 9%

What the settlement did not decide

Because it settled, Bartz never produced an appellate ruling on training. The fair use holding stands as one district judge’s view, not binding precedent. What the case really established is a price signal: acquire books lawfully and training may well be defensible; acquire them from pirate sources and the exposure is measured in billions. Every AI lab adjusted its data pipeline accordingly.

ai copyright training models on copyrighted books d fountain pen upright

Every US AI copyright dispute funnels into the fair use doctrine — a four-factor balancing test that decides whether unauthorised use of a protected work is nevertheless lawful. The factors are deliberately open-ended, which is exactly why sophisticated lawyers keep answering “it’s complicated.”

Fair use factorWhat the court asksHow it has cut in training cases
Purpose and characterIs the use transformative — does it add new purpose or meaning?Strongly pro-AI where outputs differ from the books (Anthropic, Meta); against Ross, whose product mirrored the source
Nature of the workIs the source creative or factual?Novels are highly creative, so this factor generally favours authors
Amount usedHow much was copied, and was it reasonable for the purpose?Whole books are copied, but courts have accepted that training needs full texts
Market effectDoes the use harm the market for the original?The live battleground — includes the emerging “market dilution” theory

Transformative use is the AI companies’ shield

Gellis argues Alsup’s framing was a gift to AI developers: comparing training to literary study protects the whole activity, because the law punishes copying that substitutes for the original, not learning from it. Jason Henderson, senior attorney at JWL International, drew the practical line for TechCrunch: courts favour fair use when the training does not directly compete with the works ingested, and scrutinise hard when the resulting product explicitly targets the same market.

Market harm is the authors’ sword

The fourth factor is where rights holders now concentrate their fire. Their argument: even if no single output copies a book, a market flooded with machine-generated fiction erodes the value of human-written fiction as a class. No court has yet accepted that dilution theory as a winning claim — but none has closed the door either, and it is the theory to watch in the pending cases.

ai copyright training models on copyrighted books e two blocks gap

Put the major AI copyright decisions side by side and the pattern is clearer than the headlines suggest. Courts keep asking two questions: was the data lawfully obtained, and does the product compete with the works it learned from?

CaseCourt and dateOutcomeWhy it matters
Bartz v. AnthropicN.D. California, June 2025; settlement approved July 2026Training fair use; $1.5bn settlement for pirated acquisitionLargest copyright settlement in history; made data provenance the key risk
Kadrey v. MetaN.D. California, June 2025Summary judgment for Meta on trainingAuthors showed no market harm — but the judge flagged market dilution as a stronger future theory
Thomson Reuters v. Ross IntelligenceD. Delaware, February 2025Not fair useCopying Westlaw content to build a competing legal research tool lacked “a further purpose or different character”
Thaler v. PerlmutterD.C. Circuit, March 2025; cert denied March 2026Purely AI-generated works cannot be copyrightedLocks in the human-authorship requirement for outputs
Getty Images v. Stability AI (UK)High Court, November 2025Training claim failed; narrow trademark win on watermarksTraining outside the UK left UK copyright law with little to grip
NYT v. OpenAI and MicrosoftS.D. New York, in discoveryMotion to dismiss denied; trial expected late 2026 or 2027The bellwether — outputs allegedly compete directly with the source journalism

The price of getting it wrong

US statutory damages explain why AI copyright cases settle enormous. The law allows $750 minimum per infringed work, up to $30,000 in the standard band, and up to $150,000 per work for willful infringement. Anthropic’s $3,000 per book sits at the low end of that scale — multiplied across 482,000 books it still reached $1.5 billion.

US statutory damages per work vs the Anthropic per-book payout
Statutory minimum $750
Anthropic settlement per book $3,000
Standard maximum $30,000
Willful maximum $150,000

Why Judges Disagree: Reading Machines vs Market Rivals

ai copyright training models on copyrighted books f magnifying glass

Line up Alsup, Chhabria and Bibas and you get three thoughtful judges reaching different AI copyright conclusions from the same statute. That is not judicial chaos — it is the four factors doing their job on different facts.

The Anthropic view: models read, they do not replace

Alsup’s opinion treats a language model as the ultimate student: it ingests books to learn how language works, then produces something categorically different. On those facts the first factor dominates and training wins.

The Meta caveat: win on evidence, not on principle

Judge Vince Chhabria ruled for Meta in Kadrey — but pointedly, because the thirteen authors suing had failed to build a record of market harm. His opinion went out of its way to say a better-argued case, centred on generative flooding of the very market the books occupy, could come out the other way. It was an AI copyright victory for Meta and a roadmap for the next plaintiffs.

The Ross line: compete with your source and lose

Judge Stephanos Bibas found no fair use where Ross Intelligence used Thomson Reuters’ Westlaw material to build a directly competing legal research product. The use lacked “a further purpose or different character” — the clearest statement yet that AI copyright outcomes flip when the machine’s output substitutes for exactly what it ingested.

British law answers the AI copyright training question differently — mostly by not answering it. There is no general fair use doctrine here; the Copyright, Designs and Patents Act 1988 permits only narrow “fair dealing” exceptions, and its text and data mining exception covers non-commercial research alone. Commercial training on protected works in the UK therefore has no obvious statutory shelter.

Getty v. Stability: a hollow test case

The UK’s flagship AI copyright trial fizzled in November 2025. Getty Images lost its core copyright claim against Stability AI — largely because the training happened outside the UK, beyond the Act’s reach — and salvaged only a narrow trademark win over Getty watermarks appearing in generated images. The verdict left UK training law essentially untested for any company careful about where its GPUs sit.

The government blinks: status quo, for now

The UK government’s December 2024 consultation on copyright and AI had favoured a broad text and data mining exception with a rights-holder opt-out. After fierce opposition from the creative industries, the statutory report published on 18 March 2026 under the Data (Use and Access) Act 2025 abandoned that preference entirely. No new exception, no opt-out regime — the status quo holds while working groups on licensing and transparency grind on. For UK rights holders that is a defensive win; for UK AI developers it prolongs the uncertainty their American rivals are litigating their way out of.

The closest thing to official US guidance arrived in May 2025, when the Copyright Office released Part 3 of its Copyright and Artificial Intelligence report, covering generative AI training. Its conclusion mirrors the case law: training “likely qualifies as fair use in some circumstances, but not in others” — context and degree, not categories.

A report with a body count

The report’s nuance was politically explosive. Days after its release, the White House fired Register of Copyrights Shira Perlmutter; a federal appeals court reinstated her in September 2025. The episode left the report’s formal status uncertain, but its analysis — lawful acquisition matters, competing outputs matter, licensing markets matter — tracks exactly where the courts have gone since.

Outputs are a separate question

One AI copyright issue is now genuinely settled: in March 2026 the Supreme Court declined to hear Thaler v. Perlmutter, leaving intact the rule that a work generated entirely by a machine gets no copyright at all. Human authorship remains the price of protection — which raises hard practical questions about how much human involvement is enough when AI assists a creative work, questions the Office is still answering registration by registration.

Most businesses are not training foundation models on novels — but almost every business now uses tools built by companies that did. The AI copyright fight reaches you through the tools you buy, the content you generate, and the contracts you sign.

Know your exposure as a user

Using a mainstream model is far lower-risk than building one: the AI copyright liability for training sits with the vendor, and no court anywhere has held an end user liable simply for using a model trained on protected books. The practical risks are outputs that reproduce protected material and contracts that dump liability on you.

Check whether your vendor offers copyright indemnity for generated content — the serious providers now do — and keep humans meaningfully in the loop on anything you intend to protect, since purely machine-made work belongs to no one. Our guide to private AI for UK businesses covers the parallel question of what your own data feeds into these systems.

Questions to ask before you build

If you are fine-tuning or building on top of models — the territory covered in our Copilot vs custom AI assistant comparison — provenance is now the first question, not the last. The Anthropic settlement priced pirated data at $1.5 billion; licensed data is cheaper.

If your business…Your main AI copyright riskWhat to do now
Uses AI tools for content and codeOutputs reproducing protected material; no protection for pure AI outputVendor indemnity, human editing on anything you want to own, output review for anything public
Fine-tunes models on third-party contentTraining-style infringement claims against youLicence the corpus, document provenance, prefer your own data
Creates content others might train onYour work ingested without paymentRegister key works, assert reservations, watch the licensing schemes now emerging
Signs AI vendor contractsLiability quietly shifted to youRead the IP warranty and indemnity clauses before procurement, not after a claim

Getting these questions answered early is part of any sensible adoption plan — our AI readiness assessment treats legal exposure as one of the six dimensions worth scoring before you spend, and our AI models and tools hub tracks which vendors stand behind their training data.

So is it legal to train AI models on copyrighted books?

In the US, sometimes: the two rulings on point say training itself can be fair use when the books were obtained lawfully and the model does not regurgitate or directly compete with them — but pirated source libraries cost Anthropic $1.5 billion, and no appeals court has confirmed the fair use holding yet. In the UK there is no equivalent doctrine, and commercial-scale training on protected works has no clear legal basis.

Can I copyright what an AI writes for me?

Not if the machine did all the work. After Thaler, purely AI-generated output is uncopyrightable in the US. Add genuine human creativity — selection, arrangement, substantial editing — and the human-authored contribution can be protected.

Does the Anthropic settlement mean authors get paid for AI training now?

It means authors whose books were in Anthropic’s pirated libraries get about $3,000 per work. It set no licensing rate and no precedent for lawfully bought books — though it has pushed the industry towards licensed data deals to avoid the same exposure.

Which case should I watch next?

NYT v. OpenAI. It squarely tests the scenario the other rulings dodged: a model allegedly producing outputs that compete directly with the journalism it trained on. A verdict, expected from late 2026, will do more to settle the AI copyright question than everything decided so far.

References