AI copyright is the question hanging over the entire generative AI industry, and in 2026 it still has no simple answer. Can a company feed millions of copyrighted books into a large language model without permission from the authors who wrote them? TechCrunch put that exact question to intellectual property lawyers this week in a widely shared explainer, and the honest professional answer was: it’s complicated. Courts have blessed some training runs, punished others, and produced the largest copyright settlement in American history along the way.
This article walks through where the AI copyright fight actually stands: the $1.5 billion Anthropic settlement and what it did and did not decide, the fair use doctrine every case turns on, the scoreboard of rulings so far, the very different position in the UK, and what all of it means for a business that builds on or simply uses AI models. The short version is that how the books were obtained now matters as much as what the model did with them — and, as the long-running debate over machine learning and intellectual property shows, that distinction is worth real money.
Table of contents
- The AI Copyright Question Everyone Is Asking
- AI Copyright and the $1.5 Billion Anthropic Settlement
- Fair Use: The Test at the Heart of Every AI Copyright Case
- AI Copyright Rulings So Far: The Scoreboard
- Why Judges Disagree: Reading Machines vs Market Rivals
- The AI Copyright Position in the UK
- What the US Copyright Office Says About AI Copyright
- What the AI Copyright Fight Means for Your Business
- AI Copyright FAQ
- References
The AI Copyright Question Everyone Is Asking
Strip away the technology and the AI copyright question is old-fashioned: when does using someone else’s creative work without permission become infringement? US copyright law has not been substantially rewritten since 1976 — decades before anyone imagined a model ingesting half a million novels — so judges are stretching pre-internet doctrine over a technology its drafters never saw.
What copyright actually protects
Copyright protects original expression: the sentences an author wrote, not the facts or ideas behind them. Infringement traditionally requires copying that expression. Training complicates this because a model does copy — wholesale, at ingestion — yet what it stores is statistical patterns, and what it outputs is usually something new. As IP lawyer Cathy Gellis told TechCrunch, “copyright law hinges on copying, but it doesn’t hinge on using the work.” Reading a book, learning from it, and writing something informed by it has never required a licence.
Why training sits in a grey zone
That is precisely the analogy AI companies lean on: training is reading at industrial scale. Rights holders answer that a machine ingesting 500,000 books in a week is nothing like a person reading them, and that the resulting systems compete for the same readers’ attention and money. Both arguments have now won in court, which is why every AI copyright headline seems to contradict the last one.
AI Copyright and the $1.5 Billion Anthropic Settlement
The defining AI copyright case so far is Bartz v. Anthropic, brought in 2024 by thriller novelist Andrea Bartz and two other authors over the books used to train Claude. In June 2025, Judge William Alsup handed down a split decision that has framed the debate ever since.
Training was fair use — piracy was not
Alsup ruled that training itself was lawful. The models, he wrote, trained on works “not to race ahead and replicate or supplant them — but to turn a hard corner and create something different.” That reading treats training as transformative fair use — closer to studying a library than photocopying one.
But Anthropic had also downloaded millions of books from pirate “shadow libraries,” and Alsup refused to excuse that acquisition. Facing a December 2025 trial on the piracy claims, with statutory damages of up to $150,000 per willful infringement on the table, Anthropic settled in late August 2025.
The largest copyright settlement in history
The numbers were unprecedented: $1.5 billion, roughly $3,000 per work across an estimated 500,000 books, later refined to a final list of about 482,000. Judge Alsup granted preliminary approval before retiring; Judge Araceli MartÃnez-OlguÃn gave final approval in July 2026. By then about 91% of the 482,000 covered books had been claimed by authors or publishers — roughly 439,000 works with cheques on the way, and about 43,000 still unclaimed.
What the settlement did not decide
Because it settled, Bartz never produced an appellate ruling on training. The fair use holding stands as one district judge’s view, not binding precedent. What the case really established is a price signal: acquire books lawfully and training may well be defensible; acquire them from pirate sources and the exposure is measured in billions. Every AI lab adjusted its data pipeline accordingly.
Fair Use: The Test at the Heart of Every AI Copyright Case
Every US AI copyright dispute funnels into the fair use doctrine — a four-factor balancing test that decides whether unauthorised use of a protected work is nevertheless lawful. The factors are deliberately open-ended, which is exactly why sophisticated lawyers keep answering “it’s complicated.”
| Fair use factor | What the court asks | How it has cut in training cases |
|---|---|---|
| Purpose and character | Is the use transformative — does it add new purpose or meaning? | Strongly pro-AI where outputs differ from the books (Anthropic, Meta); against Ross, whose product mirrored the source |
| Nature of the work | Is the source creative or factual? | Novels are highly creative, so this factor generally favours authors |
| Amount used | How much was copied, and was it reasonable for the purpose? | Whole books are copied, but courts have accepted that training needs full texts |
| Market effect | Does the use harm the market for the original? | The live battleground — includes the emerging “market dilution” theory |
Transformative use is the AI companies’ shield
Gellis argues Alsup’s framing was a gift to AI developers: comparing training to literary study protects the whole activity, because the law punishes copying that substitutes for the original, not learning from it. Jason Henderson, senior attorney at JWL International, drew the practical line for TechCrunch: courts favour fair use when the training does not directly compete with the works ingested, and scrutinise hard when the resulting product explicitly targets the same market.
Market harm is the authors’ sword
The fourth factor is where rights holders now concentrate their fire. Their argument: even if no single output copies a book, a market flooded with machine-generated fiction erodes the value of human-written fiction as a class. No court has yet accepted that dilution theory as a winning claim — but none has closed the door either, and it is the theory to watch in the pending cases.
AI Copyright Rulings So Far: The Scoreboard
Put the major AI copyright decisions side by side and the pattern is clearer than the headlines suggest. Courts keep asking two questions: was the data lawfully obtained, and does the product compete with the works it learned from?
| Case | Court and date | Outcome | Why it matters |
|---|---|---|---|
| Bartz v. Anthropic | N.D. California, June 2025; settlement approved July 2026 | Training fair use; $1.5bn settlement for pirated acquisition | Largest copyright settlement in history; made data provenance the key risk |
| Kadrey v. Meta | N.D. California, June 2025 | Summary judgment for Meta on training | Authors showed no market harm — but the judge flagged market dilution as a stronger future theory |
| Thomson Reuters v. Ross Intelligence | D. Delaware, February 2025 | Not fair use | Copying Westlaw content to build a competing legal research tool lacked “a further purpose or different character” |
| Thaler v. Perlmutter | D.C. Circuit, March 2025; cert denied March 2026 | Purely AI-generated works cannot be copyrighted | Locks in the human-authorship requirement for outputs |
| Getty Images v. Stability AI (UK) | High Court, November 2025 | Training claim failed; narrow trademark win on watermarks | Training outside the UK left UK copyright law with little to grip |
| NYT v. OpenAI and Microsoft | S.D. New York, in discovery | Motion to dismiss denied; trial expected late 2026 or 2027 | The bellwether — outputs allegedly compete directly with the source journalism |
The price of getting it wrong
US statutory damages explain why AI copyright cases settle enormous. The law allows $750 minimum per infringed work, up to $30,000 in the standard band, and up to $150,000 per work for willful infringement. Anthropic’s $3,000 per book sits at the low end of that scale — multiplied across 482,000 books it still reached $1.5 billion.
Why Judges Disagree: Reading Machines vs Market Rivals
Line up Alsup, Chhabria and Bibas and you get three thoughtful judges reaching different AI copyright conclusions from the same statute. That is not judicial chaos — it is the four factors doing their job on different facts.
The Anthropic view: models read, they do not replace
Alsup’s opinion treats a language model as the ultimate student: it ingests books to learn how language works, then produces something categorically different. On those facts the first factor dominates and training wins.
The Meta caveat: win on evidence, not on principle
Judge Vince Chhabria ruled for Meta in Kadrey — but pointedly, because the thirteen authors suing had failed to build a record of market harm. His opinion went out of its way to say a better-argued case, centred on generative flooding of the very market the books occupy, could come out the other way. It was an AI copyright victory for Meta and a roadmap for the next plaintiffs.
The Ross line: compete with your source and lose
Judge Stephanos Bibas found no fair use where Ross Intelligence used Thomson Reuters’ Westlaw material to build a directly competing legal research product. The use lacked “a further purpose or different character” — the clearest statement yet that AI copyright outcomes flip when the machine’s output substitutes for exactly what it ingested.
The AI Copyright Position in the UK
British law answers the AI copyright training question differently — mostly by not answering it. There is no general fair use doctrine here; the Copyright, Designs and Patents Act 1988 permits only narrow “fair dealing” exceptions, and its text and data mining exception covers non-commercial research alone. Commercial training on protected works in the UK therefore has no obvious statutory shelter.
Getty v. Stability: a hollow test case
The UK’s flagship AI copyright trial fizzled in November 2025. Getty Images lost its core copyright claim against Stability AI — largely because the training happened outside the UK, beyond the Act’s reach — and salvaged only a narrow trademark win over Getty watermarks appearing in generated images. The verdict left UK training law essentially untested for any company careful about where its GPUs sit.
The government blinks: status quo, for now
The UK government’s December 2024 consultation on copyright and AI had favoured a broad text and data mining exception with a rights-holder opt-out. After fierce opposition from the creative industries, the statutory report published on 18 March 2026 under the Data (Use and Access) Act 2025 abandoned that preference entirely. No new exception, no opt-out regime — the status quo holds while working groups on licensing and transparency grind on. For UK rights holders that is a defensive win; for UK AI developers it prolongs the uncertainty their American rivals are litigating their way out of.
What the US Copyright Office Says About AI Copyright
The closest thing to official US guidance arrived in May 2025, when the Copyright Office released Part 3 of its Copyright and Artificial Intelligence report, covering generative AI training. Its conclusion mirrors the case law: training “likely qualifies as fair use in some circumstances, but not in others” — context and degree, not categories.
A report with a body count
The report’s nuance was politically explosive. Days after its release, the White House fired Register of Copyrights Shira Perlmutter; a federal appeals court reinstated her in September 2025. The episode left the report’s formal status uncertain, but its analysis — lawful acquisition matters, competing outputs matter, licensing markets matter — tracks exactly where the courts have gone since.
Outputs are a separate question
One AI copyright issue is now genuinely settled: in March 2026 the Supreme Court declined to hear Thaler v. Perlmutter, leaving intact the rule that a work generated entirely by a machine gets no copyright at all. Human authorship remains the price of protection — which raises hard practical questions about how much human involvement is enough when AI assists a creative work, questions the Office is still answering registration by registration.
What the AI Copyright Fight Means for Your Business
Most businesses are not training foundation models on novels — but almost every business now uses tools built by companies that did. The AI copyright fight reaches you through the tools you buy, the content you generate, and the contracts you sign.
Know your exposure as a user
Using a mainstream model is far lower-risk than building one: the AI copyright liability for training sits with the vendor, and no court anywhere has held an end user liable simply for using a model trained on protected books. The practical risks are outputs that reproduce protected material and contracts that dump liability on you.
Check whether your vendor offers copyright indemnity for generated content — the serious providers now do — and keep humans meaningfully in the loop on anything you intend to protect, since purely machine-made work belongs to no one. Our guide to private AI for UK businesses covers the parallel question of what your own data feeds into these systems.
Questions to ask before you build
If you are fine-tuning or building on top of models — the territory covered in our Copilot vs custom AI assistant comparison — provenance is now the first question, not the last. The Anthropic settlement priced pirated data at $1.5 billion; licensed data is cheaper.
| If your business… | Your main AI copyright risk | What to do now |
|---|---|---|
| Uses AI tools for content and code | Outputs reproducing protected material; no protection for pure AI output | Vendor indemnity, human editing on anything you want to own, output review for anything public |
| Fine-tunes models on third-party content | Training-style infringement claims against you | Licence the corpus, document provenance, prefer your own data |
| Creates content others might train on | Your work ingested without payment | Register key works, assert reservations, watch the licensing schemes now emerging |
| Signs AI vendor contracts | Liability quietly shifted to you | Read the IP warranty and indemnity clauses before procurement, not after a claim |
Getting these questions answered early is part of any sensible adoption plan — our AI readiness assessment treats legal exposure as one of the six dimensions worth scoring before you spend, and our AI models and tools hub tracks which vendors stand behind their training data.
AI Copyright FAQ
So is it legal to train AI models on copyrighted books?
In the US, sometimes: the two rulings on point say training itself can be fair use when the books were obtained lawfully and the model does not regurgitate or directly compete with them — but pirated source libraries cost Anthropic $1.5 billion, and no appeals court has confirmed the fair use holding yet. In the UK there is no equivalent doctrine, and commercial-scale training on protected works has no clear legal basis.
Can I copyright what an AI writes for me?
Not if the machine did all the work. After Thaler, purely AI-generated output is uncopyrightable in the US. Add genuine human creativity — selection, arrangement, substantial editing — and the human-authored contribution can be protected.
Does the Anthropic settlement mean authors get paid for AI training now?
It means authors whose books were in Anthropic’s pirated libraries get about $3,000 per work. It set no licensing rate and no precedent for lawfully bought books — though it has pushed the industry towards licensed data deals to avoid the same exposure.
Which case should I watch next?
NYT v. OpenAI. It squarely tests the scenario the other rulings dodged: a model allegedly producing outputs that compete directly with the journalism it trained on. A verdict, expected from late 2026, will do more to settle the AI copyright question than everything decided so far.
References
TechCrunch: Is It Legal to Train AI Models on Copyrighted Books? It’s Complicated
TechCrunch: Anthropic’s Landmark $1.5B Copyright Settlement Is Approved
Fortune: Anthropic to Pay Authors $1.5 Billion Over Pirated Books Used to Train Claude
NBC News: AI Company Anthropic Agrees to Pay $1.5B to Settle Lawsuit With Authors
U.S. Copyright Office: Copyright and Artificial Intelligence
U.S. Copyright Office: Part 3 — Generative AI Training Report
The Authors Guild: US Copyright Office AI Report Part 3 — What Authors Should Know
GOV.UK: Copyright and Artificial Intelligence Consultation
GOV.UK: Report on Copyright and Artificial Intelligence (March 2026)
Norton Rose Fulbright: An Update on AI Copyright Cases in 2026
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.