Fair use is the two-word doctrine that decides whether the American AI industry ever has to pay for the books, articles and songs it learns from, and on Tuesday 1 September 2026 the United States government told a Manhattan federal court exactly which way it wants that question answered. In a 20-page statement of interest, the Department of Justice argued that copying copyrighted text in order to train a large language model is “extraordinarily transformative” and belongs squarely inside fair use — the statutory exception that lets one work build on another.

It is the first time the federal government has taken a formal position in any of the copyright lawsuits now aimed at AI developers. The filing landed in In re OpenAI, Inc. Copyright Infringement Litigation, the consolidated proceeding before Judge Sidney H. Stein in the Southern District of New York that now carries The New York Times’ case alongside fifteen others. The government did not ask to join the case as a party. It filed under 28 U.S.C. § 517, a one-sentence statute that lets the Attorney General send lawyers into any court to attend to the interests of the United States.

The reaction was immediate and unfriendly. A Times spokesperson said the administration was “siding with a handful of trillion-dollar AI companies at the expense of the countless American creators whose work they stole”. We covered the underlying legal question in August, when we asked whether it is legal to train AI models on copyrighted books. This piece is about what changed the day the government picked a side.

What follows is what the fair use brief actually argues, what it pointedly leaves alone, how it lines up against the rulings judges have already handed down, and what a business that builds on these models — or publishes into them — should do about it. Our AI models and tools hub tracks the wider releases as they land. Where sources disagree, including on the filing date itself, we say so rather than picking the tidier version.

What the DOJ Fair Use Brief Actually Says

fair use us government openai llm training b thread spool two flat flanges

The document is short by the standards of this litigation and blunt by the standards of government filings. It is a statement of interest, not an amicus brief and not a complaint in intervention, and it does one job: it tells Judge Stein how the executive branch reads section 107 of the Copyright Act as applied to model training.

The single sentence that matters

Strip away the framing and the brief reduces to one line: “The United States has a strong interest in this Court rejecting any argument that training LLMs on copyrighted texts violates copyright law.” Everything else in the filing is scaffolding for that sentence. The government is not asking the court to weigh a fair use defence carefully. It is asking the court to hold that the defence wins on this record.

Who signed it

The signatures matter because they show how high this went. Associate Attorney General Stanley E. Woodward Jr., Assistant Attorney General Brett Shumate and Senior Counsel Michael Weisbuch are on the filing. A fair use argument does not normally attract that seniority. Woodward framed the stakes publicly in terms that have nothing to do with copyright doctrine: “AI dominance is critical to promote national security, prosperity, and economic mobility for all Americans.”

The national-interest wrapper

The brief spends real space on competitiveness rather than statutory analysis. The United States, it says, “has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally”, and it is critical for the country to “retain global leadership in artificial intelligence”. The closing move is the sharpest: constraining model development “under a misunderstanding of fair use doctrine would thwart such creative and scientific progress while hindering American prosperity and economic mobility”.

What a statement of interest is not

None of this binds anyone. A statement of interest is persuasive advocacy filed by a non-party. It does not amend 17 U.S.C. § 107, it does not create a safe harbour for OpenAI, and Judge Stein is free to read the same four factors and reach the opposite conclusion. Federal judges have been notably unimpressed by executive-branch positions on copyright before, and the fair use analysis remains, by statute, a case-by-case exercise.

The Four Fair Use Factors, as the Government Reads Them

fair use us government openai llm training c stacked desk tray two shelves

Section 107 lists the four factors every fair use analysis has to work through. The brief leans almost entirely on the first and the fourth, treats the third as a technical question, and barely engages the second at all.

Factor one: the purpose is computational, not expressive

The government’s core move is to separate what a work is for from what a model does with it. A news article exists to communicate its content to a reader. A training run does not read the article; it derives statistical relationships across an enormous corpus, which is a different purpose and a different character of use. That is the basis for calling the copying “extraordinarily transformative”, and it is also why the brief says commercial motive “does not move the needle” — a transformative purpose has long carried more weight than a profit motive in fair use case law.

Factor three: copying the whole thing can still be reasonable

Whole-work copying normally cuts against a fair use defence, because taking everything is hard to justify. The brief argues that the amount used is still reasonable where the transformative technical function actually requires the full work and where the copy itself is never exposed to the public. A model cannot learn the shape of a sentence from a paragraph of a novel; it needs the novel. That is an engineering claim doing legal work, and it is one of the places the plaintiffs are most likely to push back.

Factor four: harm means substitution, not competition

This is the most consequential passage in the whole fair use argument. The government says the market harm factor is about substitutive competition in protected expression — an output that stands in for the original — and not about general economic pressure from machine-made work. It goes directly at the market dilution theory, the argument that flooding a market with cheap synthetic articles harms the market for real ones even when no single output copies anything. Merging training with output competition, the brief says, is a category error.

Statutory factorWhat the government arguesWhat the publishers argue
1. Purpose and characterExtraordinarily transformative; commercial motive “does not move the needle”Nothing transformative about ingesting journalism to build a rival to journalism
2. Nature of the workBarely addressed in the filingOriginal reporting sits at the core of what copyright protects
3. Amount and substantialityWhole-work copying is reasonable where the technical function needs it and the copy stays privateEvery article was taken in full, and some can be reproduced from the model
4. Effect on the marketOnly substitutive competition in protected expression counts; dilution theories failThe products substitute for the originals and destroy the licensing market

Where the Fair Use Argument Stops: Acquisition and Output

fair use us government openai llm training d kettle body spout and handle

The most useful thing about the brief is what it refuses to say. Read carefully, it draws three lines that most of the coverage collapsed into one, and every one of those lines is a place where a fair use defence can still fail.

Acquisition is a separate question

The government keeps how a copy was obtained legally distinct from what was later done with it, and fair use only ever answers the second question. An unlawfully obtained copy can generate liability even where the subsequent computational use is transformative. That is precisely the distinction that cost Anthropic $1.5bn: a judge accepted that training was fair use and still refused to excuse the shadow-library downloads that supplied the books. Nothing in this filing rescues a company that torrented its corpus.

The brief disclaims any government blessing

There is a paragraph that reads like it was written by someone anticipating the headlines. The United States expressly disclaims any contention that the alleged activities were authorised by, consented to, or undertaken for the benefit of the government. Washington is arguing a doctrine, not certifying OpenAI’s conduct, and the distinction will matter if any defendant tries to quote the filing as a character reference.

Outputs are still fair game for plaintiffs

Training is one thing; what comes out of the model is another. The filing’s logic makes memorisation and regurgitation the live battleground, because an output that reproduces protected expression is exactly the substitutive harm the government concedes would count. Anti-memorisation controls stop being a product-safety feature at that point and become evidence. If your model can be made to recite the plaintiff’s article, the fair use argument the government just made does not help you.

Why that changes the shape of the case

Put the three lines together and the fair use claim is narrower than the headline suggests. It supports one proposition — that the act of training on lawfully held text is fair use — while leaving acquisition, output and licensing entirely open. The New York Times’ strongest exhibits have always been regurgitation examples, and the government’s filing does not touch them.

How the Times, Authors and Labels Answered the Fair Use Brief

fair use us government openai llm training e weight block with ring handle

The responses arrived within hours, and they split along predictable lines.

The New York Times

The Times attacked the politics and the doctrine in the same breath. On the politics, the administration was “siding with a handful of trillion-dollar AI companies at the expense of the countless American creators whose work they stole”, and AI companies should “pay fairly for the content that makes their products possible”. On the doctrine, the paper’s position is unchanged: there is “nothing ‘transformative’ about using The Times’s content without payment to create products that substitute for The Times”.

The authors in the class

The plaintiffs consolidated into this proceeding are not only newspapers. The Authors Guild is there, and so are George R.R. Martin, John Grisham, Michael Connelly, Jonathan Franzen and Jodi Picoult, alongside the Daily News, the Chicago Tribune and the Center for Investigative Reporting. For that group the government’s brief reads as an argument that the largest unlicensed copying operation in publishing history was lawful.

The music industry’s read

Music rightsholders took it worst, and with reason. Their lawsuits against Suno, Udio and others rest heavily on how training material was acquired, not merely on whether a fair use defence covers the training itself. Trade coverage described the filing as a dispiriting development and noted that years of lobbying had “fallen on deaf ears — at least in the White House”. The counter-argument is that the acquisition carve-out above leaves those cases largely intact.

OpenAI’s own framing

OpenAI’s position has not moved: “Copyright is not a veto right over transformative technologies that leverage existing works internally.” The company did not immediately comment on the government’s filing when Reuters asked, which is unsurprising — when the Department of Justice files your argument for you, saying less is the better play.

The Case Behind the Brief: In re OpenAI, MDL 3143

fair use us government openai llm training f pencil hexagonal shaft and point

The filing does not exist in a vacuum. It arrived at a specific and quite late moment in a case that has been running for nearly three years.

How sixteen lawsuits became one

The New York Times sued OpenAI and Microsoft in December 2023. On 26 March 2025 Judge Stein denied the motions to dismiss in large part, letting the central infringement claims proceed while narrowing several Digital Millennium Copyright Act claims. A week later, on 3 April 2025, the Judicial Panel on Multidistrict Litigation transferred the scattered cases to the Southern District of New York as MDL No. 3143, twelve actions at the time and sixteen since, with Judge Stein presiding and Magistrate Judge Ona T. Wang handling discovery.

The twenty million chat logs

The discovery fight that defines the case is about outputs. OpenAI proposed in October 2025 to produce only logs surfaced by keyword searches. Magistrate Judge Wang rejected that in November 2025, and on 5 January 2026 Judge Stein affirmed: OpenAI must hand over the full sample of 20 million de-identified ChatGPT conversations. Stein’s reasoning is the mirror image of the government’s fourth-factor argument — the logs “could reveal patterns relevant to whether ChatGPT’s outputs compete with or substitute for copyrighted works”.

What happens next

Summary judgment motions are under consideration, and outside parties may file supporting briefs until 16 October 2026. Expect a queue: rightsholder coalitions, technology trade bodies, law professors and probably several state attorneys general. The government’s filing has effectively opened the amicus season on the biggest fair use question of the decade, and every serious view of fair use will now be argued in front of one judge.

DateMilestone
December 2023The New York Times sues OpenAI and Microsoft in the Southern District of New York
13 March 2025OpenAI asks the White House policy office to protect model training under section 107
26 March 2025Judge Stein denies the motions to dismiss in large part
3 April 2025The multidistrict panel consolidates the cases as MDL No. 3143
23 July 2025America’s AI Action Plan is published; the President rejects paying for training material
5 January 2026Judge Stein orders production of 20 million de-identified chat logs
20 July 2026Final approval of Anthropic’s $1.5bn settlement with authors
1 September 2026The Department of Justice files its 20-page statement of interest
16 October 2026Deadline for outside parties to file supporting briefs

Measured from the December 2023 complaint, the government’s intervention arrives 33 months in — almost at the end of the pre-trial road rather than the start of it.

Months elapsed since the December 2023 complaint, relative to the 16 October 2026 brief deadline (34 months = 100%)
Deadline for supporting briefs, 34 months 100%
Justice Department statement of interest, 33 months 97%
Order to produce 20 million chat logs, 25 months 74%
Consolidation into one proceeding, 16 months 47%
Ruling on the motions to dismiss, 15 months 44%

What Judges Have Actually Held on Fair Use So Far

The government is not writing on a blank page. Four decisions already shape how a court is likely to read this filing, and the fair use answers in them do not point the same way.

Bartz v. Anthropic split the question in two

In June 2025 Judge William Alsup held that training was protected by fair use, describing models trained “not to race ahead and replicate or supplant them — but to turn a hard corner and create something different”. He then refused to extend that to the pirated libraries Anthropic had downloaded. The company settled for $1.5bn — roughly $3,000 per work across about 500,000 books — and Judge Araceli Martínez-Olguín granted final approval on 20 July 2026, with 350 authors opting out and the release limited to conduct through 25 August 2025.

Kadrey v. Meta went the other way on the record

Also in June 2025, Judge Vince Chhabria granted summary judgment to Meta on the training question, but the reasoning was narrower than the result. He found the plaintiffs had not built a market-harm record, and he flagged market dilution as a theory that might well succeed if someone actually proved it. That is the theory the government’s brief now attacks head-on.

Thomson Reuters v. Ross rejected the defence outright

In February 2025 Judge Stephanos Bibas held that Ross Intelligence’s use of Westlaw headnotes was not fair use, because it lacked “a further purpose or different character”. It is the clearest American decision refusing a fair use defence for machine training, and it is not a large language model case — which is exactly how each side will characterise it.

Getty v. Stability shows the geography problem

In November 2025 the English High Court rejected Getty’s main copyright claim largely because the training happened outside the United Kingdom, while allowing a narrow trademark point. No fair use question arose at all, because fair use is an American doctrine. The European Union and the United Kingdom run text-and-data-mining exceptions with different shapes, so a win for OpenAI in Manhattan settles nothing in London or Brussels.

CaseDecidedHolding on training
Thomson Reuters v. RossFebruary 2025Not protected — no further purpose or different character
Bartz v. AnthropicJune 2025Protected, but pirated source copies were not excused
Kadrey v. MetaJune 2025Protected on this record; dilution theory left open
Getty v. Stability (UK)November 2025Claim failed on territory, not on doctrine
In re OpenAI (MDL 3143)PendingSummary judgment under consideration

The money is what makes the split dangerous. Statutory damages run from $750 to $30,000 per work, and up to $150,000 where infringement is wilful, which is how a training corpus turns into an existential number.

Damages per infringed work, relative to the wilful statutory maximum ($150,000 = 100%)
Statutory maximum, wilful infringement, $150,000 100%
Statutory maximum, ordinary infringement, $30,000 20%
Anthropic settlement, about $3,000 2%
Statutory minimum, $750 0.5%

Why Washington Picked This Moment

Nothing about the timing is accidental, and the policy trail behind it is public.

OpenAI asked for this in writing

On 13 March 2025 OpenAI filed a submission to the White House science and technology policy office asking for exactly this outcome — a policy environment where models may learn from copyrighted work, framed almost entirely as a race. “While America maintains a lead on AI today,” the company wrote, “DeepSeek shows that our lead is not wide and is narrowing.” Eighteen months later the Department of Justice is making that argument to a federal judge.

The AI Action Plan set the direction

America’s AI Action Plan, published on 23 July 2025, ran to 25 pages and roughly 90 federal policy actions across three pillars, accompanied by three executive orders. Copyright was not one of its formal deliverables, but the President addressed it from the stage the same day: “You can’t be expected to have a successful AI program where every single article, book or anything else that you’ve read or studied, you’re supposed to pay for.” He added that “China’s not doing it”.

The diplomatic track

The position is being exported as well as litigated. Commerce Secretary Howard Lutnick has pressed G20 officials to embrace a fair use approach while claiming to protect artists — the same posture we described when the United States pushed the G20 towards light-touch AI rules. This filing is the domestic half of a coordinated argument.

The one part of government that disagreed

The US Copyright Office is not aligned with this. Its May 2025 pre-publication report on generative training concluded that the answer is fair use in some circumstances but not in others — the case-by-case reading of fair use that the brief is trying to short-circuit. The Register of Copyrights was removed from her post days after that report appeared, then reinstated. The government does not speak with one voice here, and a court will notice.

What the Fair Use Brief Changes for Your Business

For almost everyone reading this, the honest answer is that the legal position has not changed at all. What has changed is the direction of travel, and a few things are worth doing about it now.

If you build products on foundation models

Nothing about your indemnity position improved this week, because a fair use argument filed by a non-party does not travel down your supply chain. Check what your model provider actually promises about training-data claims, and read the carve-outs — most cover outputs, not the corpus. The brief’s own acquisition line is the tell: a provider that cannot say where its training data came from is carrying a risk that no statement of interest removes.

If you fine-tune on your own corpus

The government’s fair use argument covers the act of learning from material you hold lawfully. It says nothing about material you scraped, bought from a broker with a thin provenance trail, or inherited in an acquisition. Keep an auditable record of every source and its licence terms. That record is cheap now and very expensive to reconstruct later, and it is what separates the Anthropic training ruling from the Anthropic settlement.

If you publish content

Assume the crawler question is now separate from the copyright question. Rightsholders who want control are increasingly getting it through access rules rather than lawsuits, which is the pattern we traced when AI crawlers started eating website traffic. Publish clear terms, set your robots and firewall policy deliberately, and treat licensing conversations as a commercial matter rather than a legal remedy.

If you are buying an AI vendor

Add three questions to diligence. Where did the training corpus come from, what anti-memorisation testing exists and what did it show, and is the company a defendant anywhere. The same provenance discipline that the Debian project chose over an outright ban on AI-generated code works here: you are not trying to prove purity, you are trying to prove you asked.

If you areWhat the brief changesWhat to do this quarter
Building on a hosted modelNothing legally; sentiment onlyRe-read the indemnity and its carve-outs
Training or tuning your ownHelps the training step, not acquisitionBuild a source-and-licence register
A publisher or broadcasterWeakens the dilution argumentMove control to access terms and crawler policy
A rightsholder considering suitRaises the value of acquisition and output claimsCollect regurgitation evidence, not just ingestion evidence
Procuring an AI vendorMakes provenance the differentiatorAsk for corpus provenance and memorisation test results

What We Could Not Verify

Several details in the coverage do not reconcile, and it is worth naming them rather than smoothing them over.

The filing date

Reuters reported the brief was filed on a Tuesday, which was 1 September 2026, and legal write-ups give the same date. TechCrunch’s story is dated 2 September. We have used 1 September for the filing and treated 2 September as the day the story broke, but we have not seen the docket stamp ourselves.

The page count and the exact wording

The 20-page length comes from TechCrunch. The quotations above are reproduced as they appear in press reports and legal summaries; different outlets render one phrase as “extraordinarily transformative” and another as “exceedingly transformative”, which suggests at least one is a paraphrase. We have not read the filing itself line by line.

Whether this changes any judge’s mind

There is no evidence either way yet. Judge Stein has ruled against OpenAI on discovery already, and district judges are under no obligation to defer to the executive branch on how fair use should be read. Anyone telling you this filing decides the case is guessing.

Fair Use and AI Training: Common Questions

Did the US government just make AI training legal?

No. It filed a non-binding statement of interest arguing that training a model on copyrighted text should qualify as fair use. Only the court can decide whether fair use applies here, and the statute itself is unchanged.

What exactly did the Department of Justice argue?

That the copying is “extraordinarily transformative” because the purpose is computational rather than expressive, that commercial motive “does not move the needle”, and that market harm means substitution in protected expression rather than economic competition from generated work.

Does the brief protect companies that used pirated material?

No, and it says so. It keeps acquisition separate from training, and it accepts that an unlawfully obtained copy can create liability even where the later use is transformative. That is the distinction behind Anthropic’s $1.5bn settlement.

Which case is this?

In re OpenAI, Inc. Copyright Infringement Litigation, MDL No. 3143, in the Southern District of New York, before Judge Sidney H. Stein with Magistrate Judge Ona T. Wang. Sixteen lawsuits are consolidated in it, including The New York Times’ 2023 complaint.

How did The New York Times respond?

By accusing the administration of siding with trillion-dollar AI companies at the expense of American creators, and by restating that there is nothing transformative about using its content without payment to build products that substitute for it.

Does this affect the music industry lawsuits?

Indirectly. The fair use reasoning helps every defendant, but the music cases turn heavily on how training material was acquired, which the brief expressly leaves alone. Two of the three major labels are still litigating.

Does any of this apply outside the United States?

No. Fair use is an American doctrine. The United Kingdom and the European Union rely on text-and-data-mining exceptions with different conditions, which is why Getty’s claim against Stability failed on territorial grounds rather than on the merits of training.

What should a business do differently now?

Very little legally, and one thing operationally: document the provenance of every corpus you train or tune on, and test your models for memorisation. Those two records are what a fair use defence is built from, and neither gets easier to assemble after a claim arrives.

References