NYT lawsuit filings made public on 17 September 2026 put a Microsoft scientist’s words at the centre of the case against Microsoft and OpenAI. The New York Times and four other news plaintiffs quote Microsoft’s Director of Applied Science describing AI training on people’s work as “an astonishing theft of unprecedented proportions”, and perhaps the “largest theft of labor in human history”. The headline version, carried by AFP and others, is that the two companies “knew” using news content was theft.

The filing is a 92-page public, redacted version of the News Plaintiffs’ combined summary judgment brief in the consolidated copyright litigation before Judge Sidney H. Stein in Manhattan. It was re-filed with some redactions removed under an omnibus sealing order, which is why material from depositions and internal documents about the training data behind ChatGPT is now readable for the first time.

We read the brief itself, alongside Microsoft’s own summary judgment memorandum filed the same day. The quotes are real. What they say, in context, is more specific than the headlines, and in one important case it is a forecast of how the public would see AI training rather than an admission of what the company believed. This article sets out what the NYT lawsuit brief claims, the evidence it cites, and how Microsoft answers it.

What the NYT Lawsuit Filing Actually Is

nyt lawsuit microsoft openai knew news content theft b paper shredder box with one wide top slot

Start with the document, because the procedural details explain what can and cannot be concluded from it.

A summary judgment brief, not a verdict

The News Plaintiffs are asking Judge Stein to decide parts of the NYT lawsuit without a trial, on the ground that the facts are undisputed. A brief is advocacy. Every quote in it was chosen by the plaintiffs’ lawyers, and each is keyed to a paragraph of their statement of undisputed facts, cited as “SF” followed by a number.

Who the News Plaintiffs are

The brief is filed for The New York Times Company, the Daily News plaintiffs (eight papers including the New York Daily News, Chicago Tribune and Denver Post), the Center for Investigative Reporting, The Intercept and Ziff Davis. It relates to five member cases of the multidistrict litigation, No. 25-md-3143, including the Times’ original case, No. 1:23-cv-11195.

Why it is readable now

A cover letter from Susman Godfrey’s Davida Brook, dated 17 September, explains that the attached public version removes redactions “for material that no party or third-party has sought to seal”, under paragraph 9.c of a stipulated omnibus sealing order entered on 3 September. Some passages remain blacked out, including several dollar amounts.

What it asks for

Summary judgment of liability at five stages of what it calls the AI pipeline; partial summary judgment that OpenAI intentionally removed copyright management information under the DMCA; a ruling that separate statutory damages are available for each article; and judgment against several affirmative defences.

The "Astonishing Theft" Quote in the NYT Lawsuit, in Context

nyt lawsuit microsoft openai knew news content theft c canister vacuum body on two round wheels

The phrase that led every news report comes from a Microsoft document, and its full sentence matters.

The sentence the brief quotes

In its background section the brief quotes Microsoft as recognising that “millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions”, and admitting that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”

What that sentence is, and is not

Read in full, it is a prediction about how millions of people will regard AI training. It is not, on its face, a statement that the author or Microsoft considered the training to be theft. The brief’s introduction compresses it to “as Microsoft’s Director of Applied Science put it, ‘an astonishing theft of unprecedented proportions'”, and AFP’s report compressed it further into the claim that Microsoft “knew” it was theft.

Who wrote it

AFP names the scientist as “Brent Hect”. The brief spells it Dr. Brent Hecht, Microsoft’s Director of Applied Science, and the authors’ class brief filed the same day uses the same spelling. Microsoft told AFP that his statements reflect “one employee’s individual perspective” and “do not represent the company’s views”.

Why the plaintiffs still lean on it

The NYT lawsuit brief does not rest on the prediction alone. It pairs it with Hecht’s warning that a successful fair use defence would arguably “make a complete mockery of the idea of ‘fair use'”, and with a second document, quoted below, describing a “doom loop”. Together they are offered as evidence that Microsoft understood the market harm, which is what the fourth fair use factor turns on.

Quote as reportedSpeaker per the briefWhat the full passage says
“an astonishing theft of unprecedented proportions”Microsoft document (Hecht, per the introduction)A forecast that “millions of people… will soon consider” AI training this way
“largest theft of labor in human history”Microsoft’s Director of Applied ScienceQuoted with “perhaps”; full context not in the public text
“accidental cover up”Dr. Brent HechtA concern about OpenAI’s output filter reducing rights holders’ visibility
“doom loop”Microsoft document“an end-product threatens the economic foundations of its essential suppliers”
“users won’t click”An OpenAI software engineer“no matter how prominently we show the links, users won’t click”
“existential threat”OpenAI’s Head of ChatGPTPublishers face one; products “are largely substitutive, period”

The Five Kinds of Copying the NYT Lawsuit Alleges

nyt lawsuit microsoft openai knew news content theft d crowbar lying flat with a curved claw end

The brief breaks the case into five stages, and seeks a ruling on each. Two of them are pressed against Microsoft only, for a procedural reason.

StageWhat the brief allegesAgainst
1. AcquisitionObtaining copies of articles, “often from behind paywalls or in violation of industry norms or terms of use”OpenAI and Microsoft
2. TrainingConfiguring models to predict the articles’ expressive contentOpenAI and Microsoft
3. GroundingUsing the Bing index to feed copies to models for real-time answersMicrosoft only
4. OutputsReproducing copies or derivatives in answers to usersMicrosoft only
5. “Horse trading”Supplying each other with copies “for monetary or other consideration”OpenAI and Microsoft

Why OpenAI is spared two stages, for now

A footnote says the motion seeks judgment on grounding and outputs “against Microsoft only as the issue is not ripe for determination against OpenAI until the resolution of Plaintiffs’ pending sanctions motion.” That sanctions motion is itself part of the NYT lawsuit record, and it is the first thing to watch.

The DMCA and damages requests

Separately, the brief asks the court to find that OpenAI intentionally removed copyright management information under 17 U.S.C. § 1202(b)(1), and that statutory damages can be awarded per article rather than per newspaper issue. That second point is the one with the biggest financial consequence.

The NYT Lawsuit on Scrapes, Datasets and Paywalls

nyt lawsuit microsoft openai knew news content theft e full cloth sack tied at the neck

The acquisition section is where the NYT lawsuit brief is most concrete, because it names datasets and counts.

WebText and WebText2

OpenAI built WebText for GPT-2, a scrape that “prioritiz[ed] web pages curated/filtered by humans”, and news articles were its most prevalent content type. WebText2, used for GPT-3 and GPT-3.5, contains at least 6,552 scraped works from The Times, 18,609 from the Daily News papers and 66,780 from Ziff Davis, according to the brief.

Common Crawl

An internal OpenAI Slack list of the “Top 100 domains by document count” for a filtered Common Crawl training set included 2,064,805 articles from nytimes.com and 2,182,079 from chicagotribune.com. The brief says OpenAI planned to retain Common Crawl files “forever”, and that Common Crawl’s own terms require users to respect the terms of the underlying data.

The Annotated Corpus

OpenAI obtained The New York Times Annotated Corpus, over 1.8 million articles from 1987 to 2007, from the Linguistic Data Consortium under a licence limited to non-commercial research. The brief says OpenAI staff knew it “would not be appropriate” to use it to train a model, and did so anyway.

Paywalls

OpenAI’s corporate representative testified he was unaware of “any effort to detect paywall content” in the material used to train its models. Microsoft chief executive Satya Nadella testified that “anything that is paywalled should be licensed by anyone who wants to use it”. The brief also pictures Custom GPTs in OpenAI’s store named “Bypass Paywall” and “Remove Paywall”.

Plaintiff articles counted in named datasets, per the News Plaintiffs’ brief
chicagotribune.com in filtered Common Crawl 2,182,079
nytimes.com in filtered Common Crawl 2,064,805
NYT Annotated Corpus 1.8 million+
Project Mango training set, all plaintiffs 160,903
Ziff Davis works in WebText2 66,780

Bar widths are each count divided by 2,182,079. The 1.8 million figure is the brief’s “over 1.8 million”, so its bar is a floor.

The figure we could not find

AFP’s report on the NYT lawsuit said OpenAI scraped more than 10 million articles, “nearly a third” from The Times. We could not locate either number in the public text of the brief. It may come from a sealed exhibit or another filing, so treat it as reported rather than verified.

Microsoft and OpenAI's "Horse Trading" in the NYT Lawsuit

nyt lawsuit microsoft openai knew news content theft f paperboy satchel bag with a front flap

The fifth stage is the novel one, and it is where Microsoft’s role goes beyond hosting OpenAI’s compute.

Project Taxi

From 2019 to 2022, the brief says, Microsoft provided OpenAI with a copy of the Bing Index, which Microsoft’s counsel described as “billions” of webpages gathered for traditional search. The transfer, codenamed “Project Taxi”, included the plaintiffs’ content. OpenAI’s Chief Research Officer Bob McGrew described the deal in October 2021 as “a trade”.

GPT-3’s training data went the other way

In September 2020, the brief says, OpenAI gave Microsoft its GPT-3 training data, which Microsoft accessed in 2021 to decide how to build the models into its products. Because GPT-3 was already trained, the plaintiffs argue the exchange “bore no relationship to training any model at issue”.

Project Mango

Microsoft also supplied training data through an initiative called Project Mango, assembled into a dataset containing at least 160,903 unique plaintiff works. Several details of what OpenAI gave in return remain redacted.

Why this matters legally

Fair use is judged use by use. If content moved between the companies as consideration in a commercial deal, the plaintiffs argue that is a separate use with no transformative purpose at all, which is harder to defend than training. It is the clearest attempt in the NYT lawsuit to separate Microsoft’s liability from OpenAI’s.

The Copilot Click-Through Numbers in the NYT Lawsuit

The strongest market-harm evidence in the NYT lawsuit brief is Microsoft’s own data on what happens to traffic when an answer engine replaces a list of links.

What Microsoft measured

The brief cites Microsoft data showing that the overall click-through rate difference between Copilot and traditional Bing Search was 87 to 93 per cent for The Times, 83 to 91 per cent for the Daily News papers and 51 to 94 per cent for Ziff Davis. The introduction summarises this as Microsoft “recording 83-93% drops in click-through rates” for the newspapers.

Click-through drop, Copilot vs traditional Bing Search (range, Microsoft data cited in the brief)
The Times, low end 87%
The Times, high end 93%
Daily News papers, low end 83%
Daily News papers, high end 91%
Ziff Davis, low end 51%
Ziff Davis, high end 94%

The “won’t click” line

The brief pairs the data with an OpenAI software engineer’s remark: “no matter how prominently we show the links, users won’t click.” It also quotes Copilot’s home page: “Instead of clicking through links, we can talk through whatever you’re curious about.”

Nadella’s testimony

Microsoft’s chief executive agreed under oath, the brief says, that conversing with chatbots “has substituted” for going “to the underlying source”. That is an unusually direct concession to have on the record, and the plaintiffs use it to argue substitution, the harm the Supreme Court called “copyright’s bête noire” in the Warhol case.

The NYT Lawsuit's "Accidental Cover Up" Claim

The most serious-sounding allegation in the NYT lawsuit brief concerns what OpenAI did after it was sued.

What OpenAI built

Immediately after the lawsuits began, the brief says, OpenAI searched for and copied the plaintiffs’ content most likely to be output by its models, to populate a “Bloom filter”, a data structure used here to suppress the output of that content. Unlike the natural language processing that generates an answer, a Bloom filter simply checks whether a string has been seen before.

Why the plaintiffs object

OpenAI “did not suppress the output of content from any entity that had not sued it”, according to the brief. Its argument is that the filter’s purpose “was not to prevent OpenAI’s models from infringing copyrights, but to stop Plaintiffs from gathering evidence of OpenAI’s copying for use in litigation.”

Hecht’s phrase

This is where “accidental cover up” comes from. The brief says Hecht “expressed a concern that OpenAI would use such a filter, and called this behavior an ‘accidental cover up’ because it would result in ‘people who have a right over the content having less visibility into what was used for training.'” The word “accidental” is doing real work: it describes an effect, not an intent.

How it connects to sanctions

The pending sanctions motion is why the output claims against OpenAI were held back. If the court finds that evidence of outputs was suppressed, the consequences could reach beyond this motion, which makes that ruling a pivotal moment for the NYT lawsuit.

Microsoft filed its own public, redacted summary judgment memorandum in the News Plaintiffs’ cases on the same day. It tells a very different story about the product at the heart of the NYT lawsuit.

Training is fair use

Microsoft’s memorandum opens by saying its books brief “explains why the use of copyrighted works to train an LLM is fair use”, and that the conclusion “applies equally to news articles”, because building an LLM “is new and different and far beyond any previously existing use or market”.

Copilot is designed like search

Microsoft focuses on web grounding, which pairs the chatbot with Bing search. It says Copilot outputs “provide links directly to the webpages relied upon so users can click through”, and that “any website owner that objects to web grounding with its content can simply opt out.”

The logs

“Over 8.2 million chat logs produced in discovery prove that Copilot virtually never displays even a sentence of News Plaintiffs’ content to users,” the memorandum says. It adds that “few users even use Copilot for news or current events”, and that people “are not cancelling newspaper subscriptions” because of it.

IssueNews Plaintiffs’ briefMicrosoft’s memorandum
Training on articlesSubstitutive and commercial; not fair useTransformative; fair use applies “equally to news articles”
Links in answersClick-through down 83 to 93 per cent for newspapersOutputs link to sources “so users can click through”
Verbatim outputGrounding on at least 1,820 words of one article8.2 million logs show Copilot “virtually never” displays a sentence
Publisher controlScraping continued after opt-outs for groundingAny site owner “can simply opt out”
News useNadella: chat “has substituted” for visiting sources“Few users even use Copilot for news”
Hecht’s documentsAdmissions of theft and market harm“One employee’s individual perspective” (statement to AFP)

Where the two stories meet

Both sides agree that Copilot answers questions rather than listing links. They disagree about whether that is a new market or the same market with the publisher cut out, and that disagreement is the fourth fair use factor in a sentence.

What OpenAI Staff Said About Publishers, per the NYT Lawsuit

The NYT lawsuit brief spends several pages on OpenAI statements about news specifically, because Kadrey v. Meta suggested news publishers present “even stronger arguments against fair use” than many plaintiffs.

The Head of ChatGPT

The brief says OpenAI’s Head of ChatGPT wrote that publishers face an “existential threat” from its products, which “are largely substitutive, period” and “will get more and more substitutive as they get better.” Later it names OpenAI’s Nick Turley as making the same point about ChatGPT’s browsing feature.

Greg Brockman

OpenAI’s co-founder and president, the brief says, wrote that models are “particularly good at predicting text of news articles”, “excellent at news” and “very good at any news task”.

A slide by Dario Amodei

In a presentation from his time as a top OpenAI researcher, the brief says, Dario Amodei, now Anthropic’s chief executive, listed “News Generation” among GPT-3’s skills, with the sample query “What’s the NYT saying today?”

Media Manager

In 2024, OpenAI announced a tool called Media Manager to let content owners express preferences about AI use, including opting out of scraping. The brief says it “has since abandoned the project.” For publishers weighing their own options, that is a practical note: the promised control mechanism does not exist.

The Books Case Filed Beside the NYT Lawsuit

The NYT lawsuit plaintiffs were not the only ones to file on 17 September. The authors’ class, which includes novelists such as David Baldacci, filed a public version of its own partial summary judgment brief.

LibGen and Books3

That brief alleges OpenAI “torrented more than 4 million books” from the pirate site LibGen and downloaded at least 21,000 from Books3, a 196,640-book dataset sourced from a pirate site. It says the LibGen sets were renamed “Books1” and “Books2”, a “deliberately vague” description.

The same Microsoft scientist

Hecht appears there too. The authors quote him writing that Microsoft’s failure to license data “has started a doom loop” because the LLM “end product threatens the economic foundations of its essential suppliers”, and testifying to “serious concern” that authors would suffer “economic harm”.

Why the two briefs matter together

Our earlier coverage of AI training on copyrighted books set out the Bartz and Kadrey rulings. Those courts separated how material was acquired from how it was used. The NYT lawsuit brief and the authors’ brief both press hardest on acquisition, paywalls and pirate sites, because that is where earlier defendants lost ground.

Where the NYT Lawsuit Goes Next

The filing is one step in the NYT lawsuit, and the timetable is slower than the headlines suggest.

The government’s position

In early September the US Department of Justice filed a statement of interest supporting OpenAI and Microsoft, invoking scientific progress, economic growth and national security. We covered its reasoning in detail in the US government’s fair use brief.

More publishers, same court

The multidistrict litigation keeps growing. On 4 September the Seattle Times and Newsday filed their own suit, covered in our article on the Seattle Times lawsuit, and those cases are likely to follow the rulings in this one.

Timing

AFP reports that a summary judgment ruling in the NYT lawsuit is not expected until 2027. If Judge Stein grants the plaintiffs’ motion on liability, damages would still need to be resolved. If he denies it, the NYT lawsuit heads towards trial on the disputed facts.

What to watch

Three things will move the case: the ruling on the sanctions motion over the Bloom filter, any decision on per-article statutory damages, and whether the court treats the “horse trading” exchanges as uses separate from training.

What Publishers and AI Buyers Should Take From the NYT Lawsuit

The case is about two companies, but the record now public is useful to anyone publishing online or buying AI tools.

For publishers

Opt-outs work only for the crawlers that honour them. The brief says acquisition through Common Crawl and anonymous scrapes “sidestepped” robots.txt, and that Microsoft moved Bing-crawled content to OpenAI. Blocking named AI bots is necessary but not sufficient; licensing terms and monitoring matter as much.

For AI buyers

Enterprise contracts increasingly include copyright indemnities, and this record shows why. Ask vendors which data sources their models were trained on, whether web-grounded features respect publisher opt-outs, and what they will cover if an output reproduces protected text.

For everyone reading the headlines

“Microsoft knew it was theft” is a strong summary of a narrower sentence. Hecht forecast that millions would see AI training that way and warned about its effect on creators. That is significant evidence of foreseeable market harm, and still not an admission of liability. For more on how AI rulings affect business adoption, see our AI strategy services and the AI models, tools and releases hub.

References