AI scraping was, in the words of a senior Microsoft scientist, perhaps “the largest theft of labor in human history”. That line, from Microsoft’s Director of Applied Science Brent Hecht, is one of several internal remarks made public on 17 September 2026 in a newly unredacted court filing in the copyright case brought by The New York Times and other publishers against OpenAI and Microsoft.

The quote spread fast, and some reports on AI scraping stretched it. This article looks at exactly what the filing says, how the AI scraping it describes actually worked, why “theft of labor” is a more precise charge than it first sounds, and what it means for publishers and for businesses that use AI. For our full walkthrough of the brief itself, see our report on the NYT lawsuit.

What the Unredacted Filing Says About AI Scraping

ai scraping microsoft exec largest theft of labor b ai scraping sewing machine with a hand wheel

The document is the News Plaintiffs’ combined summary judgment brief in the consolidated case In re OpenAI, Inc. Copyright Infringement Litigation, case 1:25-md-03143 in the Southern District of New York. The public version, docket entry 1977-1, runs to 92 pages.

The sentence itself

The brief opens: “This case is about, as Microsoft’s Director of Applied Science put it, ‘an astonishing theft of unprecedented proportions’; perhaps the ‘largest theft of labor in human history.'” Note the word “perhaps”. The brief quotes the phrase with a qualifier, and the full surrounding text is not public.

Two quotes, possibly two documents

The brief cites the two phrases to different exhibits. TechCrunch reports they came from a January 2023 internal memo by Hecht. Engadget says a 2023 Microsoft document stated that “millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions”. Read that way, it is a forecast of public opinion, not a confession.

What TechCrunch flagged

TechCrunch noted that “much of the new information comes from The Times’ own brief, not the underlying exhibits, which remain sealed”, and that the quotes are “presented without their original context”. That is an important caveat for every quote about AI scraping below.

Who Brent Hecht is

The brief identifies Hecht as Microsoft’s Director of Applied Science. He appears several times. Beyond the theft lines, he warned that OpenAI’s plan to filter outputs against a list of plaintiffs’ articles would amount to an “accidental cover up”, leaving “people who have a right over the content having less visibility into what was used for training”. He also agreed that “in the simplest terms: AI search engines scrape answers from websites.”

Microsoft’s response

Microsoft told AFP the statements reflect “one employee’s individual perspective”. Spokesperson Alex Haurek told The New York Times: “Microsoft’s position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers’ journalism.” OpenAI and Microsoft did not answer TechCrunch.

How AI Scraping Worked, According to the Brief

ai scraping microsoft exec largest theft of labor c vintage cash register with a pull out drawer

The headline quote is about a practice, and the brief describes that practice in unusual detail. The AI scraping it alleges ran through at least seven channels.

ChannelWhat the brief allegesScale cited
Direct crawlingBots copied pages without checking paywalls or terms of useNot stated
Common CrawlDownloaded repeatedly and kept “forever”2,064,805 nytimes.com documents in one filtered set
Bing IndexMicrosoft’s search crawl shared with OpenAIBillions of webpages
Project Taxi and Project MangoMicrosoft supplied training data to OpenAI160,903 unique plaintiff works in the Mango set
Mid-training setsNews used to improve “freshness”Over 91,692 copies of plaintiff works
Annotated CorpusResearch-only licence used for trainingAbout 1.8 million Times articles
Paywall workaroundsA “hack to get around nytimes paywall”Not stated

Crawlers, plain and simple

The brief quotes OpenAI’s own expert: “One cannot train on data without first acquiring it in digital form—typically by downloading it from the Internet—which requires making a copy of it.” It says both companies used web crawlers, or “bots”, to extract the content, and that OpenAI did not check whether it was behind a paywall or subject to terms restricting use.

Going around robots.txt

The robots.txt file is how websites tell crawlers what they may not fetch. The brief alleges that by acquiring content “through third parties such as Common Crawl and through anonymous scrapes, OpenAI also sidestepped the automated robots.txt regime”. In other words, a publisher that blocked a crawler could still be copied through a dataset someone else had built.

The paywall exchange

According to the brief, when OpenAI researcher Nick Ryder told OpenAI president Greg Brockman about “a hack to get around nytimes paywall”, Brockman replied: “ah nice.” OpenAI’s corporate representative said he was unaware of any “method for detecting paywalled content” in its datasets.

Stripping copyright notices

TechCrunch reports the filing describes deliberate efforts to strip copyright notices before data reached the model, because researchers “wouldn’t want model outputting” “copyright notices” to users. The brief turns this into a separate claim under section 1202 of the Digital Millennium Copyright Act, which protects copyright management information.

Trading data like currency

The brief says the two companies swapped publisher content through what staff called “horse-trading” deals. In September 2020, it says, OpenAI delivered its entire GPT-3 training corpus to Microsoft. “Instead of paying Plaintiffs to license their content, Defendants sold that content to each other on at least three occasions,” the brief argues.

A Short History of AI Scraping at OpenAI

ai scraping microsoft exec largest theft of labor d spinning wheel with a foot treadle

The brief also sets out how OpenAI’s appetite for news grew with each model. It reads as a timeline of AI scraping becoming more deliberate.

GPT-2 and WebText

For GPT-2, OpenAI built a new web scrape called WebText, which “emphasizes document quality by prioritizing web pages curated/filtered by humans”. According to the brief, news articles were the most common type of content in it.

GPT-3 and WebText2

For GPT-3 and GPT-3.5, OpenAI built an expanded WebText2 that “also disproportionately contained news content”. The brief says it holds at least 6,552 works from The Times, 18,609 from the Daily News papers and 66,780 from Ziff Davis.

The push for freshness

Later models needed current events. The brief says OpenAI’s goal was to “crush freshness” in “the domain of real-world news”, and that staff treated CC-NEWS, a dataset of news articles, as a “natural source of freshness”. That is AI scraping aimed squarely at recent journalism.

Memorisation

The brief quotes OpenAI’s VP of Research: “We train our networks to memorize the training data — that’s their objective.” By November 2019, it says, OpenAI worried internally about models “accidentally regenerating copyrighted works”. In June 2022 staff expected GPT-4 to be “insanely good at regurgitation”. Memorisation matters because it turns AI scraping from a question about inputs into a question about outputs.

Why AI Scraping Is Framed as a Theft of Labor

ai scraping microsoft exec largest theft of labor e water mill wheel with paddles

“Theft” gets the headlines. “Labor” is the more interesting word in the AI scraping debate, because it describes the economic harm the plaintiffs must prove.

The labour behind the text

The brief argues that AI products substitute for the “labor of the people” who produced the original content, citing Microsoft documents. Reporting a story takes interviews, travel, editing and legal checks. The finished article is the cheap part to copy. AI scraping, on this view, captures the output of that work without paying for any of it.

A supply chain that eats itself

One Microsoft document, as quoted, calls large models “a product that destroys its supply chain”. Another says: “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.'”

The “doom loop”

The same Microsoft document warned that “our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time”. This is where AI scraping becomes a business problem for the AI companies too. If AI answers stop people visiting news sites, those sites earn less, publish less, and the next model has less fresh material to learn from.

Jobs, in Microsoft’s own words

A Microsoft document states there is a “real risk” that generative AI could “significantly disrupt[] the employment of the very people who generated the data on which the foundation model was trained”. That is the “theft of labor” argument in plainer language.

The Copilot Numbers Behind the AI Scraping Claim

ai scraping microsoft exec largest theft of labor f kitchen stand mixer with a bowl

The brief offers hard numbers for the doom loop that AI scraping feeds. They come from Microsoft’s own measurement of its Copilot “answer engine” against traditional Bing search.

Drop in click-through rate, Copilot vs Bing search, per Microsoft data cited in the brief
Ziff Davis domains, worst case 94%
The Times and Daily News, worst case 93%
The Times and Daily News, best case 83%
Ziff Davis domains, best case 51%

Bar widths are the percentage drops themselves. The brief gives ranges of 83% to 93% and 51% to 94%.

Reading the range

TechCrunch reported “as much as 93%”. The brief’s fuller range is 83% to 93% for The Times and the Daily News papers. Even the best case means roughly five out of six clicks disappeared when users got an AI answer instead of a list of links.

Why links alone do not fix it

An OpenAI software engineer wrote in 2023 that “no matter how prominently we show the links, users won’t click”, according to Engadget’s account of the filing. If that is right, adding citations to AI answers does not restore the traffic that AI scraping takes away.

Microsoft's Own Words on AI Scraping, Mapped to Fair Use

Most of the new quotes about AI scraping are aimed at one legal target: the fair use defence. Here is how the plaintiffs use each one.

Quote as filedSpeakerWhat it is used to show
“largest theft of labor in human history”Brent Hecht, MicrosoftAwareness of the harm, framed with “perhaps”
“make a complete mockery of the idea of ‘fair use'”Same Microsoft executiveDoubt inside Microsoft about its own defence
“AI search engines scrape answers from websites”Hecht, agreeingRetrieval copies, not just training
“anything that is paywalled should be licensed”Satya Nadella, depositionAn industry norm that was allegedly broken
Chatbots have “substituted” for visiting the sourceSatya Nadella, under oathMarket substitution, the fourth factor
“existential threat” to publishersNick Turley, OpenAIProducts are “largely substitutive”

The fourth factor

US fair use weighs four factors. The fourth, the effect of the use on the market for the original, has become central in AI cases. Quotes about AI scraping, substitution, click-through collapse and “existential” threats go straight at it.

Nadella’s retraining remark

Nadella testified that if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall”, he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models”. That shows retraining was a remedy Microsoft itself had in mind.

A note on attribution

Some outlets attributed the “existential threat” quote to Greg Brockman. The brief attributes it to OpenAI’s Head of ChatGPT, Nick Turley. Brockman’s quoted lines are that models are “excellent at news” and “very good at any news task”.

What Microsoft Argues Back

Microsoft filed its own summary judgment memorandum the same day, and it tells a very different story about AI scraping and its products.

Training is transformative

Microsoft says using copyrighted works to train a model is fair use, and that the argument “applies equally to news articles” because building one “is new and different and far beyond any previously existing use or market”.

Copilot links to sources

Microsoft says Copilot outputs “provide links directly to the webpages relied upon so users can click through”, and that “any website owner that objects to web grounding with its content can simply opt out”.

The logs

“Over 8.2 million chat logs produced in discovery prove that Copilot virtually never displays even a sentence of News Plaintiffs’ content to users,” the memorandum says. It adds that “few users even use Copilot for news or current events”.

Where the stories collide

Both sides agree Copilot answers questions instead of listing links. They disagree about whether that is a new market or the old one with the publisher cut out. That single dispute is where the AI scraping case will be won or lost.

Why the Law on AI Scraping Is Still Unsettled

None of this decides whether AI scraping is lawful. The legal backdrop on AI scraping has so far been friendlier to AI companies than the quotes suggest.

Courts have leaned towards AI firms

Judges in earlier AI copyright cases have largely accepted that training can be fair use, while warning that the law is not settled. The brief leans on one of those rulings, which said news publishers present “even stronger arguments against fair use” than many plaintiffs.

Washington sided with OpenAI

Earlier this month the Trump administration filed a brief defending OpenAI’s unlicensed use of copyrighted material for training. We covered it in the US government siding with OpenAI on fair use.

No trial date yet

The filings are summary judgment motions, asking the judge to rule on key issues without a trial. Futurism, citing Axios, reports the judge is expected to decide whether the case goes to trial sometime in 2027. The underlying exhibits remain sealed.

Where AI Scraping Stands Today

The filing describes events from 2019 to 2024. It is worth asking what has changed since, because AI scraping has not stopped.

Purpose of AI crawling on Cloudflare’s network, 12 months to mid-2025
Training 80%
Search 18%
User actions 2%

Bar widths are Cloudflare’s published shares, which sum to 100%.

Training still dominates

Cloudflare’s 2025 data shows most AI crawling is for model training, not for answering a live question. It also found OpenAI’s GPTBot share of AI and search crawler traffic grew from 4.7% in July 2024 to 11.7% in July 2025.

The crawl-to-click gap

Cloudflare measures how many pages a platform crawls for each visitor it sends back. In July 2025 Anthropic’s crawlers visited about 38,000 pages for every referral, the highest among major AI firms. That is the doom loop measured from the outside.

Bots that ignore the rules

Most large AI firms say their crawlers respect robots.txt, and Cloudflare lists the leading ones as verified. But unverified or spoofed bots are a routine cybersecurity problem for site owners, and the brief’s claim about anonymous scrapes shows why a polite request in a text file is not a lock.

Blocking by default

On 1 July 2025 Cloudflare changed its default to block AI crawlers unless they pay, calling it “Content Independence Day”. Publishers now have more tools than they did when the scraping in the brief took place. Our guide to how websites are blocking AI crawlers covers the options in detail.

What AI Scraping Means for Publishers and AI Buyers

The case will take years. The lessons about AI scraping are available now.

For publishers: robots.txt is not enough

The brief’s own allegation is that third-party datasets let AI scraping sidestep robots.txt. Blocking named crawlers still slows AI scraping, but publishers should also track where their content appears in public datasets, keep clear terms of use and record licence terms.

For publishers: keep copyright information attached

The DMCA claim depends on copyright management information being present and removed. Clear bylines, notices and terms links make that claim stronger if it is ever needed.

For AI buyers: ask about provenance

If you build on a model, ask the vendor what its training data licences cover, whether it offers copyright indemnity, and what happens if a court orders retraining. Nadella’s own testimony shows retraining is a live possibility.

For AI buyers: watch your own reliance on answer engines

If AI answers cut clicks to publishers by 83% or more, they will do the same to your own site. Measure how much of your traffic now comes from AI assistants and plan for it.

AI Scraping FAQ

Who said “the largest theft of labor in human history”?

Brent Hecht, Microsoft’s Director of Applied Science, according to the brief. It quotes him with the qualifier “perhaps”, and TechCrunch reports the line came from a January 2023 internal memo.

Does Microsoft agree with him?

No. Microsoft says the remarks reflect one employee’s view and that its uses are transformative and lawful.

What is AI scraping?

The automated copying of web content by bots to train or feed AI systems, either directly or through datasets such as Common Crawl.

Has a court ruled on this case?

Not yet. The parties have filed summary judgment motions, and a decision on whether the case goes to trial is expected around 2027.

Can publishers stop AI scraping?

Partly. They can block known crawlers, use services that block AI bots by default, and license content on their own terms. Datasets built by third parties are harder to control.

References