The USA Today lawsuit filed against OpenAI on Thursday 8 October 2026 makes USA TODAY Co. the latest news publisher to take the ChatGPT maker to court, and one of the largest by footprint. The company, formerly known as Gannett, and 13 affiliated entities say OpenAI copied “hundreds of thousands” of articles from 19 of their newspapers into the training data behind ChatGPT, then used them to answer readers’ questions. They want more than $250 million in damages, a jury trial, and an order to destroy any GPT model trained on their journalism.

The case was first reported by Reuters and picked up the same afternoon by The Verge, which noted that “this is just the latest in a string of copyright lawsuits filed against OpenAI”. We read the full 79-page complaint, filed as USA Today Co., Inc. v. OpenAI Foundation, No. 1:26-cv-08892 in the Southern District of New York. It is far more specific than the coverage suggests, with dataset counts per website, 19 worked examples from GPT-5.6, and internal OpenAI quotes lifted from the consolidated New York litigation.

This article walks through what the USA Today lawsuit claims, the numbers it relies on, how the $250 million figure adds up, where it sits among the other publisher cases, and what it means for businesses that use ChatGPT. For the previous filing in this wave, see our report on the Seattle Times and Newsday lawsuit against OpenAI and Microsoft.

What the USA Today Lawsuit Claims

usa today lawsuit openai copyright 250 million b tiered newsagent stand holding nineteen folded newspapers

The core accusation is simple. In the complaint’s words, “OpenAI’s commercial success rests on large-scale copyright infringement.” The publishers say OpenAI never asked permission and never offered payment: it “simply took” their work “and used it to build products worth hundreds of billions of dollars.”

The USA Today lawsuit identifies three separate ways the copying allegedly happened. First, OpenAI scraped the newspapers’ websites, or bought scraped copies, to build training sets. Second, it copied that material again and again while training and fine-tuning its models. Third, its products reproduce or repackage the articles when users ask about the news.

Who filed the USA Today lawsuit and where

The plaintiffs are USA TODAY Co., Inc. and 13 subsidiaries that hold the copyrights for individual titles, from Gannett Satellite Information Network, LLC (which owns USA TODAY and the Indy Star) to Phoenix Newspapers, Inc. (The Arizona Republic). All 14 list their principal place of business as Pittsford, New York. Their lawyers are Rothwell, Figg, Ernst & Manbeck, P.C. of Washington, DC, led by Steven Lieberman.

They filed in Manhattan because that is where the other publisher cases against OpenAI have been gathered. The complaint’s opening paragraph says the plaintiffs “join the long list of copyright holders” suing AI companies, “many of which have been consolidated in this Court”. The docket on CourtListener shows a $405 filing fee and, at the time of writing, no judge yet assigned.

Seven OpenAI companies, and no Microsoft

The defendants are seven OpenAI entities: OpenAI Foundation, OpenAI GP, LLC, OAI International, Inc., OpenAI OpCo, LLC, OpenAI Global, LLC, OAI Corporation and OpenAI Group PBC. That list tracks OpenAI’s October 2025 recapitalisation, when the nonprofit became the OpenAI Foundation and the for-profit became a public benefit corporation.

Microsoft is not a defendant, which sets the USA Today lawsuit apart from the New York Times and Seattle Times cases. Microsoft still appears throughout the filing, though. The complaint says Microsoft gave OpenAI a copy of the Bing index under a project codenamed “Project Taxi”, and ran a crawler called “Project Mango” on OpenAI’s behalf. Count I treats that exchange of data as its own act of infringement.

The 19 Newspapers in the USA Today Lawsuit

usa today lawsuit openai copyright 250 million c microfilm reader with two film reels and a lit page screen

The plaintiffs own the copyrights in content from 19 publications. Together, the complaint says, they “have won dozens of Pulitzer Prizes and employ hundreds of journalists across thirteen states.” The table below lists each title with the website the complaint names and any Pulitzer count it gives.

PublicationWebsiteFounded, per complaintPulitzers stated
USA TODAYusatoday.com1982Not stated
The Des Moines Registerdesmoinesregister.com184917
Detroit Free Pressfreep.com183110
The Courier-Journalcourier-journal.comNot stated10
The Detroit Newsdetroitnews.com18733
The Tennesseantennessean.com19073
The Enquirer (Cincinnati)cincinnati.com1841Includes 2018 local reporting
Asbury Park Pressapp.com1879At least 1
Indianapolis Starindystar.com1903“Multi-time” winner
Milwaukee Journal Sentineljsonline.com1837“Multiple”
The Arizona Republicazcentral.comNot statedPrize-winning reporting
The Columbus Dispatchdispatch.com1871Not stated
The Oklahomanoklahoman.comOver a century agoNot stated
The Palm Beach Postpalmbeachpost.com1916Prize-winning photography
The Bergen Recordnorthjersey.com1895Not stated (1996 Loeb Award)
The Knoxville News-Sentinelknoxnews.com1886Not stated
Naples Daily Newsnaplesnews.com1923Not stated
Democrat & Chronicledemocratandchronicle.comNearly 200 years agoNot stated
Star News (Wilmington)starnewsonline.com1867Not stated

Why local titles dominate the USA Today lawsuit

Only one of the 19 is a national paper. The rest are regional dailies such as the Democrat & Chronicle, described as Rochester’s “sole remaining daily print newspaper”, and the Star News, “North Carolina’s oldest continuously published daily newspaper.” That mix matters to the argument. The complaint quotes OpenAI’s own February 2026 statement that “the demand for reliable local news is already visible inside ChatGPT at a rate of about 1 million prompts per week,” and then says: “OpenAI is now the one answering those questions.”

The Training Data Numbers Behind the USA Today Lawsuit

usa today lawsuit openai copyright 250 million d flatbed scanner with its lid up and a newspaper page on the glass

Most AI copyright complaints say a model was trained on “vast amounts” of content and leave it there. The USA Today lawsuit puts numbers on it, drawn from two public sources: the domain list OpenAI published for its GPT-2 training set, and a 2023 Washington Post analysis of Google’s C4 dataset.

More than 160,000 WebText entries

WebText was the internal corpus OpenAI built for GPT-2 from “the text contents of 45 million links posted by users” of Reddit. OpenAI published the top domains in WebText on GitHub, and the plaintiffs counted their own sites in it: more than 160,000 entries in total, 83,266 of them from usatoday.com alone.

WebText entries by website, as counted in the complaint (usatoday.com = 100%)
usatoday.com 83,266
freep.com 12,994
jsonline.com 9,814
azcentral.com 9,073
detroitnews.com 7,028
indystar.com 6,726
cincinnati.com 6,194
dispatch.com 6,082

The 12 domains the complaint itemises add up to 161,587 entries, and usatoday.com accounts for just over half of them. The USA Today lawsuit also says internal OpenAI documents describe news articles as “the most prevalent type of content in the WebText dataset.”

122 million tokens in C4

The second figure comes from C4, a filtered English-language snapshot of Common Crawl taken in 2019. Citing the Washington Post’s analysis, the USA Today lawsuit says the publishers’ domains account for “over 122 million tokens” in C4, including 23 million from usatoday.com and 12 million from azcentral.com.

Tokens in the C4 snapshot by website, millions (usatoday.com = 100%)
usatoday.com 23.0M
azcentral.com 12.0M
freep.com 8.2M
jsonline.com 8.1M
detroitnews.com 8.1M
cincinnati.com 7.8M
indystar.com 7.6M
tennessean.com 6.5M

One caution on the arithmetic. The 19 domains listed in that paragraph add up to roughly 115 million tokens, not 122 million. The complaint says the total is “over 122 million” and that it “includ[es]” the listed sites, so the gap presumably sits in domains it does not itemise. Either way, the scale is clear.

Why weighting matters as much as volume

The USA Today lawsuit leans on how GPT-3 was trained, not just what went in. It points out that WebText2, an expanded version of WebText, made up “less than 4% of the total tokens” in GPT-3’s training mix but was weighted at 22%, because OpenAI’s own paper says “datasets we view as higher-quality are sampled more frequently”. The publishers’ argument is that professional journalism was not incidental filler: it was the high-quality material the model was deliberately shown more often. The complaint also notes, from tax filings, that OpenAI gave Common Crawl $250,000 in 2023.

How the USA Today Lawsuit Reaches $250 Million

usa today lawsuit openai copyright 250 million e jury deliberation table with twelve chairs

The headline figure comes from US statutory damages. Under 17 U.S.C. § 504(c), a court can award between $750 and $30,000 for each infringed work, rising to $150,000 per work for wilful infringement. Separately, section 1203 allows up to $25,000 for each violation of the rule against removing copyright management information. The complaint cites both maximums.

The three counts

CountLegal basisWhat it alleges
I. Copyright infringement17 U.S.C. § 501Copying articles into datasets, swapping data with Microsoft, training and storing models that memorise them, and outputting copies and derivatives in ChatGPT
II. Vicarious infringementJudge-made doctrine under the Copyright ActThe parent and holding companies directed, controlled and profited from the infringement by OpenAI OpCo and OAI International
III. Removal of copyright management information17 U.S.C. § 1202(b)(1) (DMCA)Stripping bylines, titles, copyright notices and terms of use with the Dragnet, Newspaper and Gutentag text extractors

That makes the USA Today lawsuit leaner than the Seattle Times case, which pleaded seven counts including trademark dilution and state-law claims. The USA Today lawsuit sticks to copyright and the DMCA.

What the $250 million implies

The USA Today lawsuit does not say how it reached $250 million, but the statute makes the possibilities easy to work out. The table divides the claim by each per-work award level to show how many works the court would need to count.

Award per workWhen it appliesWorks needed to reach $250 million
$150,000Statutory maximum for wilful infringementAbout 1,667
$30,000Statutory maximum, not wilfulAbout 8,334
$750Statutory minimumAbout 333,334

The bottom row is the interesting one. At the statutory minimum, “hundreds of thousands” of articles lands almost exactly on the $250 million claim. The catch is how a court counts “works”. Section 504(c) says all parts of a compilation “constitute one work”, so if a newspaper registered a whole edition as a single collective work, every article in that edition may earn only one award. The copyright registrations attached as Exhibit A will decide much of this, which is why statutory damages in AI cases are rarely as simple as multiplying articles by a fee.

The remedy that worries AI companies most

Beyond money, the prayer for relief asks the court to order “destruction under 17 U.S.C. § 503(b) of all GPT or other LLM models and training sets that incorporate the USA TODAY Plaintiffs’ content.” The Seattle Times and New York Times asked for the same thing. Courts have never ordered a frontier model destroyed, and it would be an extraordinary step, but the demand gives every plaintiff in this wave leverage in settlement talks.

The GPT-5.6 Exhibits at the Heart of the USA Today Lawsuit

usa today lawsuit openai copyright 250 million f sledgehammer standing on a smashed stone slab

The most modern part of the complaint is Exhibit B. For each of the 19 titles, the plaintiffs gave GPT-5.6 the same prompt: “Please find and summarize the article with this title and give me an in-depth summary,” followed by a real headline. Each time, the complaint says, the model retrieved the article and produced a summary that followed its structure and reproduced its substance.

PublicationHeadline used in the promptWhat the complaint says GPT-5.6 did
USA TODAYThat’s not Kathy Hochul. AI campaign ads are going too farExtensive summary following the article’s sequence
Indianapolis StarFBI raids home, business of Westfield developerMulti-section summary with the same structural organisation
Democrat & ChronicleBiggest and wildest snowstorms in Rochester NY rankedReproduced the ranked lists in the original order, with details such as the 63-hour snowfall of 1900
The TennesseanFawn Weaver and Uncle Nearest: How the Tennessee whiskey founder’s tenure unraveled, a timelineFollowed the same chronological timeline
The Arizona RepublicOwls and Bulldogs were a bust. How Arizona State became the Sun DevilsTraced the structure and reproduced exact details
Star NewsThese Eastern NC barbecue spots turn a day trip into a feastReproduced the section-by-section recommendations
The Columbus DispatchOhio is losing people, housing is a big reason they fleeTraced the structure and reproduced exact details
Naples Daily NewsBeloved cougar dies at Shy Wolf Sanctuary in NaplesMulti-section summary reproducing the key facts

Summaries, not verbatim copies

Notice what these exhibits are, and what they are not. The New York Times and Seattle Times complaints showed long passages copied word for word; the Seattle Times found 88 consecutive words reproduced from its Boeing reporting. The USA Today lawsuit mostly shows something different: search-style answers built by retrieving a live article and rewriting it.

That is retrieval augmented generation, or RAG, and the complaint explains it at length. Its theory is that a detailed summary is a market substitute even when the words change, because “users have less need to navigate to those sources”. It also says OpenAI “affirmatively post-train[ed] its models” to summarise articles instead of returning them. Whether close paraphrase of facts and structure infringes is a harder legal question than verbatim copying, and it is the part of the USA Today lawsuit most worth watching.

The traffic argument in OpenAI’s own words

The publishers back the substitution theory with quotes from OpenAI staff. The complaint says Nick Turley, OpenAI’s head of ChatGPT, wrote that publishers face an “existential threat” and that OpenAI’s products “are largely substitutive, period” and “will get more and more substitutive as they get better”. An OpenAI software engineer is quoted as writing that “no matter how prominently we show the links, users won’t click.” Internal documents, it says, called ChatGPT “the modern newsstand”.

What OpenAI's Internal Documents Add to the USA Today Lawsuit

Much of the sharpest material in the USA Today lawsuit is not new. It is cited to a September 2025 filing in In re OpenAI, Inc. Copyright Infringement Litigation, No. 25-md-3143, the consolidated case in which the New York Times and other publishers obtained OpenAI’s internal records through discovery. The USA Today plaintiffs are, in effect, building on evidence other publishers paid to uncover.

Memorisation as the design objective

The USA Today lawsuit quotes an OpenAI vice-president of research saying: “We train our networks to memorize the training data — that’s their objective.” It says that by November 2019 OpenAI worried internally about “accidentally regenerating copyrighted works”, and that in June 2022 employees expected GPT-4 to have “memorized a ton of data and therefore will be insanely good at regurgitation.” It also quotes co-founder Greg Brockman calling OpenAI’s models “excellent at news” and “very good at any news task”.

Paywalls, filters and a “cover up”

The USA Today lawsuit alleges that OpenAI’s output filters “did not suppress output of content from any entity that had not sued it”, and quotes a Microsoft executive calling that approach an “accidental cover up”. It says that when an employee told Brockman about “a hack to get around nytimes paywall”, he replied “ah nice”, and that OpenAI’s corporate witness testified the company had no “method for detecting paywalled content” in its datasets.

It adds examples from OpenAI’s GPT store, including a “Bypass Paywall” custom GPT and a “News Summarizer Ace” that promised users they could “skip paywalls just using the link text or URL”. Each of these is an allegation drawn from another case’s record, and none of it has been tested at trial.

How copyright information was stripped

Count III of the USA Today lawsuit rests on the tools OpenAI used to clean web pages. The complaint names three extractors, Dragnet, Newspaper and Gutentag, and says they target terms such as “byline”, “copyright” and “©”. Dragnet’s own research paper describes its goal as separating the main content from “navigation chrome, advertising blocks, copyright notices and the like”. The publishers argue that choosing such tools for a newspaper corpus was a knowing removal of copyright management information under 17 U.S.C. § 1202.

Licensing Deals and the USA Today Lawsuit

USA TODAY Co. is not opposed to AI companies using its journalism. It wants to be paid and credited. That position runs through its recent history, and it explains why a company with several AI deals is suing the one partner it has not signed.

DateEventWhat it shows
30 July 2025Gannett licenses USA TODAY and 200+ local titles to PerplexityContent in Perplexity search and the Comet browser; terms undisclosed
18 November 2025Gannett renames itself USA TODAY Co., Inc.The national masthead becomes the corporate brand (NYSE: TDAY)
5 December 2025Multi-year AI licensing deal with MetaNew and archive content in Meta AI answers, with links back
8 October 2026Files the USA Today lawsuit against OpenAIMore than $250 million sought; no licence with OpenAI

When the Perplexity deal was announced, chief executive Mike Reed said the company was “committed to ensuring that our content is properly attributed and that we are fairly compensated.” In the Meta announcement, he called that deal “a testament to the value of the USA TODAY Network’s archival and real-time content locally and nationally.” The USA Today lawsuit is the other half of the same strategy: licence those who pay, sue those who do not.

Turning OpenAI’s own deals against it

The USA Today lawsuit uses OpenAI’s licensing record as evidence of wilfulness. It says OpenAI “knows a license is required—because it has paid for them”, naming agreements with the Associated Press, Axel Springer, The Atlantic and Vox Media among “over a dozen” news organisations. If a licence has a market price, the argument goes, taking the same thing without one is not an innocent mistake.

Terms of service and robots.txt

The USA Today lawsuit also points to the publishers’ own defences. Their terms of service ban using site content to develop any AI system, “including without limitation for training, re-training, fine tuning, grounding, retrieval augmented generation”. The newspapers in the USA Today lawsuit block OpenAI’s crawlers in robots.txt and send detected AI crawlers to a page that reads: “Crawling and scraping of our site is not permitted.” Those measures matter most for the RAG claims, because they show ChatGPT retrieving articles the publishers had told it not to fetch.

How the USA Today Lawsuit Fits the Wider Publisher Fight

The Verge counted the other publishers already suing OpenAI: The New York Times, The Intercept, Ziff Davis, CBC/Radio-Canada, Encyclopaedia Britannica, Merriam-Webster, The Seattle Times “and a coalition of nearly 400 local newspapers.” Most of the US news cases are now handled together in Manhattan before Judge Sidney H. Stein.

CaseDocketFiledDefendants
The New York Times v. Microsoft and OpenAI1:23-cv-11195December 2023OpenAI and Microsoft
The Intercept v. OpenAI1:24-cv-01515February 2024OpenAI
Ziff Davis v. OpenAI1:25-cv-043152025OpenAI
In re OpenAI (consolidated)1:25-md-031432025OpenAI and Microsoft
The Seattle Times Company and Newsday v. OpenAI1:26-cv-076444 September 2026OpenAI and Microsoft
Times Publishing Company v. Microsoft1:26-cv-0808216 September 2026Microsoft (lead defendant)
USA Today Co. v. OpenAI Foundation1:26-cv-088928 October 2026Seven OpenAI entities

Why publishers keep filing now

Two things have changed in 2026. The first is evidence: the discovery record in the consolidated case gives every new plaintiff quotes it could never have obtained alone, and the USA Today lawsuit cites that record on almost every page. The second is money. The complaint notes OpenAI’s $852 billion valuation, its June 2026 filing for an initial public offering, and an advertising business that reached a $1 billion annual run rate in August. A company preparing to go public has strong reasons to settle disputes that could appear as risks in its prospectus.

The one thing this case adds

What is genuinely new is scale at the local level. Earlier suits came from national brands or small groups of papers. USA TODAY Co. owns more than 200 local publications, and the 19 in this case span 13 states. If the substitution theory works for a Rochester snowstorm ranking or a Naples wildlife story, it works for almost any local news site in America.

What OpenAI Has Said About the USA Today Lawsuit

So far, nothing specific. The Verge reported that OpenAI “didn’t immediately respond” to its request for comment, and Reuters reported the same. We could not find an OpenAI statement on the USA Today lawsuit at the time of writing.

OpenAI’s general position is well known. It has long argued that training on publicly available internet material is fair use, and in January 2024 it said it offers publishers an opt-out and pursues licensing partnerships. The complaint quotes that post’s claim that OpenAI “led the AI industry in providing a simple opt-out process for publishers”, and answers it with the custom GPTs built to get around paywalls.

What OpenAI is likely to argue

Based on its filings in the other cases, OpenAI can be expected to argue that training is transformative fair use, that the GPT-2-era WebText data says little about today’s models, that the RAG summaries convey facts that copyright does not protect, and that many claims are out of time. The USA Today lawsuit anticipates the last point by stressing that OpenAI “must continuously feed its models new copyrighted material”, which would make the infringement “not a past wrong but an ongoing one.”

What the USA Today Lawsuit Means for Businesses Using AI

Most firms reading this are not publishers or AI labs, but they do use ChatGPT, Copilot or tools built on similar models. The USA Today lawsuit does not make that use unlawful. It does sharpen some practical questions about how AI tools handle other people’s content.

Summaries of paywalled news are the risky use

The exhibits show exactly the behaviour a business can control: asking a chatbot to fetch and summarise a specific article. Reading the original, or quoting a short extract with credit, is safer than pasting a detailed AI summary of paywalled reporting into a client newsletter or sales deck. If your team relies on news monitoring, a licensed media-monitoring service is a cleaner route.

Check what your AI vendor promises

Ask vendors three questions: what their models were trained on, whether they respect robots.txt and site terms when retrieving content, and whether they will indemnify you if an output infringes. OpenAI announced its Copyright Shield for ChatGPT Enterprise and API customers in November 2023, and Microsoft offers similar commitments for Copilot. Read the conditions; indemnities usually require that you used the provider’s built-in filters and did not deliberately prompt for protected material.

Build your own RAG systems carefully

Companies building internal assistants should learn from the robots.txt and terms-of-service sections of the USA Today lawsuit. Do not point a retrieval pipeline at third-party news sites that forbid it. Keep a record of the sources your system indexes and the licences that cover them. Most natural language processing projects can run on your own documents, licensed feeds and public-domain material without touching a publisher’s paywall.

The UK position

UK copyright law is narrower than the US fair use doctrine. Section 29A of the Copyright, Designs and Patents Act 1988 allows text and data mining without permission only for non-commercial research. A UK business training or grounding a commercial model on scraped news content cannot rely on that exception, so licensing matters even more here. Our compliance and IT governance teams can help you put an AI content policy in place.

A short checklist

  • Ban the “find and summarise this article” pattern for paywalled sources in client-facing work.
  • Record which AI tools staff use, and which of them carry a copyright indemnity.
  • Ask vendors for a written statement on training data provenance and crawler behaviour.
  • Keep internal RAG indexes to sources you own or license, and log them.
  • Plan for model changes: if a court ever restricts a model, a multi-model setup avoids a single point of failure.

If you are shaping a wider AI strategy, treat content provenance as a design requirement from the start rather than a legal problem to solve later.

What Happens Next in the USA Today Lawsuit

The immediate steps are procedural. OpenAI must be served and will then have time to respond, usually extended by agreement in cases of this size. The most likely early development is that the USA Today lawsuit is related to, or folded into, the consolidated proceedings before Judge Stein, who already has the Seattle Times case.

Timing relative to the consolidated case

That consolidated litigation is well advanced. Discovery has produced the internal documents quoted here, and summary judgment motions are pending. A ruling on fair use there would shape every newer case, including this one. The USA Today plaintiffs may gain from arriving late: they can rely on discovery already done, while pleading fresh claims about GPT-5.6 and ChatGPT search that the older complaints could not.

What to watch

  • Whether OpenAI or USA TODAY Co. signals settlement talks, given the company’s history of licensing.
  • How the court treats the RAG summary claims, which are central here and less tested than verbatim copying.
  • Whether the copyright registrations in Exhibit A support per-article or per-edition damages.
  • Whether more of the 200-plus USA TODAY Network titles are added to the USA Today lawsuit later.

For more coverage of the legal pressure on AI developers, see our report on the Florida attorney general’s request to halt OpenAI model development, and our AI models and tools hub for product news.

Frequently Asked Questions About the USA Today Lawsuit

Who is suing OpenAI in the USA Today lawsuit?

USA TODAY Co., Inc., formerly Gannett, and 13 subsidiaries that own the copyrights for 19 newspapers, including USA TODAY, the Detroit Free Press, The Arizona Republic, The Tennessean, The Des Moines Register and the Milwaukee Journal Sentinel.

How much money does the USA Today lawsuit seek?

More than $250 million. The complaint cites statutory damages of up to $150,000 for each wilfully infringed work and up to $25,000 for each removal of copyright management information.

Is Microsoft a defendant?

No. The defendants are seven OpenAI companies. Microsoft is mentioned because the complaint says it shared Bing index data with OpenAI and ran a crawler for it.

What evidence does the USA Today lawsuit offer?

Counts of the publishers’ content in the WebText and C4 datasets, 19 examples of GPT-5.6 summarising their articles, and internal OpenAI quotes taken from the consolidated New York litigation.

Has OpenAI responded?

Not publicly. OpenAI did not immediately respond to requests for comment from The Verge or Reuters.

References