USA Today Co seeks more than $250 million from OpenAI over alleged article copying
The federal complaint says OpenAI trained models on hundreds of thousands of articles from 19 USA Today publications and reproduced their content for users.
USA Today Co sued OpenAI in federal court in New York, alleging that the artificial-intelligence company copied hundreds of thousands of articles from 19 news publications without permission or payment, then used the material to train models and generate user outputs. The plaintiffs are seeking more than $250 million in damages and an order barring further use of their journalism.
The complaint, filed Oct. 8 in the U.S. District Court for the Southern District of New York, alleges copyright infringement by a group of OpenAI entities. It says OpenAI scraped or otherwise acquired the publishers’ work for model training, and that its systems reproduce or repackage the material in responses to users.
The publications include USA TODAY, The Tennessean, Indy Star, The Bergen Record, The Enquirer, Asbury Park Press, Democrat & Chronicle, The Knoxville News-Sentinel, Naples Daily News, The Oklahoman, Milwaukee Journal Sentinel, The Columbus Dispatch, The Arizona Republic, The Courier-Journal, The Des Moines Register, Detroit Free Press, The Detroit News, The Palm Beach Post and Star News.
According to the complaint, OpenAI’s WebText dataset, used to help train GPT-2, contained more than 160,000 entries from the plaintiffs’ news brands, including 83,266 from USA TODAY. It also points to a 2019 Common Crawl snapshot containing more than 122 million tokens from USA Today Co publications. The Courier Journal site accounted for 3.1 million of those tokens, according to reporting by WVXU.
The publishers argue that OpenAI’s models can produce verbatim or near-verbatim material when prompted, and that such answers erode the reason for readers to visit the original sites or buy subscriptions. The complaint says the alleged conduct has caused “real and continuing damage” to the publications, and accuses OpenAI of deliberately targeting news content while knowing that its models could reproduce it.
Beyond damages, the suit seeks destruction of training sets and large language models that incorporate the plaintiffs’ content. It also asks for a permanent injunction against using the articles in the future.
The case joins a wider confrontation between publishers and AI companies over training data. The U.S. Copyright Office said in a May 2025 report that generative-AI systems draw on massive collections of data, including copyrighted works, and identified data collection, training, retrieval-augmented generation and outputs as stages where copyright-relevant copying can occur. More than 30 local newspaper publishers owning almost 400 brands have separately sued OpenAI and Microsoft over comparable allegations.
USA Today Co has also been identified as a launch partner for Microsoft’s forthcoming Publisher Content Marketplace, an example of licensing arrangements developing alongside litigation as publishers seek terms for AI companies’ use of news content.