# NYT v. OpenAI (2023)

NYT v. OpenAI (2023) is a copyright infringement lawsuit filed by The New York Times against OpenAI and Microsoft in December 2023, alleging unauthorized use of millions of articles to train AI models. The case seeks damages and destruction of training data, raising pivotal questions about fair use in generative AI.

NYT v. OpenAI (2023) is a landmark copyright infringement lawsuit filed by The New York Times Company against OpenAI and its primary investor Microsoft on December 27, 2023, in the U.S. District Court for the Southern District of New York. The complaint alleges that OpenAI and Microsoft used millions of copyrighted news articles without permission to train their [large language models](https://www.wikiprompt.org/wiki/large-language-model), including GPT-4 and ChatGPT, and that these models now reproduce and "memorize" Times content in ways that compete directly with the newspaper's own products. The suit seeks statutory damages up to $150,000 per infringed work, an injunction against further use of Times content, and the destruction of all training datasets containing such material.

The case is widely seen as a bellwether for the broader conflict between the [generative AI](https://www.wikiprompt.org/wiki/generative-ai) industry and content creators, as it tests whether the "fair use" doctrine extends to mass scraping of copyrighted text for model training. The outcome could reshape the economics of [AI](https://www.wikiprompt.org/wiki/artificial-intelligence) development, forcing companies to negotiate licensing agreements or redesign training pipelines. As of early 2025, the case remains in pretrial litigation, with both sides filing motions to dismiss and counterclaims.

## Background: The Rise of Generative AI and Data Scraping

The lawsuit emerged against the backdrop of rapid advances in [deep learning](https://www.wikiprompt.org/wiki/deep-learning) and the commercial deployment of [transformer](https://www.wikiprompt.org/wiki/transformer)-based models. OpenAI, founded in 2015, released GPT-3 in 2020 and ChatGPT in November 2022, which quickly became one of the fastest-growing consumer applications in history. These models are trained on vast corpora of text scraped from the public internet, including news articles, books, and academic papers. The Times alleges that its content, which is largely paywalled, appeared in these datasets without authorization, often in quantities that made it a significant portion of the training material.

Microsoft, which invested over $13 billion in OpenAI and integrated its models into products like Bing Chat and Office 365, is named as a co-defendant for its role in deploying and profiting from the allegedly infringing systems. The complaint cites specific examples where ChatGPT produced near-verbatim excerpts from Times articles, including a 2019 investigation into the taxi industry and a 2020 piece on mortgage interest rates, demonstrating that the models can reproduce copyrighted text on demand.

## Legal Claims and Demands

The Times asserts several causes of action, including direct copyright infringement, contributory infringement, and vicarious infringement. It argues that OpenAI and Microsoft copied, stored, and distributed its works during training and inference, and that the models' outputs constitute unauthorized derivative works. The complaint emphasizes that the defendants "built their business on the back of others' work" and that the Times spent billions of dollars on journalism that the AI companies appropriated for free.

Beyond monetary damages, the Times requests a permanent injunction that would require OpenAI and Microsoft to delete all copies of Times content from their training datasets and to prevent future use. This demand is unprecedented in scale, as it would effectively force the companies to retrain their models from scratch, a process costing hundreds of millions of dollars and potentially degrading model quality. The Times also asks for an order requiring the defendants to disclose their training data sources and to implement technical measures to prevent future infringement.

## OpenAI's Defense: Fair Use and Public Interest

OpenAI filed its response in February 2024, arguing that its use of copyrighted material falls under the fair use exception, which permits limited reproduction for purposes such as criticism, comment, news reporting, teaching, and research. The company contends that training AI models is a transformative use that does not substitute for the original works, and that the models learn patterns and facts rather than memorizing specific texts. OpenAI also points to its "opt-out" mechanism, which allows website owners to block its web crawler, and to its ongoing partnerships with other publishers, such as the Associated Press and Axel Springer, as evidence of good-faith efforts to respect intellectual property.

Microsoft echoed these arguments, emphasizing that the models are tools for generating new content, not for reproducing existing articles. The defendants also note that the Times itself used AI tools for various purposes and that the newspaper's complaint is an attempt to stifle competition. Legal scholars are divided on the strength of these defenses, with some arguing that the scale of copying and the commercial nature of the use weigh against fair use, while others point to the transformative nature of machine learning as a compelling justification.

## Industry Reactions and Parallel Lawsuits

The Times lawsuit is part of a wave of litigation against AI companies. In 2023, authors including Sarah Silverman, Christopher Golden, and Richard Kadrey filed class-action suits against OpenAI and Meta, while comedian Sarah Silverman's case was later consolidated with others. Getty Images sued Stability AI in the U.K. and the U.S. over the use of its photographs in training image generators. In 2024, the Authors Guild filed a class action against OpenAI, and several newspapers, including the Chicago Tribune and the New York Daily News, joined a separate suit against OpenAI and Microsoft.

These cases have prompted a range of responses from the industry. Some companies, such as [Anthropic](https://www.wikiprompt.org/wiki/anthropic) and [Google DeepMind](https://www.wikiprompt.org/wiki/google-deepmind), have begun licensing content from publishers, while others have argued for a statutory licensing scheme similar to that used for music streaming. The U.S. Copyright Office launched a public inquiry into AI and copyright in 2023, and the European Union passed the AI Act in 2024, which includes transparency requirements for training data. The outcome of the Times case could set a precedent that influences all of these efforts.

## Key Legal Questions and Precedents

The central legal question is whether training AI models on copyrighted text constitutes fair use. Courts have historically applied a four-factor test: the purpose and character of the use, the nature of the copyrighted work, the amount and substantiality of the portion used, and the effect on the potential market. In the 2015 case Authors Guild v. Google, the Second Circuit found that Google's digitization of books for search was fair use, but that case involved limited snippets and a non-commercial purpose. The Times argues that its articles are highly creative and expressive, that the defendants copied entire works, and that ChatGPT's ability to produce summaries and excerpts directly harms the Times' licensing revenue and subscription business.

Another key issue is the "memorization" problem. OpenAI has acknowledged that large language models can sometimes reproduce training data verbatim, especially when prompted in certain ways. The Times' complaint includes examples of such outputs, which the company argues are not transformative but rather a form of wholesale copying. The defendants counter that these instances are rare and that the models are designed to generalize, not to memorize. Expert witnesses are expected to testify on both sides about the technical details of model training and the likelihood of reproduction.

## Procedural History and Current Status

After the complaint was filed, the court assigned the case to Judge Sidney H. Stein. In early 2024, OpenAI and Microsoft filed motions to dismiss, arguing that the Times had not adequately alleged direct infringement because it did not identify specific infringing outputs. The Times amended its complaint in March 2024, adding more examples and clarifying its claims. In July 2024, Judge Stein denied the motions to dismiss, allowing the case to proceed to discovery. This ruling was seen as a significant victory for the Times, as it rejected the defendants' argument that the case was too speculative.

Discovery is expected to be contentious, with the Times seeking access to OpenAI's training data and internal communications about data sourcing. OpenAI has resisted some requests, citing trade secrets and the burden of producing massive datasets. The court has also ordered the parties to discuss a potential settlement, but no agreement has been reached as of early 2025. A trial date has not been set, but legal observers estimate it could begin in late 2025 or 2026.

## Broader Implications for AI Development

The case has implications beyond the two parties. If the Times wins, AI companies may be required to obtain licenses for all copyrighted training data, which could increase costs and slow innovation. Some startups may shift to using only public domain or synthetic data, while larger firms may invest in proprietary datasets or partnerships with publishers. The case could also affect the development of open-source models, which rely heavily on scraped data and may not have the resources to comply with licensing requirements.

Conversely, if OpenAI prevails, it could legitimize the practice of training on copyrighted material without permission, potentially leading to more aggressive scraping and less compensation for creators. This outcome might accelerate the trend toward AI-generated content, putting further pressure on traditional media companies. The case is therefore closely watched by technologists, lawyers, and journalists alike, as it will help define the boundaries of intellectual property in the age of machine learning.

## Conclusion

NYT v. OpenAI (2023) is a defining legal battle for the generative AI era. It raises fundamental questions about authorship, ownership, and the value of human creativity in a world where machines can produce text at scale. The resolution of the case, whether through judicial ruling or settlement, will likely shape the future of AI research, media economics, and copyright law for decades. As the litigation unfolds, it serves as a reminder that the rapid advancement of [neural networks](https://www.wikiprompt.org/wiki/neural-network) and [machine learning](https://www.wikiprompt.org/wiki/machine-learning) is not just a technical story, but a deeply human one about who benefits from and who bears the costs of technological progress.

---
Source: https://www.wikiprompt.org/wiki/nytimes-v-openai-2023
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-12T16:23:52.775337+00:00
