# NYT v. OpenAI

The New York Times filed a copyright infringement lawsuit against OpenAI and Microsoft in December 2023, alleging unauthorized use of its articles to train large language models. The case is a landmark legal challenge to the foundations of generative AI.

The New York Times Company v. Microsoft Corporation and OpenAI, Inc. is a landmark copyright infringement lawsuit filed in the United States District Court for the Southern District of New York on December 27, 2023. The suit alleges that OpenAI and its primary investor Microsoft used millions of copyrighted news articles from The New York Times to train their [large language models](https://www.wikiprompt.org/wiki/large-language-model) without permission or compensation, thereby violating copyright law and threatening the newspaper's business model. The case is widely considered a pivotal test of how intellectual property law applies to [generative-ai](https://www.wikiprompt.org/wiki/generative-ai) systems and their underlying training data.

The lawsuit seeks to hold the defendants liable for what the Times describes as "unlawful copying and use of The Times's uniquely valuable works." It demands statutory damages up to $150,000 per infringed work, an injunction against further use of Times content, and the destruction of any datasets or models that incorporated that content. The filing came after months of negotiations between the newspaper and OpenAI failed to produce a licensing agreement, according to contemporaneous reports.

## Background and Context

The dispute arises from the rapid advancement of [artificial-intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence) systems, particularly [deep-learning](https://www.wikiprompt.org/wiki/deep-learning) models based on the [transformer](https://www.wikiprompt.org/wiki/transformer) architecture. OpenAI, founded in 2015 as a nonprofit research lab and later restructured with a for-profit arm, released GPT-3 in 2020 and ChatGPT in November 2022. These systems are trained on vast corpora of text scraped from the internet, including news articles, books, and other copyrighted material.

The New York Times, established in 1851, has long maintained a paywall and a digital subscription business, reporting over 10 million subscribers by late 2023. The newspaper argued that its journalism is a valuable, proprietary asset and that allowing AI companies to freely use it undermines its ability to recoup production costs. The Times also noted that AI chatbots could "memorize" and reproduce portions of articles, potentially competing directly with the newspaper's own distribution channels.

## Legal Claims and Arguments

The complaint asserts multiple causes of action, including direct copyright infringement, contributory infringement, and vicarious infringement. The Times provided dozens of examples where ChatGPT and other OpenAI products allegedly reproduced near-verbatim excerpts from its articles, including a 2019 investigation into taxi medallion loans and a 2020 piece on the COVID-19 pandemic. The newspaper argued that such outputs demonstrate the models were trained on its content and that the reproduction is not transformative.

OpenAI and Microsoft responded in court filings that their use of copyrighted material falls under the "fair use" doctrine, which permits limited use for purposes such as criticism, comment, news reporting, teaching, or research. They argued that training on large datasets is a transformative process that does not replace the original works but rather creates new, non-expressive statistical representations. Microsoft, which had invested over $13 billion in OpenAI by early 2023, filed a motion to dismiss in February 2024, claiming the Times had not shown specific harm.

The Times countered that fair use is a narrow exception and that the commercial scale of OpenAI's operations, combined with the direct competition posed by AI-generated summaries, weighs against a fair use finding. Legal scholars noted that the case could set precedent for how courts interpret the "purpose and character" of AI training, a question that remains unresolved in U.S. jurisprudence.

## Key Events and Timeline

- **December 27, 2023**: The Times filed its complaint in the Southern District of New York, naming both OpenAI and Microsoft as defendants.
- **January 2024**: OpenAI publicly stated it was "surprised and disappointed" by the lawsuit, noting that it had been in productive discussions with the Times. The company also began offering an opt-out mechanism for website owners who did not want their content used in future training.
- **February 2024**: Microsoft filed a motion to dismiss, arguing that the Times's claims were speculative and that the company had no direct role in training the models.
- **April 2024**: The court denied Microsoft's motion to dismiss in part, allowing the case to proceed on several claims. Judge Sidney H. Stein, who was assigned to the case, ordered both parties to engage in discovery.
- **June 2024**: OpenAI filed its own motion for summary judgment, asserting that the Times could not prove that specific articles were reproduced in a way that caused market harm. The Times opposed, citing internal OpenAI communications that acknowledged the need to license content.
- **October 2024**: The court held a preliminary hearing on discovery disputes, particularly regarding access to OpenAI's training datasets. The Times sought detailed records of which articles were included in training, while OpenAI argued that such information was a trade secret.
- **As of early 2025**: The case remained in the discovery phase, with no trial date set. Both parties continued to file motions and expert reports.

## Broader Implications

The lawsuit is part of a wave of legal actions against AI companies. In 2023, authors such as Sarah Silverman and George R.R. Martin filed class-action suits against OpenAI and Meta over similar claims. Getty Images sued Stability AI in the UK and US for using its photographs in training data. The Times case is notable because it involves a major news organization with substantial legal resources and a clear financial stake in the outcome.

The outcome could affect the economics of [machine-learning](https://www.wikiprompt.org/wiki/machine-learning) development. If the Times prevails, AI companies may be required to license content from publishers, a cost that could reshape the industry. Conversely, a ruling in favor of OpenAI could legitimize the practice of training on publicly available text, potentially accelerating the growth of [openai](https://www.wikiprompt.org/wiki/openai) and competitors like [anthropic](https://www.wikiprompt.org/wiki/anthropic) and [google-deepmind](https://www.wikiprompt.org/wiki/google-deepmind).

News organizations have responded differently. Some, like the Associated Press and Axel Springer, have signed licensing deals with OpenAI. Others, including the Times, have chosen litigation. The case has also prompted discussions about the need for new legislation, with some lawmakers proposing a "notice and opt-out" regime for AI training.

## Technical and Ethical Dimensions

The case raises questions about how [neural networks](https://www.wikiprompt.org/wiki/neural-network) store and retrieve information. Researchers have shown that large language models can memorize long passages of text, especially when those passages appear multiple times in training data. The Times argued that its articles, which are widely syndicated and quoted, are likely to be overrepresented in OpenAI's datasets, increasing the risk of verbatim reproduction.

Ethical debates center on whether AI companies have a moral obligation to compensate content creators. Proponents of fair use argue that training is analogous to a human reading and learning from articles, while critics contend that the scale and commercial nature of AI training are fundamentally different. The case has also highlighted the lack of transparency in AI training, as most companies do not publicly disclose their full data sources.

## Reactions and Commentary

Legal experts have offered divergent predictions. Some, like Stanford Law professor Mark Lemley, have argued that AI training is likely to be deemed fair use, citing precedents such as the Google Books case. Others, including copyright scholar Pamela Samuelson, have noted that the Times's evidence of direct competition could tip the scales. The [MIT Computer Science and Artificial Intelligence Laboratory](https://www.wikiprompt.org/wiki/mit-csail) and [Stanford AI Lab](https://www.wikiprompt.org/wiki/stanford-ai-lab) have both hosted symposia on the case, reflecting its academic significance.

OpenAI CEO Sam Altman has publicly stated that the company is willing to pay for content but believes that the current legal framework is unclear. In a March 2024 interview, he said, "We want to be good partners with publishers. The question is what a fair price is." The Times, for its part, has maintained that its primary goal is to protect its journalism and set a precedent for the industry.

## Current Status and Future Outlook

As of late 2024, the case had not been resolved. Both sides have engaged in extensive motion practice, and the court has yet to rule on the core fair use question. Some observers expect the case to reach the Supreme Court, given its national importance. In the meantime, OpenAI has continued to sign content deals with other publishers, including News Corp and the Financial Times, suggesting that the industry is moving toward a licensing model regardless of the lawsuit's outcome.

The Times has also pursued parallel actions, including a complaint to the UK's Information Commissioner's Office about data protection issues. The company has stated that it will continue to invest in its own AI capabilities, including a partnership with [google-cloud](https://www.wikiprompt.org/wiki/google-cloud) for data analytics, but remains committed to the litigation.

## Conclusion

The New York Times v. OpenAI case represents a critical juncture in the relationship between media and artificial intelligence. Its resolution will likely influence not only the parties involved but also the broader ecosystem of [artificial-intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence) development, content licensing, and copyright law. As the technology evolves, the legal principles established in this case may serve as a foundation for future disputes, making it one of the most closely watched lawsuits of the decade.

---
Source: https://www.wikiprompt.org/wiki/ny-times-v-openai
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-13T03:52:02.302003+00:00
