The New York Times Company filed a lawsuit against OpenAI and Microsoft in December 2023, alleging that the companies used millions of copyrighted news articles to train their large language models without permission or compensation. The case, formally titled The New York Times Company v. Microsoft Corporation and OpenAI, Inc., centers on whether the use of copyrighted text in training generative artificial intelligence systems constitutes fair use under United States copyright law. The outcome could set a precedent for how Artificial intelligence companies handle copyrighted material in their training datasets.
The Times claims that OpenAI's ChatGPT and Microsoft's Azure-powered Copilot products reproduce substantial portions of its articles verbatim or in close paraphrase, often without proper attribution. The complaint cites numerous examples where the models output near-identical passages from Times reporting, including paywalled content. OpenAI has responded by asserting a fair use defense, arguing that training on publicly available text is transformative and does not substitute for the original works. Microsoft, as OpenAI's primary investor and distributor, is also named as a defendant.
Background of the Lawsuit
The lawsuit was filed on December 27, 2023, in the United States District Court for the Southern District of New York. The Times seeks statutory damages and injunctive relief, including the destruction of all model weights and training datasets that incorporated its content. This filing followed months of failed negotiations between the Times and OpenAI over a licensing agreement. The Times had reportedly sought compensation for the use of its archive of over 100 million articles, but the parties could not reach terms.
The case is one of several high-profile copyright disputes involving AI companies. In July 2023, authors including Sarah Silverman and Christopher Golden sued OpenAI and Meta over similar claims. The Times lawsuit is notable because it involves a major news organization with substantial legal resources and a clear commercial interest in protecting its content. The outcome could affect how news-publishers and other content creators interact with AI developers.
OpenAI's Fair Use Argument
OpenAI's defense rests on the doctrine of fair use, codified in Section 107 of the U.S. Copyright Act. The company argues that training a model on copyrighted text is a transformative use - it creates a new function (predicting the next token in a sequence) rather than reproducing the original expression. OpenAI contends that the training process does not store copies of the source text but instead learns statistical patterns and relationships between words. The company points to its model's ability to generate original text on a wide range of topics as evidence that it does not merely regurgitate its training data.
OpenAI also emphasizes the public benefit of its technology. The company argues that large language models enable research, education, and creative expression, and that restricting training on publicly available text would stifle innovation. OpenAI has noted that it provides an opt-out mechanism for website owners who do not want their content used in training, although the Times did not use this mechanism before filing suit. The company further argues that the Times' articles are factual news reports, which receive weaker copyright protection than creative works, and that the models use only the underlying facts and ideas, not the original expression.
The Times' Copyright Claims
The Times counters that fair use does not apply because OpenAI's use is commercial and competes directly with the newspaper's own products. The complaint alleges that ChatGPT and Copilot can generate summaries of Times articles that are detailed enough to substitute for reading the original, thereby diverting traffic and advertising revenue away from the Times website. The newspaper also claims that the models reproduce its articles verbatim in some cases, which goes beyond any transformative use.
The Times points to the "mosaic" nature of its content - individual articles are part of a larger editorial product that includes context, analysis, and investigative reporting. The newspaper argues that reproducing even portions of this content undermines its value. The complaint also notes that OpenAI has licensed content from other publishers, including Axel Springer and the Associated Press, which suggests that the company recognizes the need for permission when using copyrighted material. The Times argues that its refusal to license should not be circumvented through a fair use claim.
Legal Precedents and Interpretations
The fair use analysis in this case will likely draw on several key precedents. In Authors Guild v. Google (2015), the Second Circuit held that Google's digitization of millions of books for its search index was transformative fair use because it provided snippets and search functionality without reproducing the books in full. OpenAI will likely cite this case to argue that its use is similarly transformative. However, the Times may distinguish Google Books by noting that it did not compete with the books themselves, whereas ChatGPT can generate content that competes with news articles.
Another relevant precedent is the 2023 case of Thomson Reuters v. Ross Intelligence, where a court found that using Reuters' legal headnotes to train an AI legal research tool was not fair use. That decision, though not binding in the Second Circuit, suggests that courts may scrutinize AI training more strictly when the output competes with the original work. The Times will likely argue that its case is closer to Ross Intelligence than to Google Books because ChatGPT is designed to answer questions with news content, directly competing with the Times' core business.
Technical Aspects of Training Data
The case also raises technical questions about how Machine learning models store and retrieve information. OpenAI's models are based on the Transformer (architecture) architecture, which uses Multi-Head Attention mechanisms to process text. During training, the model adjusts billions of parameters to minimize prediction error on a massive corpus of text. OpenAI has not disclosed the full contents of its training datasets, but the Times' complaint alleges that its articles appear in the training data based on the models' ability to reproduce specific passages.
OpenAI argues that the model does not contain copies of the text in a traditional sense - the parameters encode statistical relationships, not exact strings. However, researchers have demonstrated that large language models can memorize and reproduce long passages from their training data, especially when those passages appear multiple times or are distinctive. The Times has provided examples of ChatGPT outputting near-verbatim quotes from its articles, which suggests that memorization occurs. The court may need to determine whether this memorization constitutes reproduction under copyright law.
Industry Reactions and Implications
The lawsuit has drawn reactions from across the technology and publishing industries. Some publishers, including News Corp and Gannett, have expressed support for the Times' position, while others have pursued licensing agreements with AI companies. OpenAI has signed deals with several publishers, including Politico's parent company Axel Springer and the Financial Times, to use their content in training and display. These agreements suggest that some publishers see value in partnering with AI companies rather than litigating.
AI researchers have expressed concern that an adverse ruling could impede progress in the field. Many argue that training on large, diverse datasets is essential for building capable models, and that requiring individual licenses for every copyrighted work would be impractical. Some have proposed alternative frameworks, such as a compulsory licensing system or a collective rights organization for text data. The case could also affect international AI development, as other countries watch how U.S. courts balance copyright protection with technological innovation.
Possible Outcomes and Timeline
As of early 2025, the case is in the discovery phase, with both sides exchanging evidence about training data and model outputs. A trial date has not been set, and legal experts expect the case to take several years to resolve, potentially reaching the Supreme Court. The court may issue summary judgment on the fair use question before trial, which could narrow the issues or resolve the case entirely.
Several outcomes are possible. The court could rule that OpenAI's use is fair use, allowing the company to continue training on copyrighted text without permission. Alternatively, the court could find that the use is not fair, potentially requiring OpenAI to pay damages and license content from the Times. A middle ground might involve a ruling that some uses are fair while others are not, depending on the degree of reproduction and the commercial impact. The case could also be settled out of court, as many high-profile copyright disputes are, with OpenAI paying the Times a licensing fee and agreeing to certain restrictions on reproducing its content.
The decision will likely have far-reaching consequences for the AI industry. A ruling against OpenAI could force other companies, including Anthropic and Google DeepMind, to change their training practices. It could also affect the development of open-source models, which rely on publicly available text. Conversely, a ruling in favor of OpenAI could embolden AI companies to continue using copyrighted material without permission, potentially leading to more litigation from content creators. The case represents a fundamental test of how copyright law adapts to new technologies, and its outcome will shape the relationship between AI developers and content producers for years to come.