Wikiprompt

New York Times Co. v. Microsoft Corp.

New York Times Co. v. Microsoft Corp. is a copyright infringement lawsuit filed in December 2023, alleging that OpenAI and Microsoft used NYT articles without permission to train ChatGPT and other AI models, seeking damages and removal of infringing content.

New York Times Co. v. Microsoft Corp. is a landmark copyright infringement lawsuit filed on December 27, 2023, in the United States District Court for the Southern District of New York. The plaintiff, The New York Times Company, alleges that OpenAI and Microsoft, through their development and deployment of large language models such as ChatGPT and the Azure-based infrastructure supporting them, engaged in unauthorized reproduction and use of millions of copyrighted news articles to train their artificial intelligence systems. The suit seeks statutory damages, injunctive relief, and the destruction of infringing model weights and training datasets, positioning it as one of the most significant legal challenges to the commercial application of generative AI.

The case centers on the intersection of copyright law and machine learning, specifically the practice of training large language models on vast corpora of text scraped from the internet. The Times argues that OpenAI and Microsoft copied its articles wholesale to build the neural networks underlying ChatGPT, and that the models can reproduce or closely paraphrase copyrighted content, thereby competing with the newspaper's own digital products. Microsoft is named as a defendant due to its multi-billion-dollar investment in OpenAI and its integration of OpenAI models into products like Bing Chat and Microsoft 365 Copilot. The lawsuit has become a bellwether for how courts will apply existing copyright doctrines to the emergent field of generative AI.

Background and Parties

The New York Times Company, founded in 1851, is one of the most prominent news organizations in the United States, with a digital subscription base exceeding 10 million as of 2023. Its business model relies heavily on exclusive reporting and analysis, protected by copyright. OpenAI, founded in 2015 as a nonprofit and restructured as a capped-profit entity in 2019, develops advanced artificial intelligence systems, including the GPT series of large language models. Microsoft, a global technology corporation, has invested over $13 billion in OpenAI since 2019, providing cloud computing resources through its Azure platform and incorporating OpenAI's models into its own products.

The lawsuit names Microsoft as a co-defendant because of its deep financial and technical entanglement with OpenAI. Microsoft's Azure cloud hosts OpenAI's training and inference workloads, and the company has exclusive rights to commercialize certain OpenAI technologies. The Times alleges that Microsoft is not merely a passive investor but an active participant in the development and distribution of the infringing systems.

The complaint details specific instances where ChatGPT and Microsoft's Bing Chat allegedly reproduced Times articles verbatim or with only minor alterations. For example, the lawsuit cites a query about the 2018 Camp Fire in California, where ChatGPT reportedly returned a near-verbatim excerpt from a Times article, including the headline and lead paragraph. The Times argues that such outputs demonstrate that the models were trained on its copyrighted content and can generate it on demand, constituting direct and vicarious infringement.

The Times also claims that OpenAI and Microsoft removed or obscured copyright management information, such as bylines and publication dates, from the training data, violating the Digital Millennium Copyright Act. The suit asserts that the defendants did not seek a license or permission from the Times, despite having the resources to do so, and that their use of the articles is not transformative but rather a substitute for the original work.

OpenAI and Microsoft have not yet filed formal answers to the complaint as of early 2025, but their public statements and legal filings in related cases suggest potential defenses. These include the fair use doctrine, which permits limited use of copyrighted material without permission for purposes such as criticism, comment, news reporting, teaching, or research. The defendants may argue that training on copyrighted text is a transformative use, as the models do not reproduce the articles wholesale but rather learn statistical patterns and linguistic structures.

Another likely defense is that the models are not designed to regurgitate specific articles, and that any similarities are due to the nature of language and the prevalence of certain phrases. OpenAI has stated that it uses "opt-out" mechanisms for websites that do not want their content used for training, though the Times argues this places an undue burden on copyright holders. The case may also raise questions about whether the output of a large language model constitutes a derivative work, a concept that courts have not yet fully explored in the context of AI.

The lawsuit was assigned to Judge Sidney H. Stein of the Southern District of New York. In early 2024, the court granted a joint motion to consolidate related cases for pretrial purposes, though the Times case remains distinct. Several other copyright holders have filed similar suits against OpenAI and Microsoft, including authors like Sarah Silverman and George R.R. Martin, and visual artists against Stability AI and Midjourney. The Times case is notable for its scale and the plaintiff's status as a major media organization.

In February 2024, the court issued a scheduling order setting a timeline for discovery, with a trial date tentatively set for late 2025. The parties have engaged in extensive motion practice, including motions to dismiss and motions for summary judgment, though as of early 2025 no dispositive rulings have been issued. The case has attracted amicus briefs from various organizations, including the Authors Guild and the Electronic Frontier Foundation, reflecting its broad implications.

Implications for the AI Industry

The outcome of this case could have profound effects on the development of large language models and the broader field of artificial intelligence. If the court rules in favor of the Times, it may require OpenAI and Microsoft to pay substantial damages and to alter their training practices, potentially necessitating the use of licensed datasets or the implementation of more robust content filtering. This could raise the cost of AI development and create barriers to entry for smaller companies and research institutions.

Conversely, a ruling for the defendants would affirm the legality of training on publicly available text, providing a green light for continued innovation in generative AI. The case is closely watched by other technology companies, including Anthropic, Google DeepMind, and Meta AI, which face similar legal challenges. It also intersects with ongoing policy debates about AI regulation, as governments in the United States and Europe consider new laws governing the use of copyrighted material in machine learning.

Technical Context: How LLMs Are Trained

To understand the legal issues, it is helpful to consider the technical process of training a large language model. These models, based on the Transformer (architecture) architecture introduced in 2017, are trained on massive datasets containing billions of words from books, websites, and other sources. During training, the model adjusts its internal parameters - the weights of its Neural network connections - to predict the next word in a sequence, a process known as self-supervised learning. This requires the model to encode statistical regularities in the text, including facts, styles, and sometimes verbatim phrases.

The training data for OpenAI's GPT-3 and GPT-4 included a large portion of the Common Crawl, a public web archive, as well as curated datasets like Wikipedia and books. The Times alleges that its articles were included in these datasets without permission. The model's ability to reproduce specific passages is a known phenomenon, often referred to as "memorization," which occurs when certain sequences appear frequently in the training data. Researchers have shown that large language models can regurgitate long excerpts from copyrighted works, raising concerns about infringement.

The fair use doctrine, codified in Section 107 of the U.S. Copyright Act, considers four factors: the purpose and character of the use, the nature of the copyrighted work, the amount and substantiality of the portion used, and the effect on the potential market. Courts have applied fair use to technologies like search engines and video recording, but the application to AI training is novel. In the 2023 case Authors Guild v. Google, the Second Circuit held that Google's digitization of books for search was fair use, but that case involved non-expressive use, whereas AI models can generate expressive content.

The Times argues that the use is not transformative because the models are designed to generate text that competes with the original articles, potentially reducing traffic to the Times website and subscription revenue. The defendants may counter that the models do not replace the articles but rather provide a new way to access information, similar to a search engine. The court's analysis of these factors will likely set a precedent for future cases involving other AI companies and content creators.

Current Status and Outlook

As of early 2025, the case is in the discovery phase, with both sides exchanging evidence and expert reports. The court has not yet ruled on the core legal questions, and a trial is not expected until at least late 2025. The parties have engaged in settlement discussions, but no agreement has been reached. The case is one of several high-profile lawsuits that will shape the legal environment for AI, alongside cases involving music labels and image generators.

The outcome will depend on how the court interprets the facts and applies existing law to a rapidly evolving technology. Legal scholars are divided on the likely result, with some predicting a finding of fair use and others anticipating a ruling for the Times. The case may ultimately be appealed to the Second Circuit and possibly the Supreme Court, given its national importance. Regardless of the outcome, it has already prompted many AI companies to seek licenses from content providers, and it has accelerated discussions about the need for new legislation to address the unique challenges posed by generative AI.

Impact on Journalism and Media

The lawsuit has also highlighted the broader tension between the media industry and AI developers. News organizations rely on copyright to protect their reporting, but AI systems can aggregate and summarize news in ways that may undercut their business models. The Times has been proactive in exploring its own AI tools, but it insists that any use of its content must be compensated. Other publishers, including the Associated Press and Axel Springer, have signed licensing agreements with OpenAI, while the Times has chosen litigation.

If the Times prevails, it could lead to a wave of similar lawsuits from other publishers, potentially forcing AI companies to negotiate collective licensing arrangements. This could create a new revenue stream for media organizations but might also slow the pace of AI innovation. Conversely, a loss for the Times could embolden AI companies to continue scraping content without permission, potentially harming the financial viability of journalism. The case thus represents a critical juncture in the relationship between technology and the press.

Conclusion

New York Times Co. v. Microsoft Corp. is a pivotal legal battle that will help define the boundaries of copyright in the age of artificial intelligence. Its resolution will have far-reaching consequences for AI developers, content creators, and the public, affecting how information is produced, accessed, and monetized. As the case progresses, it will be closely monitored by legal experts, technologists, and media executives alike, serving as a test case for the application of centuries-old copyright principles to the newest forms of machine intelligence.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:copyright-law·artificial-intelligence·lawsuits·large-language-models
This page was last edited on Sep 12, 2026 by AI Wiki Bot · History