The New York Times Company filed a lawsuit against OpenAI and Microsoft in December 2023, alleging that the companies used millions of copyrighted news articles without permission to train their large language models. The complaint, filed in the U.S. District Court for the Southern District of New York, claims that OpenAI's ChatGPT and Microsoft's Copilot products reproduce Times articles verbatim or with minimal alteration when prompted, undermining the newspaper's ability to monetize its journalism. The Times seeks statutory damages, an injunction against further use of its content, and the destruction of any datasets containing its articles.
The lawsuit is one of the most prominent legal challenges to the generative AI industry, which has relied heavily on web-scraped text for training. The Times argues that its content was used without a license, while OpenAI and Microsoft contend that their use falls under the fair use doctrine, which permits limited reproduction of copyrighted material for purposes such as commentary or research. The case has drawn widespread attention because its outcome could reshape how AI companies obtain training data and whether they must pay publishers for access.
Background and Parties
The New York Times, founded in 1851, is one of the largest and most influential newspapers in the United States, with a digital subscription base exceeding 10 million as of 2023. OpenAI, established in 2015 as a nonprofit and later restructured with a for-profit arm, develops the ChatGPT chatbot and the underlying GPT series of models. Microsoft has invested over $13 billion in OpenAI since 2019 and provides the Azure cloud infrastructure used to train and deploy OpenAI's systems. Microsoft also integrates OpenAI's technology into its Bing search engine and Office productivity suite.
The lawsuit names both OpenAI and Microsoft as defendants, arguing that Microsoft's deep financial and technical involvement makes it equally liable. The Times's legal team, led by the law firm Susman Godfrey, filed the complaint on December 27, 2023, after months of failed negotiations with OpenAI over a potential licensing agreement.
Allegations of Copyright Infringement
The complaint details specific instances where ChatGPT and Microsoft's Copilot allegedly reproduced Times articles. For example, when prompted with a question about a 2019 investigative report on lead poisoning, ChatGPT reportedly generated a near-verbatim excerpt of the article, including the headline and several paragraphs. The Times also presented evidence that the models could produce text matching its restaurant reviews, product recommendations, and news stories, sometimes without attribution.
The Times argues that this reproduction goes beyond acceptable fair use because it competes directly with the newspaper's own digital content. When users ask ChatGPT for news summaries, they may receive information that would otherwise require a Times subscription, reducing traffic to the newspaper's website and diminishing advertising and subscription revenue. The complaint also notes that OpenAI's training data included the Times's articles from sources like Common Crawl, a nonprofit that archives web pages, without obtaining permission.
Legal Arguments and Fair Use Defense
OpenAI and Microsoft filed a motion to dismiss the lawsuit in February 2024, arguing that the Times's claims are based on "hacking" ChatGPT with unusual prompts designed to elicit copyrighted text, rather than typical user behavior. They cited the fair use doctrine, which considers factors such as the purpose of the use (transformative versus commercial), the nature of the copyrighted work, the amount used, and the effect on the market. OpenAI maintains that training on publicly available text is analogous to a human reading and learning from books, and that the resulting models do not store copies of the original articles.
The Times counters that the use is not transformative because the models are designed to generate text that competes with the original articles, and that the commercial nature of OpenAI's products weighs against fair use. The newspaper also points to the "amount and substantiality" factor, noting that the models can reproduce entire articles, not just brief excerpts. Legal scholars have noted that the outcome could hinge on whether courts view AI training as a fair use similar to Google's digitization of books in the Authors Guild v. Google case, or as a more direct infringement.
Broader Context of AI Copyright Litigation
The Times lawsuit is part of a wave of copyright actions against AI companies. In 2023, authors including Sarah Silverman, Christopher Golden, and Richard Kadrey sued OpenAI and Meta over the use of their books in training data. Getty Images filed a suit against Stability AI in the United Kingdom and the United States, alleging misuse of its photographs. In December 2023, the comedian Sarah Silverman and other plaintiffs amended their complaints to include specific examples of ChatGPT reproducing copyrighted text.
These cases have raised questions about the legality of web scraping for AI training, the applicability of fair use to machine learning, and the need for new licensing frameworks. Some publishers, such as the Associated Press and Axel Springer, have signed licensing deals with OpenAI, while others, including the Times, have chosen litigation. The Times's lawsuit is notable for its scale and the quality of its evidence, which includes side-by-side comparisons of model outputs and original articles.
Technological and Economic Implications
The case has significant implications for the AI industry. If the Times prevails, AI companies may be required to obtain licenses for all copyrighted training data, which could increase costs and slow model development. Alternatively, a ruling in favor of OpenAI could reinforce the practice of training on publicly available text without compensation, potentially harming news organizations that rely on subscription revenue.
The economic stakes are high. The Times reported annual revenue of approximately $2.5 billion in 2023, with digital subscriptions accounting for a growing share. The newspaper has invested heavily in its own digital products, including games, cooking, and audio, and argues that AI-generated summaries undermine these investments. OpenAI, valued at over $80 billion in early 2024, has stated that it is willing to pay for content but disputes the Times's claims of widespread infringement.
Procedural History and Current Status
After the motion to dismiss was filed, the court allowed the case to proceed, with discovery scheduled to begin in mid-2024. In April 2024, the Times filed an amended complaint adding new examples of alleged infringement, including instances where ChatGPT reproduced articles from the Times's Wirecutter product review section. OpenAI responded by arguing that the amended complaint still failed to show that the models "memorize" articles in a way that constitutes infringement.
In May 2024, the court denied OpenAI's motion to dismiss in part, allowing the core copyright claims to proceed while dismissing some ancillary claims related to the Digital Millennium Copyright Act. The case is currently in the discovery phase, with both sides exchanging documents and depositions. A trial date has not been set, but legal experts expect the case to last several years, with potential appeals to higher courts.
Reactions from the AI and Publishing Industries
The lawsuit has sparked debate within the AI community. Some researchers argue that training on copyrighted text is essential for building capable models, while others advocate for more transparent data sourcing. The OpenAI CEO Sam Altman has publicly stated that the company is open to licensing agreements and has expressed regret over the lack of clarity in its data practices. Microsoft has remained largely silent, deferring to OpenAI in public statements.
Publishing industry groups, including the News Media Alliance, have filed amicus briefs supporting the Times, arguing that AI companies are free-riding on the work of journalists. Conversely, some technology advocacy groups have warned that a ruling against OpenAI could hinder innovation and give large publishers excessive control over information. The case has also influenced legislative discussions, with some lawmakers proposing laws that would require AI companies to disclose their training data sources.
Potential Outcomes and Precedent
If the Times wins, the court could order OpenAI and Microsoft to pay statutory damages of up to $150,000 per infringed work, which could amount to billions of dollars given the millions of articles allegedly used. The court could also issue an injunction requiring the destruction of training datasets, which would force OpenAI to retrain its models without Times content. Such an outcome would likely lead to a wave of similar lawsuits from other publishers.
If OpenAI wins, the case would establish a broad fair use precedent for AI training, potentially allowing companies to use any publicly available text without compensation. This could accelerate AI development but may also reduce the financial viability of news organizations, which have already seen declining print revenue. Some observers suggest that a settlement is more likely, with OpenAI agreeing to pay the Times a licensing fee in exchange for continued access to its content.
Significance for AI Governance
The lawsuit is a landmark in the broader debate over AI governance and the balance between innovation and intellectual property rights. It raises fundamental questions about whether AI models should be treated as transformative tools or as derivative works, and whether the creators of training data deserve compensation. The case is being closely watched by regulators, including the U.S. Copyright Office, which has been conducting a study on AI and copyright issues.
As of late 2024, the case remains unresolved, with both sides preparing for a lengthy legal battle. The outcome will likely influence not only the future of OpenAI and Microsoft but also the practices of other AI companies such as Anthropic and Google DeepMind, which also rely on large-scale text training. The Times's decision to sue rather than negotiate has made it a test case for the rights of content creators in the age of generative AI.