The New York Times v. OpenAI is a copyright infringement lawsuit filed by The New York Times Company against OpenAI and Microsoft in December 2023. The case centers on the alleged unauthorized use of millions of copyrighted newspaper articles to train large language models powering ChatGPT and other products. The newspaper seeks billions in damages and demands the destruction of AI models and datasets that contain or derive from its content. As of 2025, the litigation remains ongoing in the United States District Court for the Southern District of New York, with a trial date set for late 2025 or early 2026.
Background: Generative AI and Copyright
The rise of Generative AI based on large language models created a new frontier in copyright law. These systems, built on transformer architectures, ingest vast corpora of text scraped from the internet, including news articles, books, and websites. The OpenAI's ChatGPT, first released in November 2022, demonstrated the ability to generate human-like responses, often reproducing or closely paraphrasing copyrighted text. This behavior raised legal questions about whether training AI on copyrighted material constitutes fair use or actionable infringement. Several lawsuits followed from authors and rightsholders, but the Times's case became the most prominent due to its scale and the involvement of major technology corporations.
Parties and Legal Claims
The plaintiff, The New York Times Company, is a major US newspaper with a history of vigorous copyright enforcement. The defendants include OpenAI and its commercial partner Microsoft, which provided cloud infrastructure and exclusive cloud hosting for OpenAI's products. Also named is Microsoft's GitHub Copilot, although the primary focus remains on ChatGPT and the underlying models.
In its complaint filed at the Southern District of New York on December 27, 2023, the Times asserted several claims: direct copyright infringement, contributory infringement, vicarious infringement, and unfair competition under the Lanham Act. The Times alleged that OpenAI and Microsoft copied millions of its articles without permission, used them to train models, and enabled the models to generate outputs that "memorize," "regurgitate," or paraphrase the original text, often with attribution removed. The complaint included dozens of examples where ChatGPT produced near-verbatim excerpts from Times articles, sometimes including entire paragraphs.
The newspaper sought statutory damages up to $150,000 per infringed work, which could amount to billions given the number of articles at issueians. It also demanded the deletion or destruction of all training data and models that incorporated its copyrighted content, as well as an injunction preventing future use.
OpenAI and Microsoft's Response
OpenAI and Microsoft filed motions to dismiss in February 2024, arguing that the use of copyrighted material for training constituted fair use under US copyright law. They contended that the models transform the content into new, non-expressive structures of statistical relationships, and that the outputs at issue were either lawful summaries or the result of user prompts that elicited the reproduction. They also argued that the Times's claims of "memorization" were overstated and that the company had not proven actual harm.
In their legal briefs, the defendants noted that the Times had not suffered a decline in readership specifically due to AI, and that the newspaper itself had experimented with using AI-generated content. They pointed to the industry practice of web scraping as customary. The motions to dismiss were partially successful: in September 2024, Judge Sidney H. Stein dismissed claims related to vicarious infringement but allowed the direct infringement and contributory claims to proceed. The Lanham Act claim for false endorsement was also dismissed, with the court suggesting the Times refile.
Key Legal Issues and Precedents
The case touches on several unresolved doctrines. The central question is whether training on copyrighted works is transformative under the fair use four-factor test, which examines the purpose of use, the nature of the work, the amount used, and the effect on the market. Courts have historically held that some types of digital reproduction, such as Google's book scanning project, qualify as fair use. However, the Times argues that generative AI differs because its outputs can directly compete with the original articles, potentially reducing subscription revenue and licensing opportunities.
Another issue is the "memorization" phenomenon. Research has shown that large language models can recreate training data when prompted, especially for obscure or repeated text. The Times documented instances where ChatGPT reproduced its articles almost verbatim, suggesting that the model had not fully transformed the content. OpenAI maintained that such outputs are rare and that they have implemented technical measures to prevent lengthy reproductions.
Broader Impact on the AI Industry
The lawsuit has significant implications for the broader AI ecosystem. Many companies, including Anthropic, Google DeepMind, and Apple, have faced or are facing similar suits or have entered into licensing agreements with publishers. The outcome could set a precedent for whether AI companies must obtain permission and pay for training data or can continue scraping the web freely.
In response to the litigation, several media organizations have signed deals with OpenAI and other companies. For example, the Associated Press, Axel Springer, and News Corp have reportedly reached licensing agreements, while others like National Public Radio and the Chicago Tribune have explored alternatives. The case may accelerate the trend toward content licensing, as AI firms seek to reduce legal risk. Conversely, a ruling in favor of the Times could force a fundamental shift in how training datasets are compiled, potentially raising costs for startups and smaller players.
Legal Proceedings and Timeline
The case has progressed through pre-trial phases. In 2024, the parties engaged in extensive discovery, with the court ordering OpenAI to produce training data logs and internal communications. The Times also obtained records showing that OpenAI employees had concerns about using copyrighted material without permission. In November 2024, the court denied a motion from the Times to include additional claims related to Microsoft's Copilot, but allowed amendments to the complaint.
A notable development occurred in February 2025 when Judge Stein approved a scheduling order setting trial for December 15, 2025. He also denied a request by OpenAI to delay the trial pending a related case in the Second Circuit, which involves a separate author lawsuit. The trial is expected to last several weeks and will likely feature expert testimony on AI technology and economics.
Economic and Strategic Considerations
For the Times, the suit is not just about damages. It represents a defense of its journalism and a push to establish a licensing market for news content used by AI. The newspaper has argued that AI chatbots could replace traditional search and direct news consumption, undermining its subscription model. It also values its brand and expects attribution when its reporting is used.
OpenAI, meanwhile, has invested heavily in legal defense, hiring prominent law firms and outside experts. The company has also engaged in public outreach, publishing papers on copyright and fair use. Microsoft, as a major shareholder and infrastructure provider, has its own legal team but has largely deferred to OpenAI's strategy. The financial stakes are high: if the Times wins, it could claim billions in statutory damages, and a verdict against the defendants might trigger a wave of follow-on suits from other publishers.
Industry Reactions and Constitutional Context
Reactions have been divided. Many scholars and AI researchers argue that training on copyrighted text is necessary for advancing the field and that over-restrictive rulings would hinder innovation. Others, including journalists and authors, support the Times, viewing the case as essential for protecting creative livelihoods. The US Copyright Office has weighed in, issuing guidelines that acknowledge the complexity, but has refrained from making definitive rulings on AI training.
The case also touches on First Amendment considerations, as AI-generated content relies on ingesting news to provide information. The Supreme Court has not yet addressed these issues, leaving district courts to navigate uncharted territory. Some have drawn parallels to the landmark cases of the internet era, such as Sony v. Universal and Authors Guild v. Google, which balanced technological progress with copyright protection.
What Is at Stake
If the court rules that the use of the Times's articles is not fair use, it could order the destruction of certain OpenAI models or require a licensing scheme. Such an outcome might force AI companies to rely on public domain or explicitly licensed data, slowing development but potentially creating a new business for content creators. Conversely, a fair use ruling would validate the current practices of many AI labsasi, although it might not preclude voluntary agreements.
As of late 2025, the trial has not yet begun, and both sides have signalled openness to settlement. OpenAI has licensed content from other publishers, suggesting a willingness to negotiate. However, the Times has stated its commitment to litigating, arguing that licensing alone would not resolve the "existential threat" posed by AI. The case remains one of the most watched in the intersection of law and technology, with implications for the future of journalism and the digital economy.