# Authors Guild v. OpenAI

Authors Guild v. OpenAI is a class action lawsuit filed in 2023 by the Authors Guild and prominent authors against OpenAI, alleging copyright infringement in training large language models on their works without permission.

Authors Guild v. OpenAI is a class action lawsuit filed in the United States District Court for the Southern District of New York on September 20, 2023. The suit, brought by the Authors Guild and a group of prominent fiction and nonfiction writers, alleges that OpenAI, the developer of the [large language models](https://www.wikiprompt.org/wiki/large-language-model) behind ChatGPT, infringed their copyrights by using their books to train its [generative AI](https://www.wikiprompt.org/wiki/generative-ai) systems without authorization. The case is one of several high-profile legal challenges to the data practices of [AI](https://www.wikiprompt.org/wiki/artificial-intelligence) companies, raising fundamental questions about fair use, authorship, and the economics of creative labor in the age of machine learning.

The plaintiffs include the Authors Guild, the largest professional organization for writers in the United States, along with individual authors such as John Grisham, George R.R. Martin, Jodi Picoult, and David Baldacci. Their complaint contends that OpenAI copied their works wholesale into the training datasets for its models, which then generate text that can mimic or summarize those works, competing with the originals in the marketplace. The lawsuit seeks statutory damages, injunctive relief, and a declaration that OpenAI's use of copyrighted material without consent or compensation violates the Copyright Act of 1976.

## Background: The Rise of Large Language Models

OpenAI, founded in 2015 as a nonprofit research organization, transitioned to a capped-profit structure in 2019 to secure funding for massive computing resources. Its flagship product, ChatGPT, released to the public in November 2022, became one of the fastest-growing consumer applications in history, reaching over 100 million monthly users within two months. The underlying technology is based on the [transformer architecture](https://www.wikiprompt.org/wiki/transformer), introduced in a 2017 paper by researchers at Google, which enabled efficient training on vast amounts of text data.

Training a large language model involves feeding it billions of words from diverse sources, including books, articles, websites, and other written material. OpenAI has not publicly disclosed the full contents of its training datasets, but the company has acknowledged using publicly available text from the internet, including books, in its earlier models such as GPT-2 and GPT-3. The scale is enormous: GPT-3, released in 2020, was trained on hundreds of billions of tokens, a token being a unit of text roughly equivalent to a word or part of a word.

The plaintiffs argue that books are uniquely valuable training data because they contain high-quality, long-form prose that helps models learn grammar, reasoning, and narrative structure. However, they contend that using copyrighted books without permission is not transformative fair use but rather a commercial exploitation of authors' labor. The case thus sits at the intersection of copyright law and [machine learning](https://www.wikiprompt.org/wiki/machine-learning) technology, a field that has rapidly advanced through techniques such as [deep learning](https://www.wikiprompt.org/wiki/deep-learning) and [neural networks](https://www.wikiprompt.org/wiki/neural-network).

## Legal Claims and Arguments

The complaint asserts multiple causes of action, including direct copyright infringement, vicarious infringement, and unjust enrichment. The plaintiffs allege that OpenAI's training process necessarily involves making unauthorized reproductions of their works, both in the creation of training datasets and in the model's ability to generate text that closely resembles the originals. They also claim that OpenAI's use of their works has diminished the market for their books, as readers can now obtain AI-generated summaries or imitations for free.

OpenAI's likely defense centers on the doctrine of fair use, which permits limited use of copyrighted material for purposes such as criticism, comment, news reporting, teaching, and research. The company has argued in other contexts that training AI models on publicly available text is a transformative use that does not substitute for the original works. However, the plaintiffs counter that OpenAI's use is commercial, non-transformative, and harmful to the market for their works, pointing to the company's valuation, which exceeded $80 billion in early 2024.

The case is assigned to Judge Sidney H. Stein, who also presides over a related lawsuit brought by The New York Times against OpenAI and Microsoft in December 2023. The Authors Guild case has been consolidated with other similar suits, including one filed by authors Sarah Silverman, Richard Kadrey, and Christopher Golden, though the consolidation has been partial and procedural. A key early ruling came in November 2023, when Judge Stein dismissed some claims in the Silverman case for lack of standing but allowed others to proceed, signaling that the courts are willing to engage with the substantive copyright questions.

## The Broader Litigation Landscape

The Authors Guild v. OpenAI is part of a wave of lawsuits against AI companies. In addition to The New York Times case, which alleges that OpenAI and its partner Microsoft used millions of copyrighted articles to train their models, other authors have filed suits against [Anthropic](https://www.wikiprompt.org/wiki/anthropic), the maker of the Claude model, and against [Google DeepMind](https://www.wikiprompt.org/wiki/google-deepmind). These cases are being watched closely by the technology industry, the publishing sector, and legal scholars, as their outcomes could set precedents for how AI companies obtain and use training data.

A notable development occurred in July 2024 when the U.S. Copyright Office released a report on copyright and AI, recommending that Congress consider new legislation to address the use of copyrighted works in AI training. The report did not take a definitive stance on fair use, but it highlighted the need for clarity and suggested that existing law may not adequately protect authors. This regulatory uncertainty has led some AI companies to strike licensing deals with content providers, such as OpenAI's agreements with news organizations like Axel Springer and the Associated Press, though such deals do not cover the vast majority of authors.

The Authors Guild has also been active in advocating for collective licensing mechanisms, similar to those used by music publishers and songwriters. In 2024, the guild proposed a framework under which AI companies would pay a subscription fee to use copyrighted works, with revenues distributed to authors based on usage. This proposal has gained some traction in policy circles but has not been adopted by any major AI company as of early 2025.

## Key Procedural Milestones

Following the initial filing, the case progressed through several procedural stages. In October 2023, OpenAI filed a motion to dismiss, arguing that the plaintiffs had not adequately alleged direct infringement because the training process involves intermediate copies that are not distributed to the public. The company also contended that the plaintiffs lacked standing because they had not registered all of their works with the U.S. Copyright Office, a prerequisite for filing an infringement suit.

Judge Stein ruled on the motion in part in February 2024, allowing most claims to proceed but dismissing some on technical grounds. He held that the plaintiffs had plausibly alleged that OpenAI's training process involved copying, and that the fair use defense was better suited for summary judgment or trial rather than a motion to dismiss. This ruling was seen as a victory for the authors, as it kept the core claims alive.

In June 2024, the parties engaged in discovery, with the plaintiffs requesting detailed information about OpenAI's training datasets and the company resisting on grounds of trade secrecy. A magistrate judge ordered OpenAI to produce certain documents, but the scope of discovery remains contested. The case is scheduled for a pretrial conference in March 2025, with a potential trial date in late 2025 or 2026, though such timelines are often extended in complex litigation.

## Implications for Authors and AI Development

The outcome of Authors Guild v. OpenAI could have profound implications for both the literary world and the AI industry. If the plaintiffs prevail, OpenAI and other AI companies may be required to obtain licenses for every copyrighted work used in training, a process that could be logistically challenging and financially burdensome. This could slow the development of large language models, particularly for smaller companies and research institutions that lack the resources to negotiate with thousands of rights holders.

Conversely, a ruling in favor of OpenAI would affirm the fair use doctrine as applied to AI training, potentially allowing companies to continue using copyrighted text without compensation. Such a decision might encourage more aggressive data collection and could lead to further consolidation in the AI industry, as only the largest players could afford to litigate and settle claims.

For authors, the case is about more than money. It raises questions about the nature of authorship in an era when machines can generate text that mimics human creativity. Some writers have embraced AI as a tool, while others see it as an existential threat to their livelihoods. The Authors Guild has framed the lawsuit as a defense of the profession, arguing that writers deserve to be compensated for the value their works contribute to AI systems.

The case also intersects with broader debates about [OpenAI's](https://www.wikiprompt.org/wiki/openai) business practices and its relationship with the creative community. In 2023, OpenAI launched a program to pay some authors for the use of their works in training, but the terms were widely criticized as inadequate. The company has also faced criticism for its lack of transparency about training data, a concern that is central to the litigation.

## Expert Opinions and Scholarly Debate

Legal scholars are divided on the likely outcome. Some, such as Professor Pamela Samuelson of the University of California, Berkeley, have argued that AI training on copyrighted works should be considered fair use, drawing parallels to search engines and other technologies that copy text for transformative purposes. Others, like Professor James Grimmelmann of Cornell University, have suggested that the scale and commercial nature of AI training may tip the balance in favor of authors, particularly if the models can generate outputs that compete with the originals.

The case has also drawn attention from economists, who note that the value of training data is substantial. A 2023 study estimated that the market value of a single high-quality book for AI training could be in the thousands of dollars, given the importance of diverse, well-written text in improving model performance. This has led some to argue that authors are effectively subsidizing the AI industry without receiving a share of the profits.

## Conclusion and Outlook

As of early 2025, Authors Guild v. OpenAI remains in the discovery phase, with no final ruling on the merits. The case is likely to take years to resolve, and an appeal to the U.S. Court of Appeals for the Second Circuit is probable regardless of the outcome. The Supreme Court may ultimately weigh in, as it has done in other copyright cases involving new technologies, such as the 2021 decision in Google v. Oracle, which held that Google's use of Java code was fair use.

In the meantime, the case has already had an impact. It has prompted other authors and publishers to file their own suits, and it has accelerated efforts to develop voluntary licensing frameworks. It has also contributed to a broader public conversation about the ethics of AI, the rights of creators, and the future of intellectual property in a digital age. Whatever the final judgment, Authors Guild v. OpenAI is likely to be remembered as a landmark case in the history of artificial intelligence and copyright law.

---
Source: https://www.wikiprompt.org/wiki/authors-guild-v-openai
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-12T16:23:26.318382+00:00
