# Authors Class Action against OpenAI

A class action lawsuit filed by authors against OpenAI alleging copyright infringement in training large language models, seeking damages and injunctive relief for unauthorized use of copyrighted works.

The Authors Class Action against OpenAI is a consolidated legal proceeding in which a group of authors alleged that OpenAI, the developer of large language models such as GPT-3 and GPT-4, infringed their copyrights by using their books and other written works to train its artificial intelligence systems without permission. The case, filed in the United States District Court for the Southern District of New York, became a focal point for broader debates about the legality of training generative AI on copyrighted text. The plaintiffs sought statutory damages, disgorgement of profits, and an injunction against further unauthorized use, while OpenAI argued that its training practices fell under the fair use doctrine.

The litigation emerged amid a wave of similar lawsuits against AI companies, reflecting growing tension between creative industries and the rapid commercialization of [generative-ai](https://www.wikiprompt.org/wiki/generative-ai). The authors' claims centered on the premise that OpenAI's [large-language-model](https://www.wikiprompt.org/wiki/large-language-model)s, which are built on [transformer](https://www.wikiprompt.org/wiki/transformer) architectures and trained on vast internet corpora, necessarily reproduced and derived from protected expression. Unlike earlier cases involving search engines or image thumbnails, this suit targeted the foundational training process itself, raising novel questions about whether ingesting copyrighted material for machine learning constitutes infringement.

## Background and Filing

The initial complaint was filed on September 8, 2023, by authors including Sarah Silverman, Richard Kadrey, and Christopher Golden, who alleged that OpenAI's training data included their books without authorization. The plaintiffs asserted that OpenAI's [openai](https://www.wikiprompt.org/wiki/openai) models could generate summaries, paraphrases, and even verbatim excerpts of their works, demonstrating that the training process had copied and stored expressive elements. The case was later consolidated with other author suits, including those brought by Michael Chabon, Ta-Nehisi Coates, and comedian Sarah Silverman, under the caption In re OpenAI ChatGPT Litigation.

The authors were represented by the law firm Susman Godfrey, which had previously litigated high-profile intellectual property disputes. The complaint detailed how OpenAI scraped large portions of the internet, including book repositories and pirated text collections, to assemble training datasets. It cited internal OpenAI documents and public statements suggesting that the company recognized the legal risks of using copyrighted material without licenses. The plaintiffs argued that OpenAI's actions were not transformative but rather a commercial exploitation of existing works, as the models competed with the original books in the marketplace.

## Legal Claims and Arguments

The core legal claim was direct copyright infringement under the U.S. Copyright Act, specifically 17 U.S.C. § 501. The plaintiffs alleged that OpenAI violated their exclusive rights to reproduce, prepare derivative works, and distribute copies. They also raised claims for vicarious infringement and unjust enrichment, arguing that OpenAI profited from the unauthorized use of their creative labor. The complaint sought statutory damages of up to $150,000 per work, which, given the number of works allegedly infringed, could have amounted to billions of dollars.

OpenAI's defense rested primarily on the fair use exception, codified in 17 U.S.C. § 107. The company argued that training a model is a transformative use because the model does not reproduce the original text but rather learns statistical patterns and linguistic structures. OpenAI contended that the purpose of training is to create a general-purpose tool, not to supplant the authors' works, and that the output of the models rarely resembles the source material. The company also noted that it had implemented filters to prevent the generation of verbatim copyrighted text, though the plaintiffs disputed the effectiveness of these measures.

A key legal question was whether the intermediate copying of copyrighted works during training constitutes infringement. The plaintiffs drew an analogy to the landmark case of Authors Guild v. Google, where the court held that Google's digitization of books for search indexing was fair use. However, the authors argued that Google's use was limited to making text searchable, whereas OpenAI's use involved the creation of a competing product that could generate new text in the style of the original authors. The distinction between "search" and "generation" became a central point of contention.

## Discovery and Technical Evidence

During the discovery phase, the plaintiffs sought access to OpenAI's training data and model weights, arguing that this evidence was necessary to prove that copyrighted works were actually copied. OpenAI resisted, citing trade secret protections and the enormous computational resources required to produce such data. The court, however, ordered limited discovery, allowing the plaintiffs to inspect metadata and logs that showed which sources were included in the training corpus. This led to revelations that OpenAI had used datasets such as Books3, a collection of pirated books from the shadow library Bibliotik, which contained thousands of copyrighted titles.

The technical evidence centered on the architecture of [neural-network](https://www.wikiprompt.org/wiki/neural-network)s and the nature of [deep-learning](https://www.wikiprompt.org/wiki/deep-learning). The plaintiffs' experts testified that large language models, including those based on the [transformer](https://www.wikiprompt.org/wiki/transformer) architecture, store compressed representations of training examples. They argued that these representations allow the model to reproduce text that is not merely a statistical coincidence but a direct derivation from the original. OpenAI's experts countered that the models are stochastic and that any resemblance to specific texts is a result of the model's ability to generalize, not to memorize. The debate over "memorization" versus "generalization" became a recurring theme in the case.

## Judicial Rulings and Motions

In February 2024, Judge Araceli Martínez-Olguín denied OpenAI's motion to dismiss the case, allowing the copyright claims to proceed. The judge ruled that the plaintiffs had plausibly alleged that OpenAI's training process involved copying protected expression, and that the fair use defense could not be resolved at the pleading stage. This decision was significant because it rejected OpenAI's argument that the case was barred by the statute of limitations or that the plaintiffs lacked standing. The ruling also noted that the plaintiffs had adequately alleged a connection between the training data and the model's outputs, even if the outputs were not verbatim copies.

A subsequent ruling in May 2024 narrowed the scope of the case by dismissing some of the plaintiffs' claims, including those related to state law and unjust enrichment, but preserved the core copyright claims. The court also denied OpenAI's motion to compel arbitration, finding that the plaintiffs had not agreed to any arbitration clause. These rulings set the stage for a potential trial, though both parties indicated a willingness to settle. The case was assigned to the same judge who was overseeing other AI copyright cases, including those against [anthropic](https://www.wikiprompt.org/wiki/anthropic) and [google-deepmind](https://www.wikiprompt.org/wiki/google-deepmind), suggesting that the court might address common legal issues in a coordinated manner.

## Industry and Policy Implications

The lawsuit had far-reaching implications for the [artificial-intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence) industry. It prompted other AI companies, including [anthropic](https://www.wikiprompt.org/wiki/anthropic) and [google-deepmind](https://www.wikiprompt.org/wiki/google-deepmind), to review their training practices and negotiate licensing agreements with publishers and authors. In late 2023, OpenAI announced partnerships with several media companies, including Axel Springer and the Associated Press, to license news content for training. However, these deals covered only a small fraction of the books and articles used in training, leaving many authors uncompensated.

The case also influenced legislative efforts. In the U.S. Congress, lawmakers introduced bills such as the AI Foundation Model Transparency Act, which would require companies to disclose their training data sources. The European Union's AI Act, which was finalized in 2024, included provisions requiring developers of general-purpose AI models to document their training data and to comply with copyright law. The authors' class action served as a test case for whether existing copyright law could effectively regulate AI training, or whether new legislation was needed.

## Reactions from the Creative Community

The lawsuit galvanized authors and other creators, who saw it as a defense of their livelihoods. Organizations such as the Authors Guild filed amicus briefs supporting the plaintiffs, arguing that AI models threatened to devalue human writing and undermine the economic incentives for creative work. Some authors, however, expressed support for AI as a tool for inspiration and collaboration, and a few even licensed their works to AI companies. The divide reflected broader societal debates about the role of automation in creative fields.

Prominent authors who joined the suit included Jonathan Franzen, Jodi Picoult, and George R.R. Martin, though some later withdrew or expressed reservations. The case also drew attention from academics, who debated whether training on copyrighted text is analogous to human learning. Scholars such as [melanie-mitchell](https://www.wikiprompt.org/wiki/melanie-mitchell) and [brendan-lake](https://www.wikiprompt.org/wiki/brendan-lake) argued that machines do not "read" in the human sense, but the legal system had to decide whether the mechanical process of copying and transforming data constitutes infringement.

## Settlement and Aftermath

As of 2025, the case remained unresolved, with both sides engaged in settlement negotiations. Reports indicated that OpenAI had offered to pay a licensing fee to a collective of authors, but the plaintiffs demanded a higher amount and ongoing royalties. The court had scheduled a status conference for late 2025, and observers speculated that a settlement was likely given the high costs of litigation and the potential for an adverse ruling. A settlement would have established a precedent for how AI companies compensate creators, though it would not necessarily bind other defendants.

The case also spawned related litigation. In 2024, a group of nonfiction authors filed a separate class action against OpenAI, alleging that the company had used their works without permission. The New York Times filed its own lawsuit against OpenAI and Microsoft, which was pending in the same court. These cases, along with the authors' class action, created a complex legal landscape that would shape the future of AI development. The outcome of the authors' case was widely expected to influence not only the U.S. but also jurisdictions around the world that were grappling with similar questions.

## Broader Context and Future Outlook

The Authors Class Action against OpenAI was part of a larger movement to hold AI companies accountable for their training data. It highlighted the tension between technological innovation and intellectual property rights, and it forced a public reckoning with the fact that AI systems are built on the unpaid labor of countless writers, artists, and journalists. The case also underscored the need for transparency in AI development, as the plaintiffs' discovery requests revealed the opaque nature of training data collection.

Looking ahead, the legal principles established in this case could have lasting effects. If the court ruled that training on copyrighted works without permission is infringement, AI companies would need to overhaul their data acquisition strategies, potentially relying on public domain works or licensed content. If the court ruled in favor of fair use, it would give AI developers broad latitude to use existing text, but it might also spur legislative action to create a compulsory licensing scheme. Either way, the case represented a pivotal moment in the history of [machine-learning](https://www.wikiprompt.org/wiki/machine-learning) and the commercialization of [generative-ai](https://www.wikiprompt.org/wiki/generative-ai).

The authors' class action also raised questions about the nature of authorship itself. As AI models become capable of producing text that rivals human writing, the line between original creation and derivative work becomes blurred. The case forced courts to consider whether a machine that has "read" millions of books is fundamentally different from a human author who has been influenced by the same books. While the legal system had not yet provided a definitive answer, the case ensured that the question would be debated for years to come.

---
Source: https://www.wikiprompt.org/wiki/authors-class-action
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-12T16:23:56.016833+00:00
