The Authors Class Actions Against OpenAI are a set of consolidated United States federal lawsuits filed by fiction writers against OpenAI, the developer of the large language models underlying products such as ChatGPT. The plaintiffs allege that OpenAI copied their copyrighted books without permission to train its models, constituting direct and indirect copyright infringement, and that the company profited from that unauthorized use. The litigation, consolidated in the U.S. District Court for the Southern District of New York, has become a central test case for how copyright law applies to the training of generative AI systems.
The cases emerged in the wake of the commercial release of ChatGPT in late 2022 and the subsequent expansion of OpenAI's model family. Authors argued that their works were included in massive text corpora used for training, without licensing or compensation, and that the models could reproduce or closely paraphrase passages from their books. OpenAI responded that its use of copyrighted material fell under the fair use doctrine, citing the transformative nature of machine learning and the non-expressive statistical patterns learned by the models.
Procedural History and Consolidation
The first major author lawsuit was filed in June 2023 by comedian and author Sarah Silverman, along with authors Richard Kadrey and Christopher Golden, in the Northern District of California. That case, Silverman v. OpenAI, alleged that OpenAI's training data included pirated book repositories. In July 2023, a separate class action was filed in the Southern District of New York by authors Paul Tremblay and Mona Awad, represented by the same law firm, against OpenAI and its partner Microsoft.
In September 2023, the Judicial Panel on Multidistrict Litigation consolidated several overlapping author suits into a single proceeding, In re OpenAI ChatGPT Litigation, assigned to Judge Jed S. Rakoff in the Southern District of New York. The consolidated action included claims from authors such as George R.R. Martin, John Grisham, Jodi Picoult, and others who joined later, though some authors filed separate suits that were later folded into the main case. By early 2024, the lead plaintiffs filed a consolidated amended complaint, adding claims for unjust enrichment and negligence, and naming OpenAI and its CEO Sam Altman as defendants.
Core Legal Claims
The plaintiffs' primary claim is direct copyright infringement under 17 U.S.C. § 501, arguing that OpenAI reproduced and distributed their works during training and in model outputs. They also alleged vicarious and contributory infringement, asserting that OpenAI had the right and ability to control the infringing activity and profited from it. The complaint cited specific instances where ChatGPT allegedly generated near-verbatim excerpts from copyrighted books, which the plaintiffs argued demonstrated that the models stored and retrieved expressive content.
A second major claim concerned the removal of copyright management information (CMI), such as author names and titles, from training data, in violation of the Digital Millennium Copyright Act (DMCA). The plaintiffs argued that stripping CMI facilitated infringement and hindered authors' ability to enforce their rights. OpenAI moved to dismiss this claim, arguing that the DMCA provision required intent to induce infringement, which the plaintiffs had not plausibly alleged.
Key Rulings
In February 2024, Judge Rakoff dismissed the DMCA claim, finding that the plaintiffs had not adequately alleged that OpenAI intentionally removed CMI to conceal infringement. He allowed the core copyright claims to proceed, rejecting OpenAI's argument that the claims were preempted by state law and that the plaintiffs lacked standing because they had not registered every work.
In November 2024, Judge Rakoff issued a pivotal ruling on fair use. He denied OpenAI's motion to dismiss the direct infringement claims, holding that the fair use defense could not be resolved at the pleading stage. The court emphasized that whether the use was transformative, and whether the models competed with the original works, were questions of fact requiring discovery. The ruling also noted that OpenAI's own documents, including internal communications about the scarcity of high-quality text data, could bear on the purpose and character of the use.
However, in a separate ruling in the same month, Judge Rakoff dismissed the unjust enrichment and negligence claims, finding that they were duplicative of the copyright claims and barred by the economic loss rule. He also struck certain allegations that relied on speculative theories of harm, such as the claim that the models' outputs would reduce book sales.
Fair Use Arguments
OpenAI's fair use defense rests on four statutory factors: the purpose and character of the use, the nature of the copyrighted work, the amount and substantiality used, and the effect on the potential market. The company argued that training a neural network is a non-expressive use, because the model learns statistical patterns rather than reproducing specific expressions. It cited the Supreme Court's decision in Google v. Oracle (2021), which held that copying code for a transformative purpose was fair use, and the Second Circuit's ruling in Authors Guild v. Google (2015), which found that Google's book scanning for search was fair use.
The plaintiffs countered that OpenAI's use was commercial and that the models could generate text that competes with the original works, harming the market for licensed reproductions. They also argued that the scale of copying - billions of words from millions of books - was not analogous to the limited snippets in the Google Books case. The court's November 2024 ruling left these issues for trial, noting that discovery would be needed to assess the actual outputs and the economic impact.
Related Litigation and Context
The authors' cases are part of a broader wave of copyright lawsuits against AI companies. In 2023, visual artists filed class actions against Stability AI, Midjourney, and DeviantArt, and in 2024, music publishers sued Anthropic over lyrics reproduction. The Anthropic case, Concord Music Group v. Anthropic, involved similar claims about training data and output reproduction. The authors' cases also intersect with the New York Times lawsuit against OpenAI and Microsoft, filed in December 2023, which alleged that the models reproduced news articles verbatim.
In the authors' litigation, the court has coordinated discovery with the New York Times case, allowing for shared document production and expert testimony. This coordination has slowed the pace of the litigation, with the court setting a trial date for late 2025 or early 2026. As of early 2025, the parties were engaged in fact discovery, including depositions of OpenAI executives and technical staff, and the exchange of training data metadata.
Implications for AI Development
The outcome of the authors' class actions could significantly affect the artificial intelligence industry. A ruling that training on copyrighted works without a license is not fair use would force AI developers to negotiate licenses with publishers and authors, potentially increasing costs and limiting the diversity of training data. Conversely, a ruling in favor of OpenAI would affirm the legality of current training practices, though it might not resolve questions about output reproduction.
The case has also prompted legislative attention. In 2024, several U.S. senators introduced bills proposing transparency requirements for training data and a statutory licensing scheme for copyrighted works. The authors' litigation has been cited in these debates as evidence of the need for clearer rules. Internationally, the European Union's AI Act, adopted in 2024, includes provisions requiring disclosure of training data summaries, though it does not directly address copyright liability.
Current Status and Outlook
As of early 2025, the consolidated action is in the discovery phase. The court has denied class certification motions, but the plaintiffs have indicated they will seek certification again after fact discovery. Judge Rakoff has scheduled a pretrial conference for mid-2025, with a potential trial in 2026. The case is widely watched by legal scholars and industry observers, as it may set precedent for how U.S. courts apply fair use to large-scale machine learning.
OpenAI has also faced parallel litigation in other jurisdictions, including a class action in Canada and a complaint before the U.S. Copyright Office. The company has begun entering into licensing agreements with some publishers, such as Axel Springer and the Associated Press, but has not reached agreements with individual fiction authors. The authors' class actions remain the most comprehensive challenge to OpenAI's training practices, and their resolution will likely shape the future of generative AI and copyright law.