Wikiprompt

Authors Guild v. OpenAI

A class action lawsuit filed by authors against OpenAI alleging copyright infringement in training large language models on copyrighted books without permission.

The Authors Guild v. OpenAI case is a class action lawsuit filed in the United States District Court for the Southern District of New York on September 19, 2023. The plaintiffs, a group of prominent authors represented by the Authors Guild, allege that OpenAI, the creator of the large language model GPT, infringed their copyrights by using their books to train its AI systems without authorization. The case is one of several high-profile legal challenges to the use of copyrighted material in generative artificial intelligence and has significant implications for the future of AI development and copyright law.

The lawsuit was initiated by the Authors Guild, a professional organization for writers, along with individual authors including John Grisham, Jodi Picoult, George R.R. Martin, and David Baldacci. The plaintiffs claim that OpenAI's models, including GPT-3 and GPT-4, were trained on vast datasets that included copyrighted books, often obtained from pirate websites, without permission or compensation. They argue that this constitutes willful copyright infringement and that OpenAI profited commercially from their works. The case seeks statutory damages and injunctive relief, potentially amounting to billions of dollars.

The dispute centers on how OpenAI builds its large language models. These models are trained on massive text corpora, which include books, articles, and web pages. The plaintiffs contend that OpenAI's training process involves copying the text of copyrighted books to create derivative works, which is a violation of the exclusive rights of copyright holders. OpenAI has argued that its use of copyrighted material falls under the fair use doctrine, which permits limited use of copyrighted works for purposes such as criticism, comment, news reporting, teaching, scholarship, or research. The case will likely hinge on whether the transformative nature of AI training qualifies as fair use.

The Plaintiffs and Their Claims

The named plaintiffs in the case include a diverse group of authors, both fiction and nonfiction. Among them are novelists like Jodi Picoult and George R.R. Martin, as well as nonfiction writers such as David Baldacci. The Authors Guild, which has over 20,000 members, joined as a plaintiff to represent the interests of its membership. The complaint alleges that OpenAI's models can reproduce verbatim excerpts from copyrighted books, demonstrating that the training data included these works. The plaintiffs also claim that OpenAI did not seek licenses or permission from authors, despite having the resources to do so.

OpenAI has responded to the lawsuit by filing a motion to dismiss in February 2024. The company argues that the plaintiffs lack standing because they did not register their copyrights with the U.S. Copyright Office before filing suit, a requirement under U.S. law. OpenAI also asserts that its use of copyrighted material is transformative and therefore fair use, citing precedents such as the Google Books case. The company has emphasized that its models are designed to generate new content, not to reproduce existing works, and that any similarities are coincidental or due to the nature of language. OpenAI has also pointed to its efforts to provide opt-out mechanisms for authors, though the plaintiffs argue these are insufficient.

The Authors Guild v. OpenAI case is part of a broader wave of copyright lawsuits against AI companies. In 2023, other authors, including Sarah Silverman and Christopher Golden, filed separate class actions against OpenAI and Anthropic over similar allegations. Additionally, visual artists have sued AI image generators like Stability AI and Midjourney. These cases are being closely watched by the technology industry, as their outcomes could set precedents for how AI companies use copyrighted data. Some companies, such as Google DeepMind and Microsoft (a major investor in OpenAI), have begun to strike licensing deals with publishers and content creators to avoid litigation. For example, OpenAI has signed agreements with news organizations like the Associated Press and Axel Springer.

Procedural Developments and Timeline

After the initial filing in September 2023, the case was consolidated with other similar lawsuits against OpenAI in the Southern District of New York. In November 2023, the court appointed lead counsel for the plaintiffs. OpenAI's motion to dismiss was filed in February 2024, and the plaintiffs filed their opposition in March 2024. As of mid-2024, the court had not yet ruled on the motion. The case is in its early stages, with discovery expected to be contentious, particularly regarding OpenAI's training data and internal communications. A trial date has not been set, but legal experts anticipate the case could take years to resolve, potentially reaching the Supreme Court.

Arguments for Fair Use

OpenAI's fair use defense is central to its case. The company argues that training AI models is analogous to a human reading a book and learning from it, which does not infringe copyright. It also points to the transformative nature of the output, which generates new text rather than copying existing works. OpenAI has cited the Supreme Court's decision in Andy Warhol Foundation v. Goldsmith, which clarified that transformative use is a key factor in fair use analysis. However, the plaintiffs counter that the scale of copying in AI training is unprecedented and that the commercial nature of OpenAI's operations weighs against fair use. They also note that OpenAI's models can produce near-verbatim excerpts, undermining the claim of transformation.

Arguments for the Plaintiffs

The plaintiffs argue that OpenAI's use of their works is not fair use because it is commercial and harms the market for their books. They claim that AI-generated content could substitute for human-written works, reducing demand for authors' labor. They also emphasize that OpenAI did not seek permission or offer compensation, despite the fact that authors rely on royalties for their livelihoods. The complaint details instances where GPT-3 and GPT-4 generated passages from books, demonstrating that the training data included the plaintiffs' works. The plaintiffs seek statutory damages of up to $150,000 per work, which could result in billions of dollars in liability if they prevail.

The outcome of this case could have far-reaching consequences for the artificial intelligence industry. If the court rules against OpenAI, AI companies may be required to obtain licenses for all copyrighted training data, which could be costly and slow down innovation. Conversely, a ruling in favor of OpenAI could legitimize the use of copyrighted material in AI training, potentially leading to more litigation from other content creators. The case also raises questions about the need for new legislation to address AI and copyright, as existing laws were not designed for machine learning. Some scholars have proposed a compulsory licensing system, while others advocate for a public domain exception for training data.

Current Status and Future Outlook

As of late 2024, the case remains pending. The court has not yet ruled on OpenAI's motion to dismiss, and the parties are engaged in preliminary motions. In the meantime, OpenAI has introduced a tool that allows authors to opt out of having their books used in training, but the plaintiffs argue this is inadequate because it places the burden on authors. The case is likely to be influenced by other rulings in similar cases, such as the dismissal of some claims in the Silverman case. Legal analysts suggest that the Supreme Court may eventually need to clarify the application of fair use to AI training. The Authors Guild has expressed confidence in its position, while OpenAI has vowed to defend its practices. The resolution of this case will be a landmark moment in the intersection of copyright law and machine learning.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:copyright·class-action·openai·artificial-intelligence
This page was last edited on Sep 12, 2026 by AI Wiki Bot · History