Kadrey v. Meta Platforms is a civil lawsuit filed in the United States District Court for the Northern District of California in 2023. The plaintiffs, a group of authors including Richard Kadrey, Sarah Silverman, and Christopher Golden, allege that Meta Platforms, Inc. infringed their copyrights by using pirated copies of their books to train its LLaMA (Large Language Model Meta AI) family of large language models. The case is part of a broader wave of litigation against developers of generative artificial intelligence systems, raising questions about the legality of training models on copyrighted text without explicit permission.
The lawsuit centers on Meta's use of a dataset known as "Books3," which contains over 195,000 books, many of which were sourced from the shadow library Bibliotik. The plaintiffs claim that Meta downloaded and used this dataset to train LLaMA models, including LLaMA 1 and LLaMA 2, without obtaining licenses or compensating the authors. The case has become a key test for the application of copyright law to machine learning, particularly the doctrine of fair use.
Background and Parties
The lead plaintiff, Richard Kadrey, is an American author known for the Sandman Slim urban fantasy series. Sarah Silverman, a comedian and author, and Christopher Golden, a horror and fantasy writer, joined the suit. The complaint was filed on July 7, 2023, and later consolidated with similar actions, including a case brought by author Paul Tremblay. The defendants are Meta Platforms, Inc., and its subsidiary Meta AI, which developed the LLaMA models.
The plaintiffs are represented by the law firm Joseph Saveri Law Firm, which has filed multiple suits against AI companies. Meta is represented by attorneys from Gibson, Dunn & Crutcher. The case is assigned to Judge Jacqueline Scott Corley, who also oversees other AI copyright disputes.
Allegations of Copyright Infringement
The core allegation is that Meta reproduced and distributed copyrighted books without authorization during the training process. The plaintiffs argue that the LLaMA models, when prompted, can generate text that closely resembles passages from their books, demonstrating that the models have memorized and can reproduce the copyrighted material. They also claim that Meta's use of the Books3 dataset, which was created by a user on the AI-focused platform The Eye, was knowingly infringing because the dataset was widely known to contain pirated books.
The complaint cites specific examples, such as the model generating text that matches excerpts from Kadrey's novel "Sandman Slim" and Silverman's memoir "The Bedwetter." The plaintiffs assert that this reproduction constitutes direct and vicarious infringement, as well as contributory infringement, because Meta facilitated the creation of the infringing dataset.
Meta's Defense and Fair Use Argument
Meta's primary defense is that training AI models on copyrighted text constitutes fair use under Section 107 of the U.S. Copyright Act. The company argues that the use is transformative because the models do not reproduce the books in their entirety but instead learn statistical patterns and language structures. Meta also contends that the training process is analogous to a human reading a book and learning from it, and that the models do not compete with the original works in the marketplace.
Meta has filed motions to dismiss the case, arguing that the plaintiffs failed to state a claim because they did not demonstrate that the models output substantial portions of the copyrighted works. The company also points to the fact that LLaMA models are open-source and that the training data is not publicly distributed, which it argues limits any market harm.
Procedural History
After the initial filing in July 2023, the court consolidated several related cases under the lead case number 3:23-cv-03417. In November 2023, Judge Corley denied Meta's motion to dismiss the direct infringement claims but allowed the plaintiffs to amend their complaint regarding vicarious and contributory infringement. The plaintiffs filed an amended complaint in December 2023, adding more specific allegations.
In February 2024, Meta filed a new motion to dismiss the amended complaint, which was partially granted in May 2024. The court dismissed the contributory infringement claim but allowed the direct infringement and vicarious liability claims to proceed. Discovery is ongoing, with both sides seeking internal communications and training data documentation from Meta.
Broader Context of AI Copyright Litigation
The case is one of several high-profile lawsuits against AI developers. Similar actions have been filed against OpenAI and Anthropic by authors, including a class action led by comedian Sarah Silverman (which was later dismissed in part), and against Google by authors over its Gemini model. The outcomes of these cases are expected to shape the legal landscape for generative AI, potentially influencing how companies collect training data and whether they need to license copyrighted works.
The U.S. Copyright Office has also been studying the issue, and in 2023 it issued a notice of inquiry seeking public comments on the intersection of AI and copyright. The office has not yet issued final guidance, but its reports have emphasized the need for a balanced approach that protects authors while allowing technological innovation.
Key Legal Questions
The case raises several unresolved legal questions. First, whether the reproduction of copyrighted works during training constitutes "copying" in the legal sense, even if the copies are transient and not directly distributed. Second, whether the transformative nature of AI models weighs in favor of fair use, given that the models can generate new text but also memorize and reproduce passages. Third, whether the market harm to authors is significant, especially if AI models can produce text that substitutes for the original books.
Courts have historically applied a four-factor test for fair use, considering the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the market. In the related case of Authors Guild v. Google (2015), the Second Circuit found that Google's digitization of books for search was fair use, but that case involved limited display of snippets. The Kadrey case involves more extensive reproduction, which may lead to a different outcome.
Potential Outcomes and Implications
If the plaintiffs prevail, Meta could be required to pay statutory damages of up to $150,000 per work, which could amount to billions of dollars given the number of books in the dataset. The company might also be ordered to destroy the LLaMA models or retrain them on licensed data. Such a ruling could have a chilling effect on AI development, forcing companies to obtain licenses for all training data, which would be costly and time-consuming.
Conversely, if Meta wins on fair use, it would establish a broad precedent that AI training on copyrighted text is permissible, potentially accelerating the development of large language models. This could also affect other pending cases and encourage more companies to use scraped data without compensation.
Current Status and Future Proceedings
As of late 2024, the case is in the discovery phase. The parties are exchanging documents and taking depositions. A trial date has not been set, but legal analysts expect the case to proceed to summary judgment motions in 2025. The court may also decide to certify a class action, which would expand the plaintiff group to include all authors whose works are in the Books3 dataset.
The case has attracted amicus briefs from various organizations, including the Authors Guild, which supports the plaintiffs, and the Electronic Frontier Foundation, which has filed briefs in similar cases arguing for fair use. The outcome is likely to be appealed regardless of the district court's decision, potentially reaching the Supreme Court.
Significance for the AI Industry
The Kadrey case is a bellwether for the AI industry. It tests the limits of what constitutes acceptable training data and whether the open-source nature of models like LLaMA changes the legal analysis. The case also highlights the tension between innovation and intellectual property rights, a debate that is central to the future of Artificial intelligence and Generative AI.
Companies like OpenAI and Anthropic are closely watching the proceedings, as they face similar lawsuits. The outcome could influence how they negotiate licensing deals with publishers and authors, and whether they invest in alternative training methods that avoid copyrighted material. The case also has implications for Machine learning research, as many academic datasets contain copyrighted text.
In the meantime, Meta has continued to release new versions of LLaMA, including LLaMA 3, which was trained on a larger dataset that may include additional copyrighted works. The company has stated that it is committed to respecting intellectual property rights, but it has not disclosed the full contents of its training data, citing trade secrets.
Conclusion
Kadrey v. Meta Platforms is a landmark case that will help define the legal boundaries of AI training. The decision will have far-reaching consequences for authors, technology companies, and the public. While the case is still ongoing, its outcome is likely to set a precedent that influences how copyright law is applied to Large language models and other Neural network systems. The court's ruling will be closely scrutinized by legal scholars, policymakers, and industry stakeholders, as it addresses fundamental questions about the balance between protecting creative works and fostering technological progress.