Getty v. Stability (2023) is a copyright infringement lawsuit filed by Getty Images against Stability AI, the developer of the Stable Diffusion image-generation model. The case, initiated in January 2023 in the United States District Court for the District of Delaware, alleges that Stability AI copied and processed millions of images from Getty's database without authorization to train its Generative AI model. The lawsuit seeks statutory damages, injunctive relief, and the destruction of infringing training data.
The dispute centers on the use of web-scraped images in training Deep learning systems. Getty Images, a major stock photo agency, claims that Stability AI's actions violate its copyrights and those of its contributors. Stability AI has argued that its use of publicly available images falls under fair use, a defense that has been central to several similar cases in the Artificial intelligence field.
Background of the Case
In 2022, Stability AI released Stable Diffusion, an open-source Neural network model capable of generating detailed images from text prompts. The model was trained on a dataset called LAION-5B, which contained billions of image-text pairs scraped from the internet. Getty Images, which licenses millions of photographs, alleged that a substantial portion of its proprietary images were included in this dataset without permission.
Getty Images filed its complaint on January 13, 2023, in the U.S. District Court for the District of Delaware. The lawsuit names Stability AI, as well as its CEO, Emad Mostaque, and the company's U.S. subsidiary, Stability AI Ltd. The complaint asserts direct, contributory, and vicarious copyright infringement, as well as violations of the Digital Millennium Copyright Act (DMCA) for removing or altering copyright management information.
Legal Claims and Arguments
Getty Images' primary claim is that Stability AI reproduced and stored millions of copyrighted images during the training process. The company argues that this reproduction is not transformative and that the commercial nature of Stability AI's use weighs against fair use. Getty also contends that Stable Diffusion can generate images that closely resemble the original photographs, including watermarks and other identifying marks, which further undermines the fair use defense.
Stability AI has responded by asserting that training on publicly available images is a lawful fair use, similar to a human artist learning from existing works. The company has also argued that the images generated by Stable Diffusion are new, transformative works that do not infringe on any specific copyright. Stability AI has sought to dismiss the case, but the court has allowed the claims to proceed, indicating that the legal questions are substantial.
Related Lawsuits and Industry Impact
The Getty v. Stability case is part of a broader wave of litigation against Generative AI companies. In January 2023, three artists - Sarah Andersen, Kelly McKernan, and Karla Ortiz - filed a class-action lawsuit against Stability AI, Midjourney, and DeviantArt, alleging similar copyright violations. That case, Andersen v. Stability AI, was later dismissed in part but allowed to proceed on certain claims. Other lawsuits have been filed by authors against OpenAI and Anthropic over the use of copyrighted text in training Large language models.
These cases have significant implications for the Artificial intelligence industry. If courts rule against the AI companies, they may be required to license training data, which could increase costs and alter the development of future models. Conversely, a ruling in favor of fair use could solidify the legality of training on publicly available data, encouraging further innovation.
The Role of Training Data
The core of the dispute lies in how Machine learning models are trained. Stable Diffusion is a type of Generative AI model that uses a U-Net architecture and a process called diffusion to generate images. The training process involves feeding the model millions of images and their associated text descriptions, allowing it to learn the statistical relationships between visual features and language.
Getty Images argues that this process involves unauthorized reproduction of its copyrighted works. The company has pointed out that Stable Diffusion can sometimes reproduce near-exact copies of images from its database, including the Getty Images watermark. This, Getty claims, demonstrates that the model has memorized and can output copyrighted content, which is not a fair use.
Procedural History and Current Status
After Getty filed its complaint, Stability AI moved to dismiss the case, arguing that the claims were speculative and that the company had not directly infringed any copyright. In a ruling in late 2023, the court denied the motion to dismiss, allowing the case to proceed to discovery. The court found that Getty had plausibly alleged that Stability AI had copied its images and that the fair use defense was not sufficient to dismiss the claims at this stage.
As of early 2025, the case is still in the discovery phase. Both parties have been engaged in extensive document production and depositions. A trial date has not yet been set. The outcome of this case could set a precedent for how copyright law applies to Artificial intelligence training, affecting not only Stability AI but also other companies like OpenAI, Anthropic, and Google DeepMind.
Broader Context: AI and Copyright Law
The Getty v. Stability case is one of several high-profile disputes that have prompted legal and regulatory discussions about AI and copyright. In the United States, the Copyright Office has issued guidance stating that works generated entirely by AI are not eligible for copyright protection, but the question of whether training on copyrighted works is infringement remains unresolved. In the European Union, the AI Act includes provisions requiring transparency about training data, but it does not explicitly address copyright.
Some experts argue that current copyright law is ill-equipped to handle the complexities of AI training. They suggest that new legislation may be needed to balance the interests of copyright holders and AI developers. Others believe that fair use is sufficient and that courts will ultimately uphold the legality of training on publicly available data.
Potential Outcomes and Implications
If Getty Images prevails, Stability AI could be ordered to pay substantial damages and to remove or destroy its training datasets. This could have a chilling effect on the development of open-source AI models, as other developers may be reluctant to train on web-scraped data without licenses. It could also lead to a market for licensed training data, where companies like Getty Images sell access to their archives for AI training purposes.
If Stability AI wins, it would reinforce the notion that training on publicly available data is a fair use, which could accelerate AI development. However, it might also lead to more aggressive scraping practices, potentially harming photographers and other content creators who rely on licensing fees.
The case is closely watched by the Artificial intelligence community, legal scholars, and content creators. Its resolution will likely influence how AI companies approach data collection and whether they seek licenses from copyright holders. As of now, the case remains pending, and no final judgment has been issued.