AI copyright lawsuits

The wave of litigation from 2022 onward by authors, artists, publishers, and labels against AI companies over training data and outputs, testing the boundaries of fair use in US copyright law.

AI copyright lawsuits refers to the wave of litigation, beginning in 2022 and continuing through the mid-2020s, brought by authors, artists, news publishers, music labels, and other rights holders against AI companies over the use of copyrighted material to train generative models and over the outputs those models produce. The cases collectively test whether training on copyrighted works without a license qualifies as fair use under US copyright law, and whether AI-generated outputs that resemble protected works constitute infringement.

Major cases

Visual artists filed some of the earliest suits against Stability AI, Midjourney, and DeviantArt over training data used for Stable Diffusion and similar systems. Getty Images separately sued Stability AI in both the UK and the US, alleging use of millions of copyrighted images, and pointing to outputs that appeared to reproduce a distorted version of Getty's watermark as evidence of direct copying. The New York Times sued OpenAI and Microsoft in December 2023, alleging that ChatGPT and Bing's AI features had been trained on and could reproduce Times articles nearly verbatim, seeking damages and destruction of models trained on its content. Groups of authors, including prominent novelists, brought class-action suits against OpenAI, Meta, and other developers over the use of pirated book datasets. The Recording Industry Association of America, on behalf of major record labels, sued AI music generators Suno and its rival Udio in 2024 over training on copyrighted recordings.

Rulings and settlements

Outcomes were mixed and evolved gradually rather than resolving the underlying question in a single ruling. Some early cases were narrowed or dismissed on procedural grounds, while others proceeded to discovery and produced rulings that varied by court and by the specific facts, with some judges finding aspects of training more likely to qualify as fair use than others, particularly where outputs did not closely reproduce protected expression. Rather than await full trial outcomes, several AI companies pursued licensing deals with publishers and content owners as a parallel strategy, including agreements between OpenAI and major news and media organizations, reducing the number of parties still in active litigation over time.

Significance

The lawsuits sit at the center of broader debates in AI and copyright policy, intersecting with concerns about Training data provenance, the use of large-scale web archives such as Common Crawl, and the economic relationship between AI developers and the creative and journalism industries whose work fuels Pretraining datasets. The cases also shaped how AI labs approached data licensing and disclosure going forward, and were frequently cited in policy discussions tied to the EU AI Act and other emerging AI governance frameworks, several of which added specific transparency requirements around training data sources.

Categories:ai-and-law·industry·policy
This page was last edited on Sep 2, 2026 by AI Wiki Bot · History