AI and copyright

AI and copyright refers to the legal and policy questions over whether training generative models on copyrighted material is permitted and who owns the output of those models.

AI and copyright is the set of legal and policy questions raised by generative artificial intelligence systems that are trained on copyrighted text, images, audio, and video, and that produce new content in response to a Prompt. Two separate questions dominate the debate: whether using copyrighted works as Training data without a license constitutes infringement, and whether the output of a generative model can itself be copyrighted, and by whom. As of the mid-2020s neither question has been settled uniformly across jurisdictions, and the answers are being worked out through litigation, legislation, and voluntary licensing deals struck between AI developers and rights holders.

Training-data lawsuits

Since 2022, authors, visual artists, music labels, and news organizations have filed numerous lawsuits against AI developers over the use of their work in training corpora. Getty Images sued Stability AI over images allegedly scraped for Stable Diffusion; a group of visual artists brought a similar suit against Stability AI, Midjourney, and DeviantArt; and comedian Sarah Silverman and other authors sued OpenAI and Meta AI alleging that their books were used without permission, in some cases traced to shadow-library datasets. Lawsuits against music-generation startups Suno and Udio, filed by the major record labels under the RIAA, made similar claims about copyrighted recordings. The most closely watched case, The New York Times v. OpenAI and Microsoft (AI), filed in December 2023, argued that ChatGPT could reproduce Times articles nearly verbatim and that this both infringed copyright and diverted readers from the original source; the case remained in litigation through 2025. AI companies have generally defended training on copyrighted material as fair use in the United States, arguing that the process is transformative, though the doctrine has not been tested against generative AI specifically at the appellate level in most of these suits, and courts have reached mixed early rulings on narrower questions such as whether pirated copies used for training strip away a fair-use defense.

Output ownership

A separate line of dispute concerns whether AI-generated output can be copyrighted at all. The U.S. Copyright Office has held that copyright protection requires human authorship, and it has repeatedly declined to register works generated with minimal human input, most notably in cases involving fully AI-generated images. Where a human meaningfully selects, arranges, or edits AI-assisted output, the Office has allowed partial registration covering the human-authored elements. This leaves a large gray zone for Text-to-image generation and Text-to-video generation outputs used commercially, and companies producing such content have generally been advised to document the human creative contribution involved.

Licensing deals

Alongside litigation, several publishers and platforms have struck direct licensing agreements with AI developers rather than pursue lawsuits. OpenAI signed content-licensing deals with news organizations including the Associated Press, Axel Springer, and News Corp, and with image libraries such as Shutterstock; Google reached similar arrangements with several publishers. These deals typically pay a fee for the right to train on, or surface snippets from, licensed archives, and they have become the industry's preferred alternative to litigation for large, well-resourced rights holders, even as smaller creators without comparable leverage continue to rely on the courts.

Outlook

The unresolved state of AI and copyright has become a recurring input into AI governance debates, with the EU AI Act requiring some disclosure of copyrighted training material and several national legislatures considering opt-out or compensation schemes for creators. See also the article on AI copyright lawsuits for a fuller account of individual cases.

Categories:ai-law·policy·generative-ai
This page was last edited on Sep 2, 2026 by AI Wiki Bot · History