The United States Copyright Office (USCO), a unit of the Library of Congress, has examined the use of copyrighted works in training Artificial intelligence systems. Its findings, released in a series of reports beginning in 2024, address whether and how existing copyright law applies to the ingestion of protected materials by Machine learning models, including Large language models and other Generative AI tools. The Office's analysis has become a reference point for legislative and judicial discussions on AI and intellectual property.
The Copyright Office's role in this area stems from its broader mandate to administer U.S. copyright law and advise Congress. As part of the legislative branch, it does not enforce copyright but provides policy recommendations and registers claims. Its reports on AI training data have drawn on public comments, expert testimony, and international comparisons, aiming to clarify the legal landscape without prescribing new legislation unless Congress requests it.
Background: Copyright and AI Training
Training an AI model typically involves feeding large datasets of text, images, or other works into a Neural network or Transformer (architecture) architecture. This process, known as machine learning, allows the model to recognize patterns and generate new outputs. The datasets often include copyrighted books, articles, photographs, and artworks, raising questions about whether such use constitutes infringement or falls under exceptions like fair use.
Before the USCO's involvement, courts had begun addressing these issues in cases involving companies like OpenAI and Anthropic, but no uniform standard had emerged. The Copyright Office's reports aim to provide a comprehensive overview of the legal principles and practical implications, helping to guide both policymakers and stakeholders.
The 2023 Notice of Inquiry
In August 2023, the Copyright Office issued a notice of inquiry in the Federal Register, soliciting public comments on copyright and AI, specifically focusing on training data. The inquiry asked about the nature of AI training, the extent to which copyrighted works are used, and the legal arguments for and against liability. Over 10,000 comments were submitted from individuals, companies, academic institutions, and advocacy groups, reflecting the high stakes of the issue.
The Office also held listening sessions with stakeholders, including creators, technology firms, and legal experts. These sessions highlighted divergent views: some argued that training on copyrighted works without permission is a transformative fair use, while others contended that it requires licensing or compensation. The comments and sessions informed the subsequent report.
The 2024 Report: Copyright and Artificial Intelligence
In March 2024, the Copyright Office released the first part of its report, titled "Copyright and Artificial Intelligence," focusing on digital replicas and the use of copyrighted works in training. The report concluded that existing copyright law, particularly the fair use doctrine, can apply to AI training, but it declined to endorse a blanket exemption or a new compulsory license. Instead, it emphasized a case-by-case analysis, considering factors such as the purpose of the use, the nature of the copyrighted work, the amount used, and the effect on the market.
The report noted that the Deep learning process often involves copying entire works into training datasets, which may weigh against fair use, but it also acknowledged that the transformative nature of AI outputs could support a defense. The Office recommended that Congress monitor developments and consider targeted legislation if needed, but it stopped short of proposing specific statutory changes.
Key Findings on Training Data
The report's key findings on training data include:
- Ingestion as reproduction: Training an AI model typically involves making copies of works, which implicates the reproduction right under copyright law.
- Fair use is not automatic: The Office rejected the argument that AI training is inherently fair use, emphasizing that each case must be evaluated on its facts.
- Market impact: The report considered whether AI training harms or benefits the market for original works, noting that this factor is often contested.
- Transparency and provenance: The Office encouraged voluntary disclosure of training data sources, but it did not mandate such disclosure, citing practical challenges.
These findings have been cited in ongoing litigation, such as the lawsuits filed by authors against AI companies, and in congressional hearings on AI policy.
The 2025 Update and Ongoing Work
In early 2025, the Copyright Office issued a follow-up report addressing digital replicas and the intersection with state publicity rights. It also announced plans to examine other aspects of AI, including the copyrightability of AI-generated outputs. The Office's work has been complicated by a leadership dispute: in May 2025, President Donald Trump claimed to have fired Register Shira Perlmutter, but the position is an employee of Congress, and Perlmutter's authority has been upheld in court. This uncertainty has raised questions about the continuity of the Office's AI initiatives.
Despite the dispute, the Office has continued to publish guidance and participate in international discussions, such as those at the World Intellectual Property Organization. Its reports remain a key resource for understanding how U.S. copyright law applies to AI training.
Implications for Industry and Policy
The USCO's findings have significant implications for AI developers, content creators, and policymakers. For companies like OpenAI, Anthropic, and Google DeepMind, the reports clarify that they cannot assume fair use protection, potentially increasing the need for licensing agreements or changes to training practices. For creators, the reports affirm that their works are protected, but they also highlight the difficulty of enforcing rights against opaque training processes.
Legislatively, the reports have informed proposed bills, such as the Generative AI Copyright Disclosure Act, which would require AI companies to disclose their training data. The Copyright Office has not endorsed this specific bill, but it has supported the idea of transparency as a policy goal.
Criticisms and Debates
The Copyright Office's approach has faced criticism from both sides. Some technology advocates argue that the reports are too cautious and could stifle innovation, while some creator groups contend that they do not go far enough in protecting rights. Legal scholars have also debated the Office's interpretation of fair use, with some arguing that the case-by-case approach is impractical for the scale of AI training.
Another point of contention is the role of the Copyright Office itself. As a part of the Library of Congress, it operates within the legislative branch, but its advisory role is not binding on courts. The reports are influential, but they do not have the force of law, and courts may reach different conclusions.
Future Directions
Looking ahead, the Copyright Office is expected to continue its work on AI, with additional reports on copyrightability and potential legislative recommendations. The outcome of the leadership dispute may affect the Office's priorities, but its mandate to advise Congress remains unchanged. As AI technology evolves, the Office will likely revisit its findings to address new challenges, such as the use of synthetic data and the impact of AI on creative industries.
For now, the USCO's reports on training data provide a foundational framework for navigating the complex intersection of copyright and artificial intelligence. They underscore the need for a balanced approach that respects both innovation and the rights of creators, a goal that will require ongoing dialogue among all stakeholders.