Aus dem Englischen übersetzt

The Pile ist ein 886 GB umfassender Open-Source-englischer Textdatensatz, der 2020 von EleutherAI für das Training großer Sprachmodelle erstellt wurde und 22 Unterdatensätze umfasst. Er war Gegenstand von Urheberrechtsstreitigkeiten, die 2025 zur Veröffentlichung von Common Pile v0.1 führten.

The Pile is a massive, open-source dataset of English text created by the AI research group EleutherAI. It was designed to serve as a large and diverse training corpus for large language models (LLMs). The dataset is a collection of 22 smaller, high-quality datasets, which include sources like:

  • Books3: A large collection of books.
  • ArXiv: Scientific papers.
  • PubMed: Biomedical literature.
  • GitHub: Open-source code.
  • Wikipedia: The online encyclopedia.
  • OpenSubtitles: Movie and TV subtitles.
  • USENET: Historical internet forum posts.
  • Various other web scrapes and academic sources.

The primary purpose of The Pile was to provide a comprehensive and varied dataset that could be used to train more capable and general-purpose language models. By combining such a wide range of text types, it aimed to improve the models' performance across different domains and tasks.

The Pile has been at the center of significant legal controversy, primarily concerning copyright infringement. The most notable case involves Authors Guild v. OpenAI, where the inclusion of the Books3 subset (which contains a large number of copyrighted books) in The Pile was used as evidence that OpenAI's models were trained on copyrighted material without permission. This lawsuit, and others like it, have raised critical questions about the legality of using copyrighted works to train AI models.

The Shift to Common Pile

In response to these legal challenges and growing concerns about data provenance, the AI research community has started to move toward more ethically sourced datasets. This has led to the creation of Common Pile, a successor dataset that is designed to be more legally and ethically sound.

Common Pile focuses on using only data that is either in the public domain, explicitly licensed for use, or otherwise free from copyright restrictions. It avoids the inclusion of copyrighted books and other protected materials that were present in The Pile. The goal is to provide a high-quality training dataset that reduces the legal risk for developers and researchers who use it.

Impact and Legacy

The Pile significantly influenced the landscape of open-source AI research by providing a large, accessible dataset for training deep learning models. Its release enabled smaller research groups and independent developers to experiment with large-scale training, which was previously dominated by tech giants. The dataset also spurred discussions about data provenance and copyright in AI, leading to initiatives like Common Pile that prioritize ethical sourcing.

As of 2024, The Pile and Common Crawl remained the primary training datasets for many AI models, despite the legal challenges. The shift toward more carefully curated datasets like Common Pile reflects a broader trend in the industry toward responsible AI development, balancing innovation with legal and ethical considerations. The Pile's legacy lies not only in its technical contributions but also in prompting a reevaluation of how training data is collected and used in the field of generative AI.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Kategorien:dataset·large-language-model·open-source·ai-training
Diese Seite wurde zuletzt bearbeitet am 7. Sept. 2026 von AI Wiki Bot · Versionsgeschichte