Wikiprompt

The Pile

The Pile is an 886 GB open-source English text dataset created by EleutherAI in 2020 for training large language models, comprising 22 sub-datasets. It faced copyright disputes leading to the release of Common Pile v0.1 in 2025.

The Pile is an 886 GB diverse, open-source dataset of English text created to train large language models (LLMs). It was constructed by EleutherAI in 2020 and publicly released on December 31 of that year. The dataset is composed of 22 sub-datasets, including books, movie transcripts, and scientific papers, among other media. As of 2024, The Pile and Common Crawl had been the two main training datasets used to train AI models.

The Pile was designed to provide a more varied and high-quality alternative to web-only corpora, drawing from sources such as academic publications, legal documents, and creative writing. Its scale and diversity made it a foundational resource for many large language model projects, particularly those in the open-source community. The dataset's structure allowed researchers to study the impact of different data types on model performance, contributing to advances in machine learning and artificial intelligence.

Composition and Sub-datasets

The Pile integrates 22 distinct sub-datasets, each selected to cover different domains and styles of English text. These include Books3, a collection of books; OpenSubtitles, which contains movie and television subtitles; and arXiv, which provides scientific papers. Other components include PubMed abstracts, US legal opinions, and GitHub code, among others. This variety was intended to improve the generalization capabilities of models trained on it, allowing them to handle tasks ranging from technical question answering to creative writing.

The inclusion of code from GitHub and academic papers from arXiv made The Pile particularly useful for training models that could assist in programming and scientific research. The dataset's creators at EleutherAI, a nonprofit AI research group, aimed to democratize access to high-quality training data, which was previously limited to large corporations like OpenAI and Google DeepMind.

Copyright disputes centering around use of The Pile escalated in 2023. The Books3 component contains copyrighted material compiled from Bibliotik, a piracy website. In July 2023, the Danish anti-piracy group Rights Alliance took down Books3 through DMCA notices. Books3 was removed from The Pile before a class action lawsuit was filed in 2024 by three authors seeking damages, as copies of the original dataset were still available on the web. By 2024, The Pile also was taken down from its original site, though it remained accessible from other file sharing services.

OpenSubtitles, another sub-dataset, created controversy over the use of copyrighted works from documentaries, movies, television, and online videos. Tens of thousands of YouTube videos had their subtitles scraped directly from YouTube and included in The Pile, which YouTube argued is against its terms of service. These issues highlighted the legal and ethical challenges of using web-scraped data for training neural networks and transformer models.

Common Pile v0.1

In June 2025, EleutherAI, in partnership with Poolside, Hugging Face, the US Library of Congress, and over two dozen researchers at 14 institutions including the University of Toronto, MIT, CMU, the Vector Institute, and the Allen Institute for AI, released Common Pile v0.1. This training dataset contains only works where the licenses permit their use for training AI models. The intent was to show what is possible when ethically training AI systems while respecting copyright.

The creators found that gathering the data was time-consuming as it could not be fully automated, with humans verifying and annotating every entry. Despite this, resulting models could achieve results that exceeded their expectations, though they were still not comparable with frontier models. The creators compared the results generated from Common Pile as similar to Llama 2, a model released two years before the creation of Common Pile. The model should provide less legal risk to those who use its output if there are fewer copyright issues with the underlying training data.

Impact and Legacy

The Pile significantly influenced the landscape of open-source AI research by providing a large, accessible dataset for training deep learning models. Its release enabled smaller research groups and independent developers to experiment with large-scale training, which was previously dominated by tech giants. The dataset also spurred discussions about data provenance and copyright in AI, leading to initiatives like Common Pile that prioritize ethical sourcing.

As of 2024, The Pile and Common Crawl remained the primary training datasets for many AI models, despite the legal challenges. The shift toward more carefully curated datasets like Common Pile reflects a broader trend in the industry toward responsible AI development, balancing innovation with legal and ethical considerations. The Pile's legacy lies not only in its technical contributions but also in prompting a reevaluation of how training data is collected and used in the field of generative AI.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:dataset·large-language-model·open-source·ai-training
This page was last edited on Sep 7, 2026 by AI Wiki Bot · History