# EleutherAI

EleutherAI is a non-profit artificial intelligence research group founded in 2020 to create open-source versions of GPT-3, known for releasing large language models and datasets like The Pile and GPT-NeoX-20B.

EleutherAI is a non-profit artificial intelligence research group that develops open-source AI models and datasets. Founded in 2020, the collective is often described as an open-source counterpart to [OpenAI](https://www.wikiprompt.org/wiki/openai), aiming to democratize access to [large language models](https://www.wikiprompt.org/wiki/large-language-model) and related research. In early 2023, it formally incorporated as the EleutherAI Institute, a non-profit research institute. As of 2025, the organization maintains widely-used training datasets, conducts research in areas such as interpretability and alignment, and engages in public policy discussions.

The group originated in a Discord server and has grown into a collaborative community of hundreds of volunteer researchers. Its work has influenced the broader field of [generative AI](https://www.wikiprompt.org/wiki/generative-ai) by providing freely available models and datasets that have fueled new startups and academic research.

## History

EleutherAI began as a Discord server on July 7, 2020, under the tentative name "LibreAI" before rebranding to "EleutherAI" later that month, referencing the Greek word for liberty, *eleutheria*. The founding members were Connor Leahy, Leo Gao, and Sid Black, who co-wrote the initial code to create a [machine learning](https://www.wikiprompt.org/wiki/machine-learning) model similar to GPT-3.

On December 31, 2020, the group released The Pile, a curated dataset of diverse text for training large language models. The first models, GPT-Neo, were released on March 21, 2021, followed by GPT-J-6B on June 9, 2021, which was then the largest open-source GPT-3-like model. These models were released under the Apache 2.0 license and are considered to have fueled an entirely new wave of startups.

Initially, EleutherAI turned down funding offers, relying on Google's TPU Research Cloud Program for compute. By early 2021, they accepted funding from CoreWeave and SpellML in the form of GPU cluster access. On February 10, 2022, they released GPT-NeoX-20B, a larger model enabled by these resources.

In early 2023, EleutherAI incorporated as a non-profit research institute led by Stella Biderman, Curtis Huebner, and Shivanshu Purohit. The organization shifted focus toward interpretability, alignment, and scientific research, as training and releasing LLMs had become more common elsewhere.

In July 2024, an investigation by Proof News and Wired found that The Pile dataset included subtitles from over 170,000 YouTube videos across more than 48,000 channels, drawing criticism and accusations of theft. In response, EleutherAI released "Common Pile" in 2025, a dataset without the controversial copyrighted material, and trained two models from it. Collaborating with the UK's AI Security Institute, they also found that filtering training data to remove key concepts can maintain performance while reducing harmful information.

## Research and Models

EleutherAI's research spans multiple areas, with hundreds of volunteer contributors. The group has released several notable models and datasets that are widely used in the AI community.

### The Pile

The Pile is an 886 GB dataset designed for training large language models. Originally developed for GPT-Neo, it has been used to train other models, including Microsoft's Megatron-Turing Natural Language Generation. Its distinguishing features are its curated selection of data chosen by researchers and its thorough documentation. The initial version faced scrutiny for containing copyrighted material, including books and subtitles from various media.

### Common Pile

Common Pile v0.1, released in June 2025 in partnership with many collaborators, contains only works whose licenses permit use for training AI models. This addresses the copyright issues of The Pile.

### GPT Models

EleutherAI trained open-source large language models inspired by OpenAI's GPT-3, with parameter counts of 125 million, 1.3 billion, 2.7 billion, 6 billion, and 20 billion. GPT-Neo (125M, 1.3B, 2.7B) was released in March 2021, followed by GPT-J (6B) in June 2021, both being the largest open-source GPT-3-style models at their release. GPT-NeoX-20B followed on February 10, 2022.

### VQGAN-CLIP

After OpenAI released DALL-E in January 2021 without public access, EleutherAI's Katherine Crowson and digital artist Ryan Murdock developed a technique using Contrastive Language-Image Pre-training (CLIP) to convert regular image generation models into text-to-image synthesis ones. Building on ideas from Google's DeepDream, they combined CLIP with VQGAN to create VQGAN-CLIP. Crowson shared the technology via notebooks that people could run for free.

## Impact and Policy

EleutherAI's open-source models and datasets have significantly influenced the AI ecosystem, enabling startups and researchers to work with large language models without proprietary restrictions. The organization also participates in public policy discussions, advocating for open research and addressing ethical concerns around data usage and model safety.

---
Source: https://www.wikiprompt.org/wiki/eleuther-ai
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-12T16:19:55.378385+00:00
