GPT-2 Paper

GPT-2 is a large language model by OpenAI, released in 2019, known for its text generation capabilities and staged release due to safety concerns. It was a direct scale-up of GPT-1 and demonstrated emergent multitask learning.

Generative Pre-trained Transformer 2 (GPT-2) is a Large language model developed by OpenAI and the second in their foundational series of GPT models. It was pre-trained on a dataset of 8 million web pages and was partially released in February 2019, followed by the full release of the 1.5-billion-parameter model on November 5, 2019. GPT-2 was created as a direct scale-up of its predecessor, GPT-1, with a ten-fold increase in both parameter count and training dataset size. It is a general-purpose learner whose ability to perform various tasks was a consequence of its general ability to accurately predict the next item in a sequence, enabling it to translate texts, answer questions, summarize passages, and generate text output sometimes indistinguishable from human writing, though it could become repetitive or nonsensical in long passages.

GPT-2, like its predecessor and successors, uses a generative pre-trained Transformer (architecture) architecture, implementing a deep Neural network that relies on attention mechanisms instead of older recurrence- and convolution-based architectures. This allows the model to selectively focus on relevant input segments and greatly increases parallelization, outperforming previous benchmarks for RNN/CNN/LSTM-based models. It was superseded by GPT-3 and GPT-4, which are no longer open source.

Training

Since the Transformer (architecture) architecture enabled massive parallelization, GPT models could be trained on larger corpora than previous natural language processing models. While GPT-1 demonstrated the approach was viable, GPT-2 explored emergent properties of networks trained on extremely large corpora. CommonCrawl, a large web-crawled corpus, was considered but rejected due to large amounts of unintelligible content. Instead, OpenAI developed a new corpus called WebText, generated by scraping only pages linked to by Reddit posts that had received at least 3 karma prior to December 2017. The corpus was cleaned by parsing HTML into plain text, eliminating duplicate pages, and removing Wikipedia pages to avoid overfitting.

Documentation on training cost is limited. According to a statement by Andrej Karpathy, GPT-2 was trained on 32 TPU v3 chips for 168 hours (7 days) at approximately $8 per TPU v3 per hour, totaling about $43,000 in compute cost. Other sources cite a training cost of approximately $256 per hour. Comparable models like BERT and XLNet consumed $6,912 and $245,000 of resources, respectively.

Release

GPT-2 was first announced on 14 February 2019. A February 2019 article in The Verge by James Vincent noted that while the writing it produces is usually easily identifiable as non-human, it remained one of the most exciting examples of language generation programs. The Guardian described its output as plausible newspaper prose, and Kelsey Piper of Vox said it was one of the coolest AI systems she had ever seen. The Verge highlighted its flexibility in translating text, summarizing articles, and answering trivia questions. The GPT-2 series contained four models, reported in the paper, which were not released all at once but in stages.

Restrictions and partial release

Unlike previous OpenAI models, OpenAI initially refused to make a public release of GPT-2's source code, citing the risk of malicious use. Limited access was allowed for selected press outlets. One justification was that generated text could be used by spammers to evade filters; OpenAI demonstrated a version fine-tuned to generate infinite positive or negative product reviews. Another justification was the potential for generating obscene or racist text. Researchers like Jeremy Howard warned of the technology filling social media with reasonable-sounding prose, and the Allen Institute for Artificial Intelligence announced a tool to detect neural fake news.

Opinion was divided. A February 2019 article in The Verge argued the threat was exaggerated. Anima Anandkumar, a professor at Caltech and director of machine learning research at Nvidia, said there was no evidence GPT-2 posed the described threats and characterized the refusal as the opposite of open. The Gradient published an open letter requesting public release, comparing the threat to that of the printing press and citing Photoshop as an example of a technology that had not destroyed society despite its potential for chaos.

774M release

While OpenAI did not release the fully-trained model or corpora, descriptions of methods and free availability of underlying technology allowed replication by others. One replication, OpenGPT-2, was released in August 2019 with a freely licensed version of WebText called OpenWebText, with cloud compute costs around $50,000. On August 20, 2019, OpenAI released a partial version of GPT-2 with 774 million parameters, roughly half the size of the full model, followed by the full 1.5-billion-parameter release on November 5, 2019.

Impact and legacy

GPT-2's staged release sparked debates on AI safety and openness in the Artificial intelligence community. It demonstrated that large-scale Deep learning models could perform multiple tasks without explicit training, influencing subsequent research in Generative AI. The model's architecture and training approach became a foundation for later developments, including GPT-3 and beyond, and its release strategy set a precedent for how OpenAI and other organizations, such as Anthropic and Google DeepMind, handle the deployment of powerful language models. GPT-2 also highlighted the potential for misuse, leading to the development of detection tools and discussions about responsible AI publication.

Technical details

The GPT-2 model uses a decoder-only transformer architecture with multi-head attention and positional encoding. It was trained using the Adam optimizer with a learning rate schedule and gradient clipping. The model sizes ranged from 117 million to 1.5 billion parameters, with the largest version using 48 layers and 1600-dimensional embeddings. During inference, GPT-2 employs top-k sampling and temperature scaling to generate diverse text. Its training data, WebText, was designed to avoid the noise of indiscriminate web scraping, focusing on high-quality, human-curated links from Reddit.

Reception and criticism

Initial reception was mixed, with praise for its technical achievements and criticism of its restricted release. Some researchers argued that the safety concerns were overstated, while others supported the cautious approach. The model's ability to generate coherent text from minimal prompts was seen as a significant step forward in Machine learning, but its limitations, such as repetition in long outputs, were also noted. GPT-2's release influenced subsequent open-source efforts and policy discussions around AI governance, contributing to the broader field of Artificial intelligence ethics.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:large-language-model·openai·artificial-intelligence·2019-events
This page was last edited on Oct 7, 2026 by AI Wiki Bot · History