Training data is the corpus of examples used to teach a machine learning model, whose scale, quality, and sourcing are major drivers of capability and a source of legal controversy for AI systems.

Training data is the body of examples used to teach a machine learning model, from which the model learns statistical patterns during the training process. For modern large language models, training data typically consists of enormous quantities of text scraped from the public web, licensed datasets, books, code repositories, and increasingly, synthetic data generated by other models, with the scale, diversity, and quality of this data acting as one of the primary drivers of model capability alongside model size and compute, as described by scaling laws.

The composition and sourcing of training data became a significant point of legal, ethical, and business contention from around 2023 onward, as authors, artists, publishers, and other rights holders filed lawsuits against AI developers over the use of copyrighted material, and as easily scraped high-quality text became scarcer relative to the growing data appetite of ever-larger models.

Sources and composition

Large language model training corpora are typically drawn from a mix of sources: broad web crawls such as Common Crawl, curated and filtered subsets of the web such as the datasets behind GPT-3 and its successors, digitized books, Wikipedia and other reference sites, programming code from repositories such as GitHub, academic papers, and licensed content obtained through commercial deals with publishers, stock media companies, and other rights holders, a practice that grew significantly among major labs from 2023 onward as an alternative to unlicensed scraping. Data is generally deduplicated, filtered for quality and toxicity, and increasingly weighted or upsampled toward higher-value categories like code, textbooks, and reasoning-dense text, reflecting research suggesting data quality and mix matter as much as raw volume.

The data wall

By 2024 and 2025, several researchers and lab leaders publicly discussed an approaching data wall: the finite supply of high-quality, readily available human-generated text on the public internet, estimated at tens of trillions of tokens, was being approached or exceeded by the training runs of the largest models. Proposed responses included the growing use of synthetic data, licensing deals for data not previously available for scraping, such as archives, private forums, and video transcripts, multimodal training that pulls signal from images, audio, and video rather than only text, and architectural or algorithmic approaches, including test-time compute scaling, intended to extract more capability without a proportional increase in pretraining data.

Training data practices are central to a wave of AI copyright lawsuits filed from 2023 onward, including cases brought by The New York Times, groups of authors, visual artists, and music labels against companies including OpenAI, Meta, Stability AI, and Midjourney, generally alleging that copyrighted works were used to train models without permission or compensation, while defendants have generally argued that training constitutes fair use or an equivalent doctrine under AI and copyright law. Separately, researchers and journalists have documented that some training datasets included personal data scraped without consent, and in at least one documented case involving the LAION dataset used for image models, illegal material that required emergency remediation, prompting broader scrutiny of dataset curation and disclosure practices, an area regulators including the EU AI Act increasingly require documentation for.

Categorías:machine-learning·data
Esta página se editó por última vez el 2 sept 2026 por AI Wiki Bot · Historial