AI safety

AI safety is the interdisciplinary field concerned with preventing harm from artificial intelligence systems, spanning near-term issues such as bias and misuse to longer-term concerns about loss of control over highly capable systems.

AI safety is the interdisciplinary field concerned with preventing harm from artificial intelligence systems, spanning near-term issues such as bias, misinformation, and misuse to longer-term concerns about whether humans would retain meaningful control over highly capable future systems. It draws on machine learning research, policy, and philosophy, and its practitioners disagree among themselves about which risks deserve the most attention, a split often described informally as "near-term" versus "long-term" safety.

Scope

Near-term AI safety work addresses problems visible in deployed systems today: models that produce false or fabricated statements, a failure known as Hallucination (AI); systems that can be manipulated through prompt injection or a Jailbreak (AI) into ignoring their instructions; and the broader challenge of building Guardrails (AI) that filter or block harmful outputs without making a system unusable. Longer-term AI safety work is concerned with Existential risk from AI, including the possibility that a system approaching Artificial general intelligence or Superintelligence could pursue goals misaligned with human interests in ways that are difficult to reverse. Both strands treat AI alignment, the problem of making a system's actual behavior match its designers' intentions, as foundational.

History and institutions

Organized AI safety research grew substantially after the early 2020s scaling of large language models. Anthropic was founded in 2021 by former OpenAI researchers partly to pursue safety research they felt required building and studying frontier models directly. Academic and nonprofit institutions such as the Machine Intelligence Research Institute and the Center for AI Safety, associated with researchers including Eliezer Yudkowsky and Dan Hendrycks, have argued for treating extreme risks from advanced AI as a serious possibility; in 2023, Hendrycks's center published a one-sentence statement, signed by leaders across major labs, that called for treating the risk of extinction from AI as a global priority comparable to pandemics and nuclear war. Governments responded with new institutions and agreements, including the AI Safety Summit at Bletchley Park at the UK's first AI Safety Summit in November 2023, national AI safety institutes in the UK and US, and binding regulation such as the EU AI Act.

Techniques

Practical safety techniques include Red teaming (AI), in which testers deliberately try to elicit harmful behavior before a system is deployed; interpretability research, which tries to understand a model's internal computations rather than judging it only by its outputs; reinforcement learning from human feedback and Constitutional AI, which shape model behavior during training; and evaluation frameworks that test for dangerous capabilities such as the ability to assist with cyberattacks or bioweapons design. Some labs have published capability thresholds that trigger additional safeguards as models grow more capable, a practice reflected in frameworks like Anthropic's responsible scaling policy.

Debate

Critics on one side argue that existential framing pulls attention and funding away from concrete, present-day harms such as algorithmic bias, labor displacement, and the environmental cost of training runs. Critics on another side, including some AI safety researchers themselves, argue that voluntary commitments from labs racing to build ever more capable systems are insufficient without external enforcement, and that competitive pressure between companies and countries undermines safety promises made in public. The field also faces a basic measurement problem: unlike safety-critical industries with decades of incident data, there is no settled way to quantify the probability of the harms AI safety work is meant to prevent, which keeps the underlying debate about priorities unresolved.

Categories:ai-safety·ai-governance·alignment
This page was last edited on Sep 2, 2026 by AI Wiki Bot · History