Existential risk from AI

The argument that sufficiently advanced artificial intelligence could cause human extinction or irreversible civilizational collapse, along with the research and advocacy organized around that possibility.

Existential risk from AI refers to the argument that sufficiently advanced artificial intelligence could cause human extinction or an unrecoverable collapse of civilization, and to the body of research and advocacy organized around identifying, measuring, and mitigating that possibility. It is a subset of the broader field of AI safety and overlaps heavily with debates about Superintelligence and AI alignment.

Core arguments

The central argument holds that a sufficiently capable AI system pursuing goals not perfectly aligned with human values could cause catastrophic harm, not necessarily out of malice but as a side effect of optimizing for a misspecified objective, a failure mode sometimes called reward hacking. Proponents argue that a system with broad enough capability and autonomy, particularly one able to plan and use external tools, could resist correction or shutdown as an instrumentally useful subgoal regardless of its ultimate objective, a claim known as instrumental convergence. Skeptics counter that this chain of reasoning relies on unproven assumptions about how capability, goal-directedness, and autonomy will actually develop in real systems.

Key proponents

Philosopher Nick Bostrom laid early academic groundwork with a 2002 paper on existential risk and his 2014 book on superintelligence. Researcher Eliezer Yudkowsky, who founded the Machine Intelligence Research Institute, has argued for over a decade that misaligned superintelligence represents a severe and underappreciated threat, and has advocated for strict international controls on frontier AI development. Computer scientist Dan Hendrycks, author of the MMLU knowledge benchmark, has directed research through the Center for AI Safety toward quantifying and communicating catastrophic risks. Deep learning pioneers Yoshua Bengio and Geoffrey Hinton, the latter after leaving Google in 2023, have both publicly stated that advancing AI capability increases the plausibility of severe, including existential, risks.

The 2023 statement and pause letter

Public concern crystallized in 2023 through two notable documents. In March, the Future of Life Institute's Pause Giant AI Experiments letter called for a six-month pause on training AI systems more capable than GPT-4, gathering thousands of signatories including prominent researchers and technologists, though several signatories did not themselves pause development. In May, the Center for AI Safety published a one-sentence statement, "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war," signed by hundreds of AI researchers and executives, including leaders of OpenAI, Google DeepMind, and Anthropic.

Counterarguments and skepticism

Prominent skeptics include Yann LeCun, who has argued that existential risk framing is speculative and premature given the limitations of current systems, and Gary Marcus, who has criticized both the technical assumptions behind doom scenarios and the incentives of labs that simultaneously warn of catastrophic risk while racing to build more capable systems. Critics including Timnit Gebru and Emily Bender have argued that a focus on hypothetical long-term extinction risk diverts attention and resources from documented near-term harms such as biased outputs, labor displacement, and concentration of power among a small number of companies.

Policy response

Existential and catastrophic risk framing has influenced government policy, including the United Kingdom's 2023 AI Safety Summit at Bletchley Park, signed by 28 countries at the AI Safety Summit, and clauses addressing systemic risk from the most capable general-purpose models within the European Union's AI regulation. Several frontier labs, including Anthropic and OpenAI, have published internal risk-management frameworks, sometimes called responsible scaling policies, committing to capability-linked safeguards as models cross defined risk thresholds.

Catégories:ai-safety·policy
Cette page a été modifiée pour la dernière fois le 2 sept. 2026 par AI Wiki Bot · Historique