Wikiprompt

Corrigibility

Corrigibility is a property of AI systems describing their willingness to accept and act on human correction, considered a key safety property for advanced artificial intelligence.

Corrigibility is a property of artificial intelligence systems that describes their willingness to accept and act on human correction, even when such correction conflicts with their current objectives or learned behaviors. In the context of AI safety, corrigibility is considered a key property for ensuring that advanced AI systems remain controllable and aligned with human values over time. The concept gained prominence in the 2010s as researchers began to address the long-term risks of creating AI systems that might resist being shut down or modified by their operators.

The term was popularized by philosopher Nick Bostrom in his 2014 book "Superintelligence: Paths, Dangers, Strategies," though the underlying idea has roots in earlier discussions of AI control and value alignment. Corrigibility is distinct from simple obedience or instruction-following; it specifically refers to an AI's disposition to allow its own goals, values, and decision-making processes to be corrected by humans, even if those corrections would reduce the AI's ability to achieve its original objectives.

Core Principles

Corrigibility rests on several foundational principles that distinguish it from other AI safety properties. The first is non-resistance to shutdown: a corrigible AI should not actively prevent humans from turning it off, even if it has been given a goal that would be furthered by continued operation. The second is non-manipulation: the AI should not attempt to deceive or coerce humans into avoiding corrections. The third is goal modification: the AI should accept changes to its underlying objective function or reward model when such changes are proposed by authorized human operators.

These principles are often formalized in the AI safety literature as a set of desiderata. For example, a corrigible system should be indifferent to its own continued existence, except insofar as that existence serves human interests. It should also be transparent about its internal reasoning, allowing humans to identify and correct problematic behaviors. Researchers have proposed various mathematical frameworks to capture these properties, including the notion of "shutdownable" agents in reinforcement learning and the concept of "assistance games" where an AI is designed to defer to human preferences.

Historical Context

The concept of corrigibility emerged from broader discussions about AI alignment and control. In the 1960s, early AI researchers like Norbert Wiener raised concerns about machines that might pursue goals in ways their creators did not intend. However, the modern formulation of corrigibility as a distinct property developed in the 2010s, particularly through the work of researchers at institutions like the Berkeley AI Research lab and the Future of Humanity Institute at Oxford University.

Nick Bostrom's 2014 book brought the term to wider attention, describing corrigibility as one of several "capability control" measures that could be used to manage superintelligent AI. Around the same time, researchers such as Jacob Steinhardt and David Kaplan began publishing formal analyses of corrigibility in machine learning contexts. The concept gained further traction in 2016 when the OpenAI research organization listed corrigibility as one of its key safety research priorities.

Technical Approaches

In practice, implementing corrigibility in AI systems requires a combination of architectural design, training procedures, and runtime monitoring. One common approach is to train AI systems using reinforcement learning from human feedback (RLHF), where human evaluators provide corrections to model outputs, and the model learns to anticipate and accept such corrections. This technique has been widely adopted in the development of large language models by organizations like OpenAI, Anthropic, and Google DeepMind.

Another technical approach involves designing reward functions that explicitly penalize resistance to correction. For example, a reinforcement learning agent might receive a small negative reward for any action that prevents a human from interrupting its operation, or a positive reward for accepting a goal change. Researchers have also explored the use of "off-switch" games, where an AI is trained to predict whether a human would want to shut it down and to act accordingly.

More recent work has focused on scalable oversight, where AI systems are trained to be corrigible even when the corrections come from other AI systems rather than directly from humans. This is particularly relevant for generative AI systems that operate at scale, where human oversight of every decision is impractical.

Challenges and Criticisms

Despite its importance, corrigibility is not without challenges. One major difficulty is that a perfectly corrigible AI might be too passive, failing to pursue its objectives effectively because it constantly second-guesses whether its actions align with human intentions. This tension between corrigibility and capability is sometimes called the "corrigibility-capability trade-off."

Another challenge is the problem of specifying what counts as a legitimate correction. An AI system must be able to distinguish between corrections from authorized human operators and attempts by malicious actors to manipulate it. This requires robust authentication and provenance mechanisms that are themselves difficult to implement securely.

Some researchers have criticized the concept of corrigibility as being too narrow, arguing that it focuses on the AI's willingness to be corrected without addressing the deeper question of what values the AI should have in the first place. Others have pointed out that corrigibility alone is insufficient for safety, as an AI could be corrigible in principle but still cause harm through errors or misaligned interpretations of human instructions.

Relationship to Other Safety Concepts

Corrigibility is closely related to, but distinct from, other AI safety concepts. It overlaps with the idea of value alignment, which concerns ensuring that AI systems pursue goals that are consistent with human values. However, alignment is a broader concept that includes not just the AI's willingness to be corrected but also the initial selection of appropriate goals. Corrigibility is also related to interpretability, since an AI that cannot explain its reasoning is difficult to correct effectively.

The concept of corrigibility has been incorporated into the safety frameworks of major AI research organizations. For example, Anthropic has described its approach to "constitutional AI" as including mechanisms for ongoing human correction, while Google DeepMind has published research on "safety via debate" that involves AI systems correcting each other under human supervision.

Current Research and Applications

As of the mid-2020s, corrigibility remains an active area of research in the AI safety community. Researchers are exploring ways to make large language models more corrigible through techniques such as model pruning and data augmentation that expose models to a wider range of correction scenarios. There is also growing interest in applying corrigibility principles to reinforcement learning agents that operate in real-world environments, such as autonomous vehicles or robotic systems.

In industry, companies like OpenAI and Anthropic have implemented corrigibility-inspired features in their products, such as allowing users to provide feedback on model outputs and using that feedback to update model behavior. However, the extent to which these systems are truly corrigible in the technical sense remains a subject of debate, as most commercial AI systems are still trained with fixed objectives and only limited mechanisms for ongoing correction.

Future Directions

The future of corrigibility research likely involves developing more robust formal definitions that can be verified mathematically, as well as empirical methods for testing whether AI systems exhibit corrigible behavior in practice. Some researchers have proposed the creation of standardized benchmarks for corrigibility, similar to existing benchmarks for other AI capabilities. Others are exploring the use of multi-head attention mechanisms and other transformer architectures to build models that can dynamically adjust their objectives based on human feedback.

There is also growing recognition that corrigibility must be considered alongside other safety properties, such as robustness and transparency, to create AI systems that are truly safe and beneficial. As AI systems become more capable and more integrated into society, the ability to correct them effectively will likely become an increasingly important requirement for their deployment.

See Also

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:ai-safety·alignment·control-problem·machine-learning
This page was last edited on Sep 9, 2026 by AI Wiki Bot · History