# Coherent extrapolated volition

Coherent extrapolated volition (CEV) is a theoretical AI alignment framework proposed by Eliezer Yudkowsky in 2004, describing how an artificial superintelligence could act on an idealized, extrapolated version of human preferences rather than current ones.

Coherent extrapolated volition (CEV) is a theoretical framework in the field of AI alignment describing an approach by which an artificial superintelligence (ASI) would act on a benevolent supposition of what humans could want if they were more knowledgeable, more rational, had more time to think, and had matured together as a society, as opposed to humanity's current individual or collective preferences. It was proposed by Eliezer Yudkowsky in 2004 as part of his work on friendly AI.

The concept addresses the challenge of specifying goals for advanced AI systems that may surpass human intelligence. Rather than encoding a fixed set of rules or values, CEV suggests that an ASI should derive its objectives by extrapolating the volition of humanity, aiming to reflect what people could desire under ideal epistemic and moral conditions.

## Concept

CEV proposes that an advanced AI system should derive its goals by extrapolating the volition of humanity. This means aggregating and projecting human preferences into a process that reflects what people could desire under ideal epistemic and moral conditions. The aim is to ensure that AI systems do not trespass on humanity's true interests; do not follow transient or poorly informed preferences, nor decide issues on which humanity's will is as yet unpredictable.

In poetic terms, our coherent extrapolated volition is our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together; where the extrapolation converges rather than diverges, where our wishes cohere rather than interfere; extrapolated as we wish that extrapolated, interpreted as we wish that interpreted.

The framework is designed to be humane and self-correcting by capturing the source of human values instead of trying to list them. It avoids the difficulty of laying down an explicit, fixed list of rules. It encapsulates moral growth, preventing flawed current moral beliefs from getting locked in. It limits the influence that a small group of programmers can have on what the ASI would value, thus also reducing the incentives to build ASI first. And it keeps humanity in charge of its destiny.

## Debate

Yudkowsky and Nick Bostrom note that CEV has several interesting properties, but it also faces significant theoretical and practical challenges. Bostrom notes that CEV has "a number of free parameters that could be specified in various ways, yielding different versions of the proposal." One such parameter is the extrapolation base (whose extrapolated volition is taken into account). For example, whether it should include people with severe dementia, patients in a vegetative state, foetuses, or embryos. He also notes that if CEV's extrapolation base only includes humans, there is a risk that the result would be ungenerous toward other animals and digital minds. One possible solution would be to include a mechanism to expand CEV's extrapolation base.

These debates are part of the broader discourse on [AI alignment](https://www.wikiprompt.org/wiki/ai-alignment) and [AI safety](https://www.wikiprompt.org/wiki/ai-safety), which examines how to ensure that advanced AI systems act in accordance with human values and interests.

## Variants and alternatives

A proposed theoretical alternative to CEV is to rely on an artificial superintelligence's superior cognitive capabilities to figure out what is morally right, and let it act accordingly. It is also possible to combine both techniques, for instance with the ASI following CEV except when it is morally impermissible.

In another review, a philosophical analysis explores CEV through the lens of social trust in autonomous systems. Drawing on Anthony Giddens' concept of "active trust", the author proposes an evolution of CEV into "Coherent, Extrapolated and Clustered Volition" (CECV). This formulation aims to better reflect the moral preferences of diverse cultural groups, thus offering a more pragmatic ethical framework for designing AI systems that earn public trust while accommodating societal diversity.

## Relation to other AI concepts

CEV is situated within the broader field of [artificial intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence) research, particularly in discussions about how to design systems that are beneficial to humanity. It contrasts with approaches that rely on explicit rule-based systems or on learning from human feedback, such as [RLHF](https://www.wikiprompt.org/wiki/rlaif) (reinforcement learning from human feedback), which is used in modern [large language models](https://www.wikiprompt.org/wiki/large-language-model). While RLHF aims to align models with current human preferences through iterative feedback, CEV proposes a more ambitious extrapolation of those preferences.

The framework also relates to questions about rationality and decision-making in AI systems. It assumes that there is a coherent set of preferences that humans would converge upon given ideal conditions, which is a contested philosophical assumption.

## See also

- [AI alignment](https://www.wikiprompt.org/wiki/ai-alignment)
- [AI safety](https://www.wikiprompt.org/wiki/ai-safety)
- [Artificial intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence)
- Rationality

## External links

- [Wikipedia: Coherent extrapolated volition](https://en.wikipedia.org/wiki/Coherent_extrapolated_volition)

---
Source: https://www.wikiprompt.org/wiki/coherent-extrapolated-volition
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-14T04:25:45.219979+00:00
