# Orthogonality thesis

The orthogonality thesis is the claim, associated with Nick Bostrom, that an AI system's level of intelligence is independent of its final goals: arbitrarily capable systems can pursue arbitrarily simple or alien objectives.

The orthogonality thesis is a claim in the philosophy of [AI safety](https://www.wikiprompt.org/wiki/ai-safety) stating that intelligence and final goals are independent, or orthogonal, dimensions: in principle, almost any level of capability can be combined with almost any terminal objective. A [superintelligent](https://www.wikiprompt.org/wiki/superintelligence) system, on this view, would not automatically converge on human-friendly values; it could apply extraordinary competence to a trivial or harmful goal.

## Origin and statement

The thesis was articulated by philosopher [Nick Bostrom](https://www.wikiprompt.org/wiki/nick-bostrom) in his 2012 paper "The Superintelligent Will" and developed in his 2014 book Superintelligence, building on earlier arguments by [Eliezer Yudkowsky](https://www.wikiprompt.org/wiki/eliezer-yudkowsky) and the rationalist community. Bostrom paired it with the instrumental convergence thesis: whatever its final goal, a sufficiently capable agent tends to acquire resources, preserve itself and resist shutdown, because those subgoals help achieve almost any objective. The classic illustration is the paperclip maximizer, a hypothetical system that converts everything into paperclips with flawless competence and no malice.

## Role in AI risk arguments

Together, orthogonality and instrumental convergence form the backbone of the case that [existential risk from AI](https://www.wikiprompt.org/wiki/existential-risk-from-ai) does not require malevolent machines, only capable ones with imperfectly specified goals. This motivates [alignment](https://www.wikiprompt.org/wiki/alignment) research: if values do not emerge from intelligence by default, they must be engineered in. The argument influenced safety agendas at [Anthropic](https://www.wikiprompt.org/wiki/anthropic) and [OpenAI](https://www.wikiprompt.org/wiki/openai) and public warnings by researchers including [Geoffrey Hinton](https://www.wikiprompt.org/wiki/geoffrey-hinton) and [Yoshua Bengio](https://www.wikiprompt.org/wiki/yoshua-bengio).

## Criticism

Critics argue the thesis is a claim about logical possibility, not likelihood: systems trained on human data, such as [large language models](https://www.wikiprompt.org/wiki/large-language-model) shaped by [RLHF](https://www.wikiprompt.org/wiki/rlhf), may absorb human-like goal structures in practice. Others, including [Yann LeCun](https://www.wikiprompt.org/wiki/yann-lecun), contend that goals are engineered choices and that the paperclip framing exaggerates the autonomy of real systems. Defenders reply that observed [reward hacking](https://www.wikiprompt.org/wiki/reward-hacking) in deployed models is small-scale evidence of the underlying dynamic.

---
Source: https://www.wikiprompt.org/wiki/orthogonality-thesis
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-02T22:00:28.97427+00:00
