The Center for Human-Compatible Artificial Intelligence (CHAI) is a research center at the University of California, Berkeley focused on developing methods to ensure advanced artificial intelligence systems are aligned with human values and intentions. Founded in 2016, the center addresses the long-term challenge of creating AI that remains beneficial as its capabilities grow, with particular emphasis on value alignment and controllable machine behavior.
The center was established by a group of academics led by Stuart J. Russell, a Berkeley computer science professor and co-author of the widely used textbook Artificial Intelligence: A Modern Approach. Russell has been a prominent voice in AI safety discussions, arguing that the current paradigm of optimizing fixed objectives may produce systems whose behavior diverges from genuine human preferences.
Research Focus
CHAI's core research strategy centers on value alignment, particularly through inverse reinforcement learning. In this approach, AI systems infer human values not from explicit instructions or reward functions but by observing human behavior. The goal is to create AI that can learn what people genuinely want, including preferences they may have difficulty articulating directly.
The center has also investigated the dynamics of human-machine interaction in scenarios where an intelligent agent has an "off-switch" it could override. This line of research examines how AI might be designed to remain deferential to human control, essentially teaching systems to recognize when they should allow themselves to be switched off.
Faculty and Collaborators
CHAI's faculty membership spans multiple institutions. Alongside Russell, Berkeley-based members include Pieter Abbeel and Anca Dragan. The center also includes Bart Selman and Joseph Halpern from Cornell University, Michael Wellman and Satinder Singh Baveja from the University of Michigan, and Tom Griffiths and Tania Lombrozo from Princeton University. This interdisciplinary composition brings together expertise in computer science, cognitive science, and philosophy.
The center has established partnerships beyond academia, including collaborations with the World Economic Forum and its Global AI Council. These connections aim to translate research findings into policy discussions and industry practices.
Funding History
Initial funding for CHAI came in 2016, when the Open Philanthropy Project recommended that Good Ventures provide $5,555,550 in support over five years. Since then, CHAI has received additional grants from OpenPhil and Good Ventures exceeding $12,000,000 in total. This sustained financial backing has allowed the center to support graduate students, postdoctoral researchers, and multi-institution research efforts.
Publications and Output
Researchers affiliated with CHAI have published extensively on topics including reward learning, corrigibility, and the limits of specification in AI systems. The center's work has appeared in major machine-learning and AI conferences, and its findings have informed broader debates about existential risk from artificial general intelligence. Russell's 2019 book Human Compatible synthesizes many of the center's ideas, presenting in accessible form the argument that AI should be designed with uncertainty about human preferences as a core principle.
Impact and Legacy
The center has contributed to the growth of AI safety as a recognized field within deep learning and AI research. Its emphasis on rigorous formal frameworks has influenced how researchers approach problems such as reward hacking and specification gaming. As large language models and other advanced systems have become more capable, the questions CHAI originally posed about alignment have moved from theoretical concern to practical engineering challenge, with implications for organizations like OpenAI and Anthropic.
See Also
- Existential risk from artificial general intelligence
- Machine Intelligence Research Institute
- Future of Humanity Institute
- Future of Life Institute
External Links
Official website