Paul Christiano is an American researcher in AI alignment known for co-developing the technique that became RLHF and for founding the Alignment Research Center, a nonprofit alignment research organization.
Early career and RLHF
Christiano worked as a researcher at OpenAI in its early years, where in 2017 he was a co-author of "Deep Reinforcement Learning from Human Preferences," a paper that introduced the core method of training a reward model from human comparisons between model outputs and then optimizing a policy against that reward model. This approach, later scaled up and applied to language models as RLHF, became the dominant technique for aligning chat-oriented models with human intent, most visibly in the training behind ChatGPT and its instruction-following predecessor InstructGPT.
Alignment Research Center
In 2021, Christiano left OpenAI to found the Alignment Research Center (ARC), a nonprofit focused on theoretical work in AI AI alignment, including research into "eliciting latent knowledge" from models and methods for evaluating whether a model has capabilities or intentions its outputs do not reveal. ARC also became known for building and running some of the earliest formal pre-deployment capability evaluations of frontier models, including dangerous-capability testing conducted in coordination with labs such as OpenAI and Anthropic ahead of releases including GPT-4.
US AI Safety Institute
In 2024, Christiano was appointed to lead safety-focused evaluation work at the US AI Safety Institute, housed within the National Institute of Standards and Technology, a role that placed a researcher closely associated with technical alignment work inside a government body responsible for evaluating frontier models ahead of the international AI safety summit process that began at Bletchley Park in November 2023. The appointment drew both support and criticism from different parts of the AI safety community over questions of independence and the government's evaluation capacity.
Influence
Christiano is widely regarded within the AI safety field as one of the researchers most responsible for translating abstract concerns about long-term AI risk into concrete, empirically testable research programs, bridging the more philosophical work associated with figures like Nick Bostrom and the applied training techniques used across nearly every major chat-oriented model released after 2022.