Jan Leike is an AI alignment researcher who co-led OpenAI's Superalignment team and later joined Anthropic after a high-profile resignation from OpenAI in 2024.
Career at DeepMind and OpenAI
Leike holds a PhD in reinforcement learning theory from the Australian National University. Before joining OpenAI, he worked as a research scientist at DeepMind on AI safety and AI alignment problems. At OpenAI, he worked extensively on techniques for training models to be more helpful and honest, contributing to the RLHF research program that shaped ChatGPT and later co-authoring OpenAI's work on scalable oversight and weak-to-strong generalization.
Superalignment and resignation
In July 2023, OpenAI announced the Superalignment team, co-led by Leike and OpenAI cofounder Ilya Sutskever, with a stated goal of solving the technical problem of aligning a hypothetical future superintelligent AI system within four years, backed by a commitment of significant computing resources. The team's work touched on Mechanistic interpretability, scalable oversight, and automated alignment research.
In May 2024, Leike resigned from OpenAI, publicly stating that "safety culture and processes have taken a backseat to shiny products" at the company and that his team had been struggling to get the compute resources it had been promised. His resignation came in the same week as Sutskever's own departure and drew significant attention to internal tensions at OpenAI over the balance between rapid product deployment and longer-term safety research, tensions that had also surfaced during the November 2023 OpenAI board crisis.
Move to Anthropic
Shortly after leaving OpenAI, Leike joined Anthropic to continue alignment research, joining a growing cohort of researchers, including earlier arrivals such as Jared Kaplan and Chris Olah, who had moved from OpenAI to Anthropic over disagreements about safety prioritization. At Anthropic, he has continued work related to scalable oversight and model evaluation.
Influence
Leike's public resignation letter became one of the most widely circulated documents in the AI safety community's discussion of whether frontier labs' internal incentives reliably support their stated safety commitments, a debate closely tied to broader questions in AI governance and long-term AI risk. His departure, following similar exits by other safety-focused researchers, was frequently cited as evidence that internal alignment teams at fast-moving commercial labs face structural pressure to compete with product deadlines for compute and headcount.