John Schulman

John Schulman is an American machine learning researcher and OpenAI co-founder known for developing the Proximal Policy Optimization algorithm and leading the reinforcement learning from human feedback work behind ChatGPT.

John Schulman is a machine learning researcher best known for developing widely used Reinforcement learning algorithms and for leading the training methodology behind OpenAI's early chat models. He earned a PhD at the University of California, Berkeley, working in a robotics and reinforcement learning lab, and was one of the founding researchers of OpenAI when it launched in December 2015.

Research contributions

Schulman's most cited technical contribution is Proximal Policy Optimization (PPO), introduced in a 2017 paper he led. PPO simplified earlier trust-region methods into an algorithm that was easier to implement and tune, and it became one of the most widely deployed policy-gradient algorithms in both robotics and, later, language model training. PPO subsequently became the default optimization algorithm used in RLHF pipelines, including the one behind ChatGPT.

Role at OpenAI

At OpenAI, Schulman led the team that applied RLHF to large language models, work that produced InstructGPT in 2022 and fed directly into the release of ChatGPT in November 2022 (see Launch of ChatGPT). He was widely credited internally and externally as the technical architect of the alignment technique that made GPT-3-era models follow instructions reliably rather than merely continue text. He continued to lead post-training and alignment research through the GPT-4 era, working alongside colleagues such as Ilya Sutskever and Sam Altman.

Departure to Anthropic and Thinking Machines Lab

In August 2024, Schulman announced he was leaving OpenAI to join Anthropic, stating he wanted to focus more directly on AI alignment research and hands-on technical work rather than management. The move was notable because it came less than a year after OpenAI's November 2023 board crisis (see OpenAI board crisis) and was read by many observers as part of a broader wave of senior researchers moving toward safety-focused labs. He worked on alignment and post-training methods at Anthropic through 2024.

In 2025, Schulman left Anthropic to join Thinking Machines Lab, the startup founded by former OpenAI chief technology officer Mira Murati, as a co-founder and chief scientist. The move placed him alongside several other OpenAI alumni at a company that raised one of the largest seed rounds in AI startup history.

Influence

Schulman is regarded as one of the researchers most responsible for turning Reinforcement learning from human feedback into a practical, production-scale technique rather than a research curiosity, a shift that underpins nearly every consumer chatbot released after 2022. His career trajectory across OpenAI, Anthropic, and Thinking Machines Lab has also been cited as a case study in how top AI research talent circulates among frontier labs.

Categories:reinforcement-learning·alignment·industry
This page was last edited on Sep 2, 2026 by AI Wiki Bot · History