# John Schulman

John Schulman is an American machine learning researcher and OpenAI co-founder known for developing the Proximal Policy Optimization algorithm and leading the reinforcement learning from human feedback work behind ChatGPT.

John Schulman is a machine learning researcher best known for developing widely used [reinforcement-learning](https://www.wikiprompt.org/wiki/reinforcement-learning) algorithms and for leading the training methodology behind OpenAI's early chat models. He earned a PhD at the University of California, Berkeley, working in a robotics and reinforcement learning lab, and was one of the founding researchers of [openai](https://www.wikiprompt.org/wiki/openai) when it launched in December 2015.

## Research contributions

Schulman's most cited technical contribution is Proximal Policy Optimization (PPO), introduced in a 2017 paper he led. PPO simplified earlier trust-region methods into an algorithm that was easier to implement and tune, and it became one of the most widely deployed policy-gradient algorithms in both robotics and, later, language model training. PPO subsequently became the default optimization algorithm used in [rlhf](https://www.wikiprompt.org/wiki/rlhf) pipelines, including the one behind [chatgpt](https://www.wikiprompt.org/wiki/chatgpt).

## Role at OpenAI

At OpenAI, Schulman led the team that applied RLHF to large language models, work that produced InstructGPT in 2022 and fed directly into the release of ChatGPT in November 2022 (see [chatgpt-launch](https://www.wikiprompt.org/wiki/chatgpt-launch)). He was widely credited internally and externally as the technical architect of the alignment technique that made [gpt-3](https://www.wikiprompt.org/wiki/gpt-3)-era models follow instructions reliably rather than merely continue text. He continued to lead post-training and alignment research through the [gpt-4](https://www.wikiprompt.org/wiki/gpt-4) era, working alongside colleagues such as [ilya-sutskever](https://www.wikiprompt.org/wiki/ilya-sutskever) and [sam-altman](https://www.wikiprompt.org/wiki/sam-altman).

## Departure to Anthropic and Thinking Machines Lab

In August 2024, Schulman announced he was leaving OpenAI to join [anthropic](https://www.wikiprompt.org/wiki/anthropic), stating he wanted to focus more directly on AI alignment research and hands-on technical work rather than management. The move was notable because it came less than a year after OpenAI's November 2023 board crisis (see [openai-board-crisis](https://www.wikiprompt.org/wiki/openai-board-crisis)) and was read by many observers as part of a broader wave of senior researchers moving toward safety-focused labs. He worked on alignment and post-training methods at Anthropic through 2024.

In 2025, Schulman left Anthropic to join Thinking Machines Lab, the startup founded by former OpenAI chief technology officer [mira-murati](https://www.wikiprompt.org/wiki/mira-murati), as a co-founder and chief scientist. The move placed him alongside several other OpenAI alumni at a company that raised one of the largest seed rounds in AI startup history.

## Influence

Schulman is regarded as one of the researchers most responsible for turning [reinforcement-learning](https://www.wikiprompt.org/wiki/reinforcement-learning) from human feedback into a practical, production-scale technique rather than a research curiosity, a shift that underpins nearly every consumer chatbot released after 2022. His career trajectory across OpenAI, Anthropic, and Thinking Machines Lab has also been cited as a case study in how top AI research talent circulates among frontier labs.

---
Source: https://www.wikiprompt.org/wiki/john-schulman
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-02T20:31:59.889949+00:00
