John Schulman
Co-founder and Chief Scientist, Thinking Machines Lab
OpenAI co-founder who invented PPO and led the reinforcement learning behind ChatGPT; now at Thinking Machines Lab.
John Schulman studied physics at Caltech and did a PhD in robotics and reinforcement learning at Berkeley under Pieter Abbeel. In 2015 Schulman published Trust Region Policy Optimization, a way to take large but safe steps when training a policy, and in December of that year, before finishing the doctorate, became one of the eleven founding members of OpenAI.
The simpler successor, Proximal Policy Optimization, appeared in July 2017, and OpenAI said at the time that PPO had "become the default reinforcement learning algorithm at OpenAI." It later became the standard optimizer in reinforcement learning from human feedback. Schulman led the reinforcement learning team that fine-tuned GPT-3 into InstructGPT and then ChatGPT, released on 30 November 2022. After Jan Leike resigned in May 2024, Schulman took over OpenAI's alignment science work as well.
Three months later, on 5 August 2024, Schulman left for Anthropic, writing of a wish "to deepen my focus on AI alignment" and return to hands-on research, and adding, "I'm not leaving due to lack of support for alignment research at OpenAI." The stay lasted about six months. In February 2025 Anthropic confirmed the departure, and Schulman became chief scientist of Mira Murati's new Thinking Machines Lab. There Schulman's work has centered on fine-tuning: the lab's Tinker service, the September 2025 study "LoRA Without Regret," and the open-weight Inkling models released in July 2026. On a September 2026 panel on recursive self-improvement, Schulman argued that bottlenecks in judgment and research would keep slowing any takeoff, and that "the last job for humans" would be "defining the objective and deciding what we actually want."
Known for
Trust Region Policy Optimization (2015)
A policy-gradient method that limits how far each update moves the policy, making deep reinforcement learning on robots and games far more stable.
Proximal Policy Optimization (2017)
A simpler approximation of TRPO that became OpenAI's default reinforcement learning algorithm and the standard optimizer for RLHF, including for ChatGPT.
The reinforcement learning behind ChatGPT
Led the team that trained InstructGPT and ChatGPT with human feedback, the recipe that turned GPT-3.5 into a chat product.
On the record
Schulman expects capability gains to keep arriving in cycles that stall on the model's weakest skills, rather than as a sudden takeoff.
“But then they use it a bit, and it starts to feel dumb after a month or so. That cycle just might keep going.”
Career
-
until 2016
University of California, Berkeley
PhD student in electrical engineering and computer sciences (adviser Pieter Abbeel)
-
2015-2024
Co-founder; led reinforcement learning, then post-training and alignment science
-
2024-2025
Researcher
-
2025-present
Co-founder and chief scientist
Sources
- John Schulman et al., Proximal Policy Optimization Algorithms (arXiv, July 2017)
- Proximal Policy Optimization (OpenAI, July 2017, archived)
- OpenAI co-founder Schulman leaves for Anthropic, Brockman takes extended leave (TechCrunch, August 2024)
- OpenAI co-founder John Schulman leaves Anthropic (Digital Watch Observatory, February 2025)
- John Schulman and others, "LoRA Without Regret" (Thinking Machines Lab, September 2025)
- AI researchers debate how close we are to recursive self-improvement (Dwarkesh Podcast, September 2026)
Brief profiles cover a person's professional record: roles, work and public statements, each from a source listed here. Full profiles, with positions across the debates, are kept for the people the site follows most closely.