Some people change AI through headlines, and some change it through mathematics. John Schulman belongs to the second group. Trained as a physicist (Caltech, 2010) and with a PhD in computer engineering from Berkeley (2016), he spent his early research years obsessed with a question that seemed almost philosophical: how to make an AI system learn from human feedback in a stable and efficient way, without updates destabilizing what it had already learned.
In 2015, as one of the six original co-founders of OpenAI, he published TRPO (Trust Region Policy Optimization), an algorithm that for the first time guaranteed that each update to a reinforcement learning neural network would not "break" what it had already learned. Two years later, in 2017, PPO (Proximal Policy Optimization) arrived: simpler, more robust, equally powerful. PPO quickly became the default method across the field and, more importantly, it was the engine behind RLHF (reinforcement learning from human feedback) that gives ChatGPT, Claude, and virtually all modern AI assistants their ability to follow instructions, be helpful, and not go off the rails. When you write to an AI and it responds coherently and aligned with what you asked for, there is a little piece of PPO working beneath the screen.
At OpenAI for nearly nine years, he led the post-training team for the most important models: GPT-4, GPT-4o, and ChatGPT. In August 2024, he left the company he co-founded to join Anthropic with the goal of focusing more closely on the AI alignment problem, that central challenge of making AI systems safe and genuinely beneficial. The stint was brief: in February 2025, he joined Thinking Machines Lab, Mira Murati's new company, as Chief Scientist.
The lesson I take from John Schulman is subtle but powerful: sometimes the most transformative contribution is not the one that appears on the cover of a magazine, but the one that solves a concrete technical problem that enables everything else. PPO is not famous outside the field, but without it, the path to the conversational AI we use today would have been much rockier. Every time Claude gives you a response that seems to understand exactly what you needed, it is using, at some level, the patient work of this physicist turned pioneer of reinforcement learning.
Official links for John Schulman, The architect of human reinforcement learning
More figures who shaped AI in AI Legends, or back to the news.