Some people change the history of a field from the inside, working steadily on an idea the rest of the world has not yet caught up to. Richard S. Sutton is one of them. A Canadian computer science researcher, he spent much of his career building the foundations of what we now call reinforcement learning: the branch of AI that teaches a machine to learn through trial and error, the same way a person learns when they play a game, drive a car, or negotiate. In the 1980s and 1990s, that approach was not the field's favorite. The dominant current placed its bets on systems based on logic and expert-written rules. Sutton chose a different path: let the machine learn from its own experience, rewarding what works and correcting what does not.
In 1988 he published the paper "Learning to Predict by the Methods of Temporal Differences," formalizing TD-learning (temporal difference learning): a mechanism that lets a system adjust its predictions gradually, evaluating each decision with the benefit of what came after. Together with his colleague Andrew Barto, he built the field's foundational textbook, "Reinforcement Learning: An Introduction" (first edition, 1998), used today by researchers and AI teams worldwide. And in 2019 he published "The Bitter Lesson," a short essay that shook the community. In it he argues that 70 years of AI research show that general methods leveraging computational power always win in the long run over methods that try to encode human knowledge. It was uncomfortable to read. It was also correct.
Sutton earned his PhD from the University of Massachusetts Amherst in 1984 and, after years in private research labs, became a professor at the University of Alberta, where he founded the RLAI (Reinforcement Learning and Artificial Intelligence) lab. The recognitions he accumulated speak for themselves: AAAI Fellow since 2001, Fellow of the Royal Society of Canada in 2016, Fellow of the Royal Society of London in 2021. And in 2024, ACM awarded him the A.M. Turing Award, the highest honor in computer science, together with Andrew Barto, "for developing the conceptual and algorithmic foundations of reinforcement learning." The Turing Award comes with a million dollars and, in this case, with decades of vindication.
Sutton's lesson is not only technical: it is an invitation to intellectual humility. At its core, "The Bitter Lesson" says that our intuitions about how intelligence works can be more of an obstacle than a guide. Letting data, experience, and computation speak, rather than imposing what we believe we know, is exactly the disposition we need to use today's AI tools well. Every time a system learns from its own mistakes, whether it is a robot, a chess player, or a language model improving through feedback, there is an invisible thread leading back to that Canadian researcher who chose to trust experience when almost no one else would.
Official links for Richard S. Sutton, The father of reinforcement learning
More figures who shaped AI in AI Legends, or back to the news.