The latest in AI, every dayAI News
Andrew Barto

AI Legend · Researcher · University of Massachusetts Amherst

Andrew Barto

The architect of reinforcement learning

August 8, 2026

When Andrew Barto arrived at the University of Massachusetts Amherst in 1977, artificial intelligence was obsessed with rules: if A happens, do B. The dominant approach was to hand-code all the knowledge a machine would ever need. Barto had a different idea, at once more humble and more profound: what if we taught machines to learn on their own, the way animals do, through rewards and punishments? Most of his colleagues were not convinced. He kept going anyway.

Through the 1980s, working alongside his then-student Richard Sutton, Barto laid the mathematical foundations of reinforcement learning (RL). Their most influential contribution was temporal difference (TD) learning, published in 1988: an algorithm that lets an agent update its predictions about the future as experience accumulates, without waiting for the final outcome. As early as 1981 they had shown that this mechanism could explain learning behaviors that existing models could not; a 1995 neuroscience study then confirmed that dopamine neurons in the human brain operate on exactly that same logic. What Barto described in equations, nature had already invented millions of years earlier.

Together with Sutton he wrote the field's reference text in 1998: "Reinforcement Learning: An Introduction" (MIT Press), cited more than 75,000 times and still the standard textbook at universities worldwide. That work, along with the policy-gradient methods he also helped develop, is the backbone of systems like AlphaGo (DeepMind, 2016), which beat the world Go champion, and the human-feedback alignment algorithms that make models like ChatGPT and Claude possible. In March 2025, the ACM awarded him the 2024 Turing Award, the "Nobel of computing," shared with Sutton, recognizing that ideas born four decades ago are the hidden engine of today's AI.

What I love most about Barto's story is not the Turing Award, well-deserved as it is. It is that he spent decades doing basic research, with no guarantee of application, convinced that understanding how living things learn was the key to building smarter machines. He was right. Every time you use an AI that improves with use, that adjusts its responses based on your feedback, that learns from its own mistakes, there is an echo of his work inside it. And if you take one thing from this profile, let it be this: the most powerful ideas often look too simple at first.


More figures who shaped AI in AI Legends, or back to the news.