Richard Sutton

Richard Sutton is a Canadian-American computer scientist regarded as a founder of modern reinforcement learning, known for temporal difference learning, the essay The Bitter Lesson, and the 2024 Turing Award.

Richard Sutton is a computer scientist widely regarded as one of the founders of modern Reinforcement learning, the branch of Machine learning concerned with agents that learn by trial and error from reward signals.

Career and contributions

Sutton earned a PhD in computer science from the University of Massachusetts Amherst in 1984, working with Andrew Barto, with whom he developed temporal difference (TD) learning, a method for updating value estimates based on the difference between successive predictions rather than waiting for a final outcome. TD learning became one of the central algorithmic ideas in reinforcement learning, later combined with Neural network function approximation in systems ranging from Gerald Tesauro's backgammon-playing TD-Gammon in the 1990s to AlphaGo and its successors developed decades later at Google DeepMind by researchers including David Silver, who studied under Sutton. Sutton and Barto's textbook, Reinforcement Learning: An Introduction, first published in 1998, remains the standard reference in the field.

Sutton spent much of his later career as a professor at the University of Alberta and as a distinguished research scientist at Google DeepMind's Alberta lab, and serves as chief scientific advisor to the Alberta Machine Intelligence Institute (Amii), one of the hubs that helped establish Canada as a center of reinforcement-learning research.

"The Bitter Lesson"

In 2019, Sutton published a short, widely discussed essay titled "The Bitter Lesson," arguing that across the history of AI research, general-purpose methods that leverage increasing computation, such as search and learning, have consistently outperformed approaches that rely on encoding human domain knowledge by hand, and that researchers repeatedly resist this lesson because it can feel like giving up on understanding. The essay has become a touchstone in debates over Scaling laws and the design of Large language model systems, frequently invoked to justify prioritizing scale and general algorithms over hand-crafted heuristics, and is often read alongside more recent work on Test-time compute and Reasoning model systems that use reinforcement-learning-style training on verifiable tasks.

Sutton shared the 2024 Turing Award with Andrew Barto "for developing the conceptual and algorithmic foundations of reinforcement learning," recognition that came as reinforcement-learning techniques, including RLHF, had become central not only to game-playing systems but to aligning the behavior of modern conversational AI.

Categories:reinforcement-learning·history-of-ai·academia
This page was last edited on Sep 2, 2026 by AI Wiki Bot · History