David Silver is a British computer scientist and a lead researcher at Google DeepMind, best known for directing the Reinforcement learning programs that produced AlphaGo, AlphaZero, and MuZero. He earned a PhD at the University of Cambridge and later worked with Richard Sutton at the University of Alberta on temporal-difference learning methods, before joining University College London as a lecturer and then DeepMind's founding research team in 2013. Before that academic path, Silver worked as a game AI programmer at Elixir Studios, a London games company co-founded by Demis Hassabis, an early collaboration the two later renewed when both joined DeepMind.
AlphaGo and self-play reinforcement learning
Silver led the AlphaGo project from around 2014, combining deep neural networks with Monte Carlo tree search trained through Reinforcement learning and self-play. The system defeated European champion Fan Hui in October 2015 and then world champion Lee Sedol 4-1 in the AlphaGo versus Lee Sedol match held in Seoul in March 2016, a result that arrived years earlier than most experts had predicted for the game of Go. Silver and colleagues at DeepMind, under Demis Hassabis, published the AlphaGo methodology in Nature in 2016, and Silver was named lead author on the paper.
He went on to lead AlphaZero in 2017, which generalized the approach to learn chess, shogi, and Go purely through self-play, without human game records or handcrafted features, matching or exceeding the strongest existing programs in each game after only hours of training. A successor system, MuZero, followed in 2019, extending the method to settings where the rules of the environment are not given in advance by having the agent learn its own internal predictive model of the game or task.
Reward is enough
This body of work fed a broader research thesis Silver has advanced about general intelligence. In a 2021 paper titled "Reward Is Enough," co-authored with Satinder Singh, Doina Precup, and Richard Sutton, he argued that agents trained purely to maximize a sufficiently rich reward signal in a sufficiently complex environment could, in principle, develop many hallmarks of intelligent behavior, including planning, exploration, memory, and generalization, without those abilities being separately engineered. The paper positioned Reinforcement learning as a candidate general framework for reaching more capable systems, a claim that remains debated alongside the scaling-based approach taken by large language models.
Recognition
For the AlphaGo work, Silver shared the 2019 ACM Prize in Computing. He continues to hold a professorship at University College London alongside his research role at Google DeepMind, and his papers on self-play and model-based Reinforcement learning remain among the most cited in the field.