# David Silver

[Frontier Minds profile](https://frontierminds.ai/people/david-silver/)

Last verified: 2026-09-19

Positions and narratives are editorial summaries. Quotes are attributed quotations. Source dates and profile verification dates are distinct. Missing positions mean not recorded, not agreement or neutrality.

Founder and CEO of Ineffable Intelligence; Professor of Computer Science, UCL · Ineffable Intelligence

London, UK. Born 1976

Led AlphaGo, AlphaZero and MuZero at DeepMind, co-wrote Reward is enough and The Era of Experience, and in 2026 founded Ineffable Intelligence to build a superintelligence that learns from its own experience.

## Overview

### Games first

David Silver was born on 26 September 1976 in Hanham, near Bristol, and read computer science at Christ's College, Cambridge, graduating in 1997 with the Addison-Wesley award. At Cambridge he got to know Demis Hassabis, and when Hassabis founded the London games developer Elixir Studios in July 1998, Silver became its chief technology officer and the lead programmer on Republic: The Revolution, a political strategy game that Elixir built under its three-game publishing deal with Eidos Interactive.

He later described the five years he spent programming games as his first job, and said he left because everything people were doing in games was short-term fixes rather than long-term vision. He wanted to work on intelligence itself. In 2004 he returned to academia and wrote to Richard Sutton, whose textbook he had read, asking whether Sutton would supervise a PhD on computer Go. Sutton, as Silver tells it, replied that if he was still alive he would be happy to.


### Alberta and the search for Go

At the University of Alberta, then the strongest reinforcement learning group in the world, Silver combined reinforcement learning with Monte Carlo tree search, the technique that samples thousands of random games to evaluate a position. With Sylvain Gelly he co-introduced the algorithms used in the first master-level 9x9 Go programs, and the idea that a learned value function could steer a simulation-based search ran through his 2009 thesis, Reinforcement Learning and Simulation-Based Search in Computer Go. The ACM later noted that Go had remained his continuing research interest ever since.

In 2011 he was awarded a Royal Society University Research Fellowship and became a lecturer at University College London. His 2015 UCL course, ten lectures from Markov decision processes to a case study of reinforcement learning in classic games, was recorded and posted online under a CC BY-NC licence, and it is still how many researchers first meet the field.


### DeepMind, Atari and AlphaGo

Silver consulted for DeepMind from its founding in 2010 and joined full time in 2013, going on to lead its reinforcement learning research group. His first major result there was the Deep Q-Network: a single network that learned to play Atari games from raw pixels and a score. The 2015 Nature paper describing it had been cited nearly 10,000 times by the time the ACM gave him its Prize in Computing in April 2020.

He then led AlphaGo. The program's neural networks were first trained on human expert games, then improved by reinforcement learning from self-play, and a tree search combined the two. In October 2015 it beat the European champion Fan Hui 5-0, a result the Nature paper called a feat previously thought to be at least a decade away. In March 2016, in a Seoul hotel with millions watching, it beat Lee Sedol 4-1. Move 37 of game two, a shoulder hit on the fifth line that commentators first took for a mistake, became the emblem of the match. Silver told Wired the next morning that AlphaGo had estimated a one-in-ten-thousand chance that a human would play it, and had played it anyway.


### Removing the human data

The next question was whether the human games had been necessary at all. AlphaGo Zero, published in Nature in October 2017, started from random play with no human data and no knowledge beyond the rules, and beat the version that had defeated Lee Sedol 100-0. AlphaZero, in Science in December 2018, applied the same algorithm to chess and shogi as well as Go and beat the strongest existing program in each, including the chess engine Stockfish. MuZero, in Nature in 2020, went a step further by learning its own model of the environment, so that it could plan without being told the rules, and it matched AlphaZero on board games while also mastering Atari. He co-led AlphaStar, which reached grandmaster level in StarCraft II in 2019, extending self-play to a real-time game with hidden information.

Not everything he touched was a reinforcement learning problem. Asked in 2024 about his part in AlphaFold, he said one of his biggest contributions had been to encourage the team to stop viewing protein folding as a reinforcement learning problem and treat it as supervised learning. He was also an author on AlphaProof, which in 2024 became the first program to reach medal standard at the International Mathematical Olympiad, and on the 2023 Gemini technical report. In 2021, with Satinder Singh, Doina Precup and Sutton, he published Reward is enough in the journal Artificial Intelligence. The paper's claim, that intelligence and all its associated abilities can be understood as serving the maximisation of reward, drew a twelve-author rebuttal titled Scalar reward is not enough. After the pandemic he had long COVID and, by his own account, was not well for a year or two, which is why his conference output thinned in that period.


### The era of experience

In April 2025 Silver and Sutton published Welcome to the Era of Experience, a preprint of a chapter for the MIT Press book Designing an Intelligence. Its argument is that imitating humans can reproduce human competence but not exceed it, that the high-quality human data able to improve a strong model has largely been consumed in mathematics, coding and science, and that the next generation of agents will learn predominantly from streams of their own experience, with rewards grounded in the environment rather than in human judgement. On Google DeepMind's podcast that month, as vice president of reinforcement learning, he made the same case to Hannah Fry with AlphaGo and AlphaZero as the worked examples.

Ineffable Intelligence was incorporated in London in November 2025 while Silver was on sabbatical from Google DeepMind. He was appointed its director on 16 January 2026 and never formally returned to his DeepMind role; the company told staff of his departure that month, and a spokesperson said his contributions had been invaluable. A note he wrote for himself when deciding to leave, dated 15 January 2026 and posted on the company's site, says the world needs a place where the full ambition of the reinforcement learning paradigm can flourish, and calls the mission his life's work.


### Ineffable Intelligence

By 19 February 2026 Sutton was describing Ineffable as a $4 billion company. On 27 April 2026 it announced a $1.1 billion seed round at a $5.1 billion post-money valuation, co-led by Sequoia Capital and Lightspeed Venture Partners with Nvidia, DST Global, Index, Google, EQT Ventures, the Wellcome Trust, the British Business Bank and the UK's Sovereign AI Fund among the participants, the largest seed financing in Europe. The company's stated mission is to make first contact with superintelligence by building a superlearner that discovers all knowledge from its own experience, from elementary motor skills to profound intellectual breakthroughs, and it says superintelligence can be built within years rather than decades. Silver has committed, through Founders Pledge, to give away all of the money he makes from his Ineffable equity to charities that save the most lives.

On 16 June 2026 Ineffable named Google Cloud its infrastructure partner and said it would deploy one of the largest clusters of Nvidia Vera Rubin NVL72 systems on the platform. On 7 September 2026 Fortune reported that the company had added six cofounders, four of them former Google DeepMind colleagues, including Junhyuk Oh to run reinforcement learning and Wojciech Czarnecki to oversee the science team, along with Alexandre Laterre from InstaDeep and Heather Gorham from Flying Fish. Silver remains a professor at UCL.


## Timeline

### Born in Hanham, England

September 1976

Born 26 September 1976, near Bristol.

### BA, Christ's College, Cambridge

June 1997

Graduated with the Addison-Wesley award; knew Demis Hassabis from university.

### CTO of Elixir Studios

July 1998

Joined Hassabis's new London games studio as chief technology officer and lead programmer on Republic, The Revolution.

### PhD student under Richard Sutton

September 2004

Left the games industry for the University of Alberta and worked on Monte Carlo tree search for Go with Sylvain Gelly.

### PhD, University of Alberta

June 2009

Thesis: Reinforcement Learning and Simulation-Based Search in Computer Go.

### Royal Society University Research Fellowship

October 2011

Took up the fellowship at UCL and became a lecturer; recorded his ten-lecture reinforcement learning course in 2015.

### Joined DeepMind full time

June 2013

Had consulted for the company since its founding in 2010; later led its reinforcement learning research group.

### Deep Q-Network in Nature

February 2015

One network learned dozens of Atari games from pixels; the paper had nearly 10,000 citations within five years.

### AlphaGo in Nature

January 2016

Led the project; the Nature paper reported the 5-0 win over Fan Hui.

### AlphaGo beats Lee Sedol 4-1

March 2016

The match in Seoul was the first defeat of a world champion Go player by a program.

### AlphaGo Zero

October 2017

Learned Go from random play with no human data and beat the Lee Sedol version 100-0.

### UCL inaugural lecture

May 2018

Gave his inaugural lecture as a UCL professor on AlphaZero.

### AlphaZero in Science

December 2018

One algorithm mastered chess, shogi and Go by self-play.

### AlphaStar

October 2019

Co-led the grandmaster-level StarCraft II agent.

### ACM Prize in Computing

April 2020

Received the 2019 prize, with $250,000 from Infosys.

### Fellow of the Royal Society

May 2021

Elected FRS for his contributions to Deep Q-Networks and AlphaGo.

### Reward is enough

October 2021

Published the reward-is-enough hypothesis with Singh, Precup and Sutton, which drew a formal rebuttal.

### Welcome to the Era of Experience

April 2025

Essay with Sutton on learning from experience rather than human data.

### Ineffable Intelligence incorporated

November 2025

Incorporated in London while Silver was on sabbatical from Google DeepMind.

### Left Google DeepMind

January 2026

Ended his Google DeepMind role to run Ineffable Intelligence.

### Record seed round

April 2026

Ineffable Intelligence raised $1.1 billion at a $5.1 billion post-money valuation, co-led by Sequoia Capital and Lightspeed Venture Partners.

### Google Cloud partnership

June 2026

Ineffable named Google Cloud its infrastructure partner.

### Six cofounders added

September 2026

Ineffable Intelligence added six cofounders.

## Key contributions

### [AlphaGo](https://www.nature.com/articles/nature16961)

Silver led the project that combined policy and value networks, trained first on human expert games and then by self-play, with Monte Carlo tree search. Its 5-0 win over Fan Hui in 2015 was the first defeat of a professional on a full-size board, and its 4-1 win over Lee Sedol in March 2016 made move 37 a byword for machine creativity.


### [AlphaGo Zero and AlphaZero](https://arxiv.org/abs/1712.01815)

He showed the human data could be removed. AlphaGo Zero, starting tabula rasa and trained only to predict its own moves and its own game outcomes, beat the Lee Sedol version 100-0. AlphaZero generalised the recipe to chess and shogi and beat the strongest program in each, including Stockfish, which is why the ACM cited it as a demonstration of generality in game-playing methods.


### [Deep Q-Network](https://www.nature.com/articles/nature14236)

With DeepMind colleagues he combined deep convolutional networks with Q-learning so that one agent could learn many Atari games from pixels and score alone. The 2015 Nature paper was the launch of deep reinforcement learning as a field and was cited nearly 10,000 times in five years.


### [MuZero and AlphaStar](https://www.nature.com/articles/s41586-020-03051-4)

MuZero learned a model of its environment and planned inside it, matching AlphaZero on board games without being told the rules and also mastering Atari. AlphaStar, which he co-led, reached grandmaster level in StarCraft II, a real-time game with hidden information and multiple agents.


### [Monte Carlo tree search for Go](https://www.ucl.ac.uk/engineering/events/2018/may/inaugural-lecture-david-silver)

His Alberta thesis joined reinforcement learning to simulation-based search, and with Sylvain Gelly he co-introduced the algorithms behind the first master-level 9x9 Go programs. AlphaGo's search was the descendant of this work.


### [The reward-is-enough hypothesis](https://doi.org/10.1016/j.artint.2021.103535)

With Singh, Precup and Sutton he argued in 2021 that intelligence and its associated abilities can be understood as subserving the maximisation of reward, and that powerful reinforcement learning agents could constitute a solution to artificial general intelligence. A twelve-author response argued that scalar reward cannot capture multi-objective behaviour and is unsafe as a basis for general intelligence.


### [The era of experience](https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf)

His 2025 essay with Sutton argues that the knowledge extractable from human data is approaching a limit in mathematics, coding and science, and sets out four shifts for the next generation of agents: lifelong streams of experience, grounded actions, grounded rewards and non-human reasoning. It is the research charter of Ineffable Intelligence.


### [Teaching reinforcement learning](https://davidstarsilver.wordpress.com/teaching/)

His ten UCL lectures, recorded in 2015 and released under a CC BY-NC licence, became the standard video introduction to reinforcement learning, and the accompanying Easy21 assignment is still set in courses elsewhere.


## Signature ideas

### [Reward is enough](https://doi.org/10.1016/j.artint.2021.103535)

The hypothesis, stated with Satinder Singh, Doina Precup and Richard Sutton in 2021, is that intelligence and all the abilities we associate with it, from perception and language to social intelligence and imitation, can be understood as serving the maximisation of a reward by an agent in a rich environment. On this view there is no need for a separate problem formulation for each ability; an agent that learns by trial and error to maximise reward will acquire whichever abilities the environment demands. The idea grew out of the games work: AlphaZero was given only a win-or-lose signal and developed opening theory, tactics and endgame technique on its own. The consequence Silver draws is that powerful reinforcement learning agents could be a solution to general intelligence, which is the bet behind Ineffable Intelligence. The claim is contested. Peter Vamplew and eleven co-authors argued in Scalar reward is not enough that a single number cannot represent the many competing objectives of biological or artificial agents, and that building general intelligence around scalar reward carries unacceptable risks of unsafe or unethical behaviour. Others note that the paper is a hypothesis about explanation rather than a demonstration that reward-maximising agents can be built for open-ended environments.


### [Experience beats imitation](https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf)

Silver's most repeated lesson is that a system trained to imitate humans can at best match them, and that AlphaGo Zero, which learned with no human games, beat the human-trained AlphaGo 100-0. Welcome to the Era of Experience, written with Sutton in April 2025, turns that lesson into a forecast for the whole field: the high-quality human data able to improve a strong model is nearly used up in mathematics, coding and science, so the next gains must come from data the agent generates itself by acting in an environment, at a scale that will eventually dwarf human data. The essay lists four shifts that follow: agents that live in lifelong streams of experience rather than short chats, actions and observations grounded in the environment rather than in dialogue, rewards grounded in the environment rather than in human pre-judgement, and reasoning that need not be in human terms. AlphaProof is the worked example, having generated a hundred million formal proofs after starting from a hundred thousand human ones. The counter-argument, made by Yann LeCun among others, is that reinforcement learning is sample-inefficient and should be the last, thin layer on top of learning from observation, not the foundation.


### [Search plus learning, with the system as its own teacher](https://www.nature.com/articles/nature24270)

Every one of Silver's game programs pairs a learned evaluation with a lookahead search, and uses the search to improve the learner. In AlphaGo the tree search made the neural networks stronger at play time; in AlphaGo Zero the search results became the training targets, so that the network was trained to predict what the search would choose and who would win, and the improved network then made the search better in the next iteration. MuZero closed the loop by learning the model the search plans in, so that the same recipe worked on Atari, where the rules are not given. The origin is his Alberta thesis on simulation-based search, and the consequence is a general method that has since been reused, in modified forms, in mathematical theorem proving. Gary Marcus has objected that calling these systems tabula rasa overstates the case, because the rules, the search procedure and the network architecture are all built in.


### [Creativity is discovery beyond human convention](https://www.wired.com/2016/03/two-moves-alphago-lee-sedol-redefined-future/)

Silver defines creativity as discovering something not known before, something unexpected and outside our norms, and he argues that self-play is the essence of it, because a system playing against itself is not bound by what humans have done. His evidence is move 37 of AlphaGo's second game against Lee Sedol, which the program judged a human would play with probability one in ten thousand and which turned out to be strong. The stakes of the idea are large: if machines can be creative in this sense, then the scientific breakthroughs that lie beyond current human knowledge are reachable by learning from experience, which is the promise Ineffable's mission statement makes when it says its superlearner should rediscover and then transcend language, science and mathematics. Critics reply that Go is a closed world with a perfect simulator and a crisp reward, and that the open world offers neither.


### [Choose problems you might fail at](https://www.talkrl.com/episodes/david-silver-rcl-2024/transcript)

Asked how he picks research problems, Silver said he looks for ones just within range, where he believes the chance of success is at most fifty percent, and that too much of the field aims lower. The pattern runs from Go, which he chose as a PhD topic when programs were weak, to StarCraft, to protein folding, where his contribution was to argue against his own tools, to a company whose founding note accepts a significant risk of failure for a chance of spectacular success. The habit explains the shape of his career better than any single algorithm does.


## Notable works

### [Human-level control through deep reinforcement learning](https://www.nature.com/articles/nature14236)

paper · 2015

The Deep Q-Network paper in Nature; cited nearly 10,000 times by April 2020.

### [Reinforcement Learning lecture course, UCL](https://davidstarsilver.wordpress.com/teaching/)

course · 2015

Ten lectures from Markov decision processes to reinforcement learning in classic games, with the Easy21 assignment.

### [Mastering the game of Go with deep neural networks and tree search](https://www.nature.com/articles/nature16961)

paper · 2016

The AlphaGo paper, Silver first author; reports the 5-0 win over Fan Hui.

### [Mastering the game of Go without human knowledge](https://www.nature.com/articles/nature24270)

paper · 2017

AlphaGo Zero, which learned from self-play alone and beat the original 100-0.

### [Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm](https://arxiv.org/abs/1712.01815)

paper · 2017

The AlphaZero preprint, December 2017; the peer-reviewed version appeared in Science in December 2018.

### [Grandmaster level in StarCraft II using multi-agent reinforcement learning](https://www.nature.com/articles/s41586-019-1724-z)

paper · 2019

AlphaStar, which Silver co-led.

### [David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning (Lex Fridman Podcast \#86)](https://www.youtube.com/watch?v=uPUEq8d73JI)

podcast · 2020

An hour and 48 minutes on self-play, creativity, reward and how he came to study with Sutton.

### [Mastering Atari, Go, chess and shogi by planning with a learned model](https://www.nature.com/articles/s41586-020-03051-4)

paper · 2020

MuZero plans with a model it learned itself, without being told the rules.

### [Reward is enough](https://doi.org/10.1016/j.artint.2021.103535)

paper · 2021

With Singh, Precup and Sutton in Artificial Intelligence, vol. 299.

### [Welcome to the Era of Experience](https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf)

essay · 2025

With Richard Sutton; a preprint of a chapter in the MIT Press book Designing an Intelligence.

### [Is human data enough? (Google DeepMind: The Podcast)](https://www.youtube.com/watch?v=zzXyPGEtseI)

podcast · 2025

Fifty minutes with Hannah Fry, released 10 April 2025, on moving from human data to experience.

### [Olympiad-level formal mathematical reasoning with reinforcement learning](https://www.nature.com/articles/s41586-025-09833-y)

paper · 2025

The AlphaProof paper in Nature, November 2025; Silver is among the authors, the project was led by Hubert, Mehta and Sartran.

## Where to start

### [Is human data enough? (Google DeepMind: The Podcast with Hannah Fry)](https://www.youtube.com/watch?v=zzXyPGEtseI)

podcast

Fifty minutes in which Silver explains the whole argument in plain language, from AlphaGo Zero to the era of experience. Start here.

### [Welcome to the Era of Experience](https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf)

essay

The written version of his current position, about a dozen pages including a section on safety; read it to see exactly what Ineffable Intelligence is trying to build.

### [AlphaGo (documentary, 2017)](https://www.youtube.com/watch?v=WXuK6gekU1Y)

talk

Ninety minutes inside the Lee Sedol match with Silver, Hassabis and Fan Hui on camera; the best account of what move 37 felt like in the room.

### [Lex Fridman Podcast](https://www.youtube.com/watch?v=uPUEq8d73JI)

podcast

Just under two hours covering his path from games to Alberta to DeepMind, with his definitions of self-play, creativity and reward.

### [Reinforcement Learning lecture course, UCL](https://davidstarsilver.wordpress.com/teaching/)

course

Ten lectures of roughly ninety minutes each; the standard way to learn the technical foundations his work rests on.

### [Reward is enough](https://doi.org/10.1016/j.artint.2021.103535)

paper

A readable position paper rather than a technical one; pair it with the Vamplew rebuttal to see the strongest objections.

## Awards

### Royal Academy of Engineering Silver Medal

2017

### Mensa Foundation Prize

2017

### Marvin Minsky Medal (IJCAI)

2018

Awarded for the AlphaGo work.

### ACM Prize in Computing

2019

For breakthrough advances in computer game-playing; announced 1 April 2020 with a $250,000 prize endowed by Infosys.

### Fellow of the Royal Society (FRS)

2021

### Fellow of the Association for the Advancement of Artificial Intelligence (AAAI)

2022

For significant contributions to machine learning and game theory, and the application of deep learning to game playing.

## Education

### PhD in Computing Science (reinforcement learning and simulation-based search in computer Go)

2009 · University of Alberta

### BA in Computer Science

1997 · Christ's College, University of Cambridge

## Perspectives

### Limits of learning from human data

Editorial summary: Argues that imitating human data can reproduce human competence but not exceed it, and that in mathematics, coding and science the useful human data has largely been consumed.

> A new generation of agents will acquire superhuman capabilities by learning predominantly from experience.

Source: [Welcome to the Era of Experience, with Richard Sutton, 2025](https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf)

How this view has changed: The direct descendant of the AlphaGo Zero result of 2017, when removing the human games made the program stronger, and of the 2021 reward-is-enough hypothesis.

### Reinforcement learning as the path to general intelligence

Editorial summary: Holds that reward maximisation by an agent acting in its environment is sufficient to account for intelligence, and that reinforcement learning will sit at the core of any general system.

> powerful reinforcement learning agents could constitute a solution to artificial general intelligence

Source: [Reward is enough, abstract, 2021](https://doi.org/10.1016/j.artint.2021.103535)

How this view has changed: Said in 2020 that reinforcement learning would be at the core of any human-level system; formalised it as a hypothesis in 2021; by 2026 he was building a company on it.

### Large language models and superintelligence

Editorial summary: Believes systems trained to imitate human knowledge cannot go beyond it, and that a different method is needed for superintelligence.

> We want to go beyond what humans know, and to do that we're going to need a different type of method

Source: [Fortune, quoting the Google DeepMind podcast, 2026](https://www.fortune.com/2026/01/30/google-deepmind-ai-researcher-david-silver-leaves-to-found-ai-startup-ineffable-intelligence)

### When not to use reinforcement learning

Editorial summary: Says reinforcement learning is not the right frame for every problem, and cites his own advice that AlphaFold be treated as supervised learning.

> not all problems are best suited to RL. You really have to find the problems which are natively better understood in a different way.

Source: [TalkRL interview at the Reinforcement Learning Conference, August 2024, 2024](https://www.talkrl.com/episodes/david-silver-rcl-2024/transcript)

### Self-play and creativity

Editorial summary: Argues that a system learning by trial and error can discover things no human knew, and that self-play is the essence of machine creativity.

> creativity means discovering something which wasn't known before, something unexpected

Source: [Lex Fridman Podcast, 2020](https://podcasts.happyscribe.com/lex-fridman-podcast-artificial-intelligence-ai/86-david-silver-alphago-alphazero-and-deep-reinforcement-learning)

### AI safety

Editorial summary: Accepts that agents learning from experience raise new safety risks needing research, but argues that an agent which observes and adapts to its environment can also be safer than a fixed system.

> further research is surely required to ensure a safe transition into the era of experience

Source: [Welcome to the Era of Experience, with Richard Sutton, 2025](https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf)

### Timelines

Editorial summary: States as a company belief that superintelligence can be built within years, and that its knowledge will be too profound to describe in human language.

> Superintelligence can be built within years, not decades or centuries.

Source: [Ineffable Intelligence, Beliefs, 2026](https://www.ineffable.ai/)

### Who should benefit

Editorial summary: Has pledged, through Founders Pledge, to give away all the money he makes from his Ineffable equity.

> Any money that I make from Ineffable will go to high-impact charities that save as many lives as possible.

Source: [TechCrunch, 2026](https://techcrunch.com/2026/04/27/deepminds-david-silver-just-raised-1-1b-to-build-an-ai-that-learns-without-human-data/)

## Predictions

### Superintelligence can be built within years rather than decades, and it will come from agents learning from their own experience rather than from human data.

> Superintelligence can be built within years, not decades or centuries.

Made: Apr 2026. Listed under "Beliefs" on the website of Ineffable Intelligence, the company Silver founded, which TechCrunch described as "newly launched" when the company announced its funding on April 27, 2026; the Internet Archive's first capture with this text is from May 1, 2026

[Source](https://www.ineffable.ai/)

Window closes: No stated window

Status: Too early to tell

Verdict: "Within years" sets no fixed date, so no resolution date is given; a reading of under ten years would put the test in the mid-2030s at the latest. As of September 2026 Ineffable had raised a 1.1 billion dollar seed round at a 5.1 billion dollar valuation (April 2026) and added six cofounders, four of them Silver's former Google DeepMind colleagues (September 2026), and had not published a system. The belief extends the April 2025 paper "Welcome to the Era of Experience," written with Richard Sutton, which argued that knowledge from human data "is rapidly approaching a limit" in mathematics, coding and science and that "the transition to the era of experience is imminent."


## Critics and counterpoints

### Is scalar reward enough?

Their view: A single reward signal, maximised by an agent in a sufficiently rich environment, is enough to explain and eventually to produce every ability associated with intelligence.

The case against: Peter Vamplew, Benjamin Smith, Diederik Roijers, Richard Dazeley and eight co-authors argue that biological and artificial agents balance many objectives that no single scalar captures, that the hypothesis is unfalsifiable as stated, and that pursuing general intelligence through scalar reward maximisation risks unsafe or unethical behaviour when the reward is misspecified. Alignment researchers make the related point that reward misspecification is the central failure mode of reinforcement learning, not a detail.

Where it stands: Silver and Sutton's 2025 essay concedes that rewards for open-world agents cannot come from human data alone and proposes adapting the reward function through experience, which critics read as a partial retreat from a fixed scalar signal.

[Source](https://arxiv.org/abs/2112.15422)

### Human data versus experience

Their view: Human data is a ceiling; superhuman ability in science and mathematics requires agents that generate their own data by acting in the world.

The case against: Yann LeCun agrees that today's language models will not reach human-level intelligence, but has argued since 2016 that if intelligence is a cake, self-supervised learning from observation is the bulk of it and reinforcement learning is the cherry on top, too sample-inefficient to carry the load. Others note that the largest recent gains in reasoning came from reinforcement learning applied on top of language models pretrained on human text, which suggests the two eras are complementary rather than sequential.

Where it stands: Silver's essay itself cites DeepSeek's reinforcement-learned reasoning as evidence the transition has begun inside language models; the disagreement is now about whether pretraining on human data remains the foundation or becomes a temporary scaffold.

[Source](https://arxiv.org/abs/2502.03038)

### Should anyone build an agentic superintelligence?

Their view: An endlessly learning agent that discovers knowledge from experience is the goal, it can be built within years, and it can and must be built to benefit humanity.

The case against: Yoshua Bengio and twelve co-authors argue that autonomous, goal-directed agents pose risks ranging from misuse to an irreversible loss of human control, and propose a non-agentic Scientist AI that explains and predicts without pursuing goals. On that view, an agent with grounded rewards and lifelong autonomy is the precise design safety researchers want to avoid, not a path to safety.

Where it stands: Silver's essay acknowledges that experiential learning will increase certain safety risks and calls for research on a safe transition, while arguing that an agent which notices human distress and adapts is safer than a frozen model; the two camps have not converged.

[Source](https://arxiv.org/abs/2502.15657)

### How much is really learned from scratch?

Their view: AlphaGo Zero and AlphaZero started tabula rasa, with no human data or domain knowledge beyond the rules, and reached superhuman play.

The case against: Gary Marcus argued in 2018 that the claim is overstated, because the rules, the Monte Carlo tree search procedure, the convolutional architecture and the self-play curriculum are all innate structure supplied by the designers, and that the field should study what to build in rather than pretend nothing is. Sceptics add that perfect simulators and unambiguous rewards exist for board games and few real problems.

Where it stands: Silver has not disputed that the algorithm and rules are given; the argument has shifted to whether the era-of-experience programme can supply grounded rewards and environments rich enough to stand in for the simulator.

[Source](https://arxiv.org/abs/1801.05667)

## Misconceptions

Claim: AlphaGo learned to play Go without any human data.

Correction: The AlphaGo that beat Lee Sedol was first trained by supervised learning on human expert games and then improved by self-play. The version that learned from nothing but the rules was AlphaGo Zero, published in October 2017.

[Source](https://www.nature.com/articles/nature24270)

Claim: Silver co-founded DeepMind.

Correction: DeepMind was founded in 2010 by Demis Hassabis, Shane Legg and Mustafa Suleyman. Silver consulted for it from the start and joined full time in 2013.

[Source](https://www.ucl.ac.uk/engineering/events/2018/may/inaugural-lecture-david-silver)

Claim: Ineffable Intelligence is valued at $4 billion.

Correction: The $4 billion figure was the pre-money valuation reported in February 2026 while the round was being raised. The $1.1 billion seed round announced on 27 April 2026 set a post-money valuation of $5.1 billion.

[Source](https://www.cnbc.com/2026/04/27/deepmind-ineffable-intelligence-record-seed-funding-nvidia-google.html)

Claim: Silver led AlphaProof.

Correction: He is an author on the AlphaProof paper and cites it in the era-of-experience essay, but the project was led by Thomas Hubert, Rishi Mehta and Laurent Sartran.

[Source](https://www.nature.com/articles/s41586-025-09833-y)

## Quotes

> The reinforcement learning problem is to simply take actions over time so as to maximize that reward signal.

Source: [Lex Fridman Podcast, 2020](https://podcasts.happyscribe.com/lex-fridman-podcast-artificial-intelligence-ai/86-david-silver-alphago-alphazero-and-deep-reinforcement-learning)

> the only way to address them in any complex system is to give the system the ability to correct its own errors

Source: [Lex Fridman Podcast, 2020](https://podcasts.happyscribe.com/lex-fridman-podcast-artificial-intelligence-ai/86-david-silver-alphago-alphazero-and-deep-reinforcement-learning)

> AlphaGo learned to discover new strategies for itself, by playing millions of games between its neural networks, against themselves, and gradually improving

Source: [Wired, quoting Silver at AlphaGo's unveiling, 2016](https://www.wired.com/2016/03/two-moves-alphago-lee-sedol-redefined-future/)

> I try and choose problems where I believe that the chance of success is at most 50%.

Source: [TalkRL interview at the Reinforcement Learning Conference, 2024](https://www.talkrl.com/episodes/david-silver-rcl-2024/transcript)

> We stand on the threshold of a new era in artificial intelligence that promises to achieve an unprecedented level of ability.

Source: [Welcome to the Era of Experience, with Richard Sutton, 2025](https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf)

> In key domains such as mathematics, coding, and science, the knowledge extracted from human data is rapidly approaching a limit.

Source: [Welcome to the Era of Experience, with Richard Sutton, 2025](https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf)

> The world needs a place where the full ambition of the reinforcement learning paradigm can flourish.

Source: [A Note from Dave, 15 January 2026, on the Ineffable Intelligence site, 2026](https://www.ineffable.ai/)

> Our mission is to make first contact with superintelligence

Source: [Statement reported by CNBC, 2026](https://www.cnbc.com/2026/04/27/deepmind-ineffable-intelligence-record-seed-funding-nvidia-google.html)

## Related leaders

- [demis-hassabis](https://frontierminds.ai/people/demis-hassabis/): Cambridge contemporary who hired him at Elixir Studios in 1998 and brought him to DeepMind, where Silver led AlphaGo and the reinforcement learning group for more than a decade; co-author of the AlphaGo, AlphaGo Zero and MuZero papers.

- [yann-lecun](https://frontierminds.ai/people/yann-lecun/): Agrees that scaling language models will not reach human-level intelligence, but has argued since 2016 that reinforcement learning is the cherry on the cake rather than the foundation Silver takes it to be.

- [yoshua-bengio](https://frontierminds.ai/people/yoshua-bengio/): Argues for a non-agentic Scientist AI as a safer alternative to exactly the kind of autonomous, goal-seeking superintelligence Ineffable Intelligence set out to build.

- [ilya-sutskever](https://frontierminds.ai/people/ilya-sutskever/): Co-author, from Google Brain, of the January 2016 AlphaGo paper in Nature on which Silver was lead author.

- [gary-marcus](https://frontierminds.ai/people/gary-marcus/): His January 2018 paper "Innateness, AlphaZero, and Artificial Intelligence" argued that the "tabula rasa" framing of AlphaGo Zero understated the search, architecture and rules built into it.

- [mustafa-suleyman](https://frontierminds.ai/people/mustafa-suleyman/): DeepMind co-founder who ran the lab's product and applied work while Silver led the reinforcement learning research behind AlphaGo.

## Affiliations

- Ineffable Intelligence (Founder and CEO, 2026-present; incorporated November 2025)
- University College London (Professor of Computer Science; Royal Society University Research Fellow from 2011)
- Google DeepMind (consultant from 2010; full time 2013-2026; led the reinforcement learning research group, latterly as VP of Reinforcement Learning)
- Royal Society (Fellow, 2021)
- Academy for the Mathematical Sciences (Fellow, first cohort)
- AAAI (Fellow)
- University of Alberta (PhD under Richard Sutton, 2004-2009)
- Elixir Studios (CTO and lead programmer, 1998-2004)

## Areas of focus

- reinforcement learning
- self-play and search
- world models and planning
- learning from experience
- superintelligence

## Tags

- reinforcement learning
- self-play
- game AI
- planning
- superintelligence

## Organizations

- ineffable-intelligence: Founder and CEO (2025–present)

- deepmind: Principal Research Scientist, then VP of Reinforcement Learning (2013–2026)

## Links

- [Homepage](https://www.davidsilver.uk/)

- [Ineffable Intelligence](https://www.ineffable.ai/)

- [Google Scholar](https://scholar.google.com/citations?user=-8DNE4UAAAAJ&hl=en)

- [Royal Society profile](https://royalsociety.org/people/david-silver-35033/)

- [UCL reinforcement learning course](https://davidstarsilver.wordpress.com/teaching/)

- [website](https://www.davidsilver.uk/)

- [youtube](https://www.youtube.com/watch?v=2pWv7GOvuf0)

## Sources

1. [David Silver (computer scientist) - Wikipedia](https://en.wikipedia.org/wiki/David_Silver_%28computer_scientist%29)

2. [Professor David Silver FRS - Royal Society](https://royalsociety.org/people/david-silver-35033/)

3. [ACM Prize in Computing awarded to AlphaGo developer - EurekAlert (ACM release, 1 April 2020)](https://www.eurekalert.org/news-releases/730838)

4. [Inaugural Lecture: David Silver, 23 May 2018 - UCL Engineering](https://www.ucl.ac.uk/engineering/events/2018/may/inaugural-lecture-david-silver)

5. [David Silver - Heidelberg Laureate Forum](https://www.heidelberg-laureate-forum.org/laureate/david-silver/)

6. [Professor David Silver FRS FAcadMathSci - Academy for the Mathematical Sciences](https://www.acadmathsci.org.uk/team/member/professor-david-silver-frs/)

7. [Elected AAAI Fellows - AAAI](https://aaai.org/about-aaai/aaai-awards/the-aaai-fellows-program/elected-aaai-fellows/)

8. [Elixir Studios - Wikipedia](https://en.wikipedia.org/wiki/Elixir_Studios)

9. [Mastering the game of Go with deep neural networks and tree search - Nature, 2016](https://www.nature.com/articles/nature16961)

10. [Mastering the game of Go without human knowledge - Nature, 2017](https://www.nature.com/articles/nature24270)

11. [Olympiad-level formal mathematical reasoning with reinforcement learning - Nature, November 2025](https://www.nature.com/articles/s41586-025-09833-y)

12. [Welcome to the Era of Experience - Silver and Sutton, April 2025 (PDF)](https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf)

13. [Reward is enough - Artificial Intelligence, vol. 299, 2021 (DOI)](https://doi.org/10.1016/j.artint.2021.103535)

14. [Scalar reward is not enough: A response to Silver, Singh, Precup and Sutton (2021) - arXiv](https://arxiv.org/abs/2112.15422)

15. [Lex Fridman Podcast \#86 transcript - HappyScribe](https://podcasts.happyscribe.com/lex-fridman-podcast-artificial-intelligence-ai/86-david-silver-alphago-alphazero-and-deep-reinforcement-learning)

16. [David Silver @ RCL 2024 transcript - TalkRL, 26 August 2024](https://www.talkrl.com/episodes/david-silver-rcl-2024/transcript)

17. [In Two Moves, AlphaGo and Lee Sedol Redefined the Future - Wired, March 2016](https://www.wired.com/2016/03/two-moves-alphago-lee-sedol-redefined-future/)

18. [Is human data enough? \| David Silver - Google DeepMind on YouTube, 10 April 2025](https://www.youtube.com/watch?v=zzXyPGEtseI)

19. [Google DeepMind on X, 10 April 2025 - David Silver, VP of Reinforcement Learning, on the podcast](https://x.com/GoogleDeepMind/status/1910363683215008227)

20. [Longtime Google DeepMind researcher David Silver leaves to found his own AI startup - Fortune, 30 January 2026](https://www.fortune.com/2026/01/30/google-deepmind-ai-researcher-david-silver-leaves-to-found-ai-startup-ineffable-intelligence)

21. [Google DeepMind's David Silver departs to found AI startup - The Decoder, 31 January 2026](https://the-decoder.com/google-deepmind-pioneer-david-silver-departs-to-found-ai-startup-betting-llms-alone-wont-reach-superintelligence/)

22. [Richard Sutton on X, 19 February 2026 - on Ineffable Intelligence](https://x.com/RichardSSutton/status/2024291626420752437)

23. [Ineffable Intelligence - Mission, Beliefs and A Note from Dave](https://www.ineffable.ai/)

24. [Former Google DeepMind researcher's AI startup raises record $1.1 billion seed funding - CNBC, 27 April 2026](https://www.cnbc.com/2026/04/27/deepmind-ineffable-intelligence-record-seed-funding-nvidia-google.html)

25. [DeepMind's David Silver just raised $1.1B to build an AI that learns without human data - TechCrunch, 27 April 2026](https://techcrunch.com/2026/04/27/deepminds-david-silver-just-raised-1-1b-to-build-an-ai-that-learns-without-human-data/)

26. [Ineffable Intelligence launches with record-breaking $1.1B Seed round - Tech.eu, 27 April 2026](https://tech.eu/2026/04/27/ineffable-intelligence-launches-with-record-breaking-11b-seed-round/)

27. [Ineffable Intelligence Selects Google Cloud To Power Its Superintelligence Mission - Google Cloud, 16 June 2026](https://www.googlecloudpresscorner.com/2026-06-16-Ineffable-Intelligence-Selects-Google-Cloud-To-Power-Its-Superintelligence-Mission)

28. [Ineffable Intelligence adds six cofounders - Fortune, 7 September 2026](https://fortune.com/2026/09/07/ineffable-intelligence-hires-cofounders-hiring-google-deepmind-instadeep-flying-fish/)

29. [Innateness, AlphaZero, and Artificial Intelligence - Gary Marcus, arXiv, January 2018](https://arxiv.org/abs/1801.05667)

30. [Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path? - Bengio et al., arXiv, February 2025](https://arxiv.org/abs/2502.15657)

31. [The Cake that is Intelligence and Who Gets to Bake it - arXiv, February 2025 (documents LeCun's cake analogy)](https://arxiv.org/abs/2502.03038)

## Portrait credit

AI-generated watercolor interpretation based on a reference photograph.

[Likeness reference: David Silver's official profile photograph. AI-generated watercolor interpretation.](https://davidstarsilver.wordpress.com/)
