← All topics

Reinforcement learning

What can AI learn through experience?

Start with David Silver's positions on reinforcement learning, self-play, and when other methods are more appropriate. Coverage will grow as profiles are added.

These are editorial summaries of positions recorded in the profiles. The source year describes the cited statement; the verification date describes the profile review. Inclusion does not imply agreement. Coverage is limited to the profiles available here.

David Silver

Ineffable Intelligence

Profile verified 2026-09-19

Reinforcement learning as the path to general intelligence

Holds that reward maximisation by an agent acting in its environment is sufficient to account for intelligence, and that reinforcement learning will sit at the core of any general system.

Reward is enough, abstract, 2021 ↗

Quote and context
“powerful reinforcement learning agents could constitute a solution to artificial general intelligence”

How the view has changed. Said in 2020 that reinforcement learning would be at the core of any human-level system; formalised it as a hypothesis in 2021; by 2026 he was building a company on it.

When not to use reinforcement learning

Says reinforcement learning is not the right frame for every problem, and cites his own advice that AlphaFold be treated as supervised learning.

TalkRL interview at the Reinforcement Learning Conference, August 2024, 2024 ↗

Quote and context
“not all problems are best suited to RL. You really have to find the problems which are natively better understood in a different way.”

Self-play and creativity

Argues that a system learning by trial and error can discover things no human knew, and that self-play is the essence of machine creativity.

Lex Fridman Podcast, 2020 ↗

Quote and context
“creativity means discovering something which wasn't known before, something unexpected”
Read the profile and counterpoints →

Gary Marcus

New York University

Profile verified 2026-09-22

Innate structure and reinforcement learning

Argues that systems presented as learning from scratch, such as AlphaGo Zero, depend on structure their designers built in, and that AI should study which innate machinery to include rather than minimize it.

Innateness, AlphaZero, and Artificial Intelligence, arXiv, 2018 ↗

Quote and context
“I close by arguing that artificial intelligence needs greater attention to innateness, and I point to some proposals about what that innateness might look like.”
Read the profile and counterpoints →

Ilya Sutskever

Safe Superintelligence Inc.

Profile verified 2026-09-22

Generalization and continual learning

He considers poor generalization the central weakness of current models, including those trained with reinforcement learning on narrow evaluations, and wants systems that learn on the job the way a person does, guided by something like a human value function.

Dwarkesh Podcast, November 25, 2025, 2025 ↗

Quote and context
“these models somehow just generalize dramatically worse than people. It's super obvious. That seems like a very fundamental thing.”

How the view has changed. He described the goal as a "superintelligent 15-year-old" that is deployed and then learns each job, rather than a finished system that already knows everything.

Read the profile and counterpoints →