Holds that reward maximisation by an agent acting in its environment is sufficient to account for intelligence, and that reinforcement learning will sit at the core of any general system.
Quote and context
“powerful reinforcement learning agents could constitute a solution to artificial general intelligence”
How the view has changed. Said in 2020 that reinforcement learning would be at the core of any human-level system; formalised it as a hypothesis in 2021; by 2026 he was building a company on it.
Says reinforcement learning is not the right frame for every problem, and cites his own advice that AlphaFold be treated as supervised learning.
Quote and context
“not all problems are best suited to RL. You really have to find the problems which are natively better understood in a different way.”
Argues that a system learning by trial and error can discover things no human knew, and that self-play is the essence of machine creativity.
Quote and context
“creativity means discovering something which wasn't known before, something unexpected”