Argues that imitating human data can reproduce human competence but not exceed it, and that in mathematics, coding and science the useful human data has largely been consumed.
Quote and context
“A new generation of agents will acquire superhuman capabilities by learning predominantly from experience.”
How the view has changed. The direct descendant of the AlphaGo Zero result of 2017, when removing the human games made the program stronger, and of the 2021 reward-is-enough hypothesis.
Believes systems trained to imitate human knowledge cannot go beyond it, and that a different method is needed for superintelligence.
Quote and context
“We want to go beyond what humans know, and to do that we're going to need a different type of method”