AlphaGo Zero learns Go from self-play alone

DeepMind announced AlphaGo Zero, trained only by playing itself from random play, with no human games and no hand-crafted features. After three days it beat the version that had defeated Lee Sedol, 100 games to nil. Nature published the work the next day.

Why it mattered A general learning method passed centuries of accumulated human Go strategy without reading a single human game, weakening the assumption that human data was needed for superhuman play.

The AlphaGo that beat Lee Sedol in Seoul had learned in two steps. It first studied about 30 million moves from games between strong human players, then improved by playing copies of itself. Human games were where it started.

AlphaGo Zero removed them. It began with the rules of the game and random play, and it used one neural network where the earlier system had used two, so that proposing a move and judging a position became a single act. After three days of self-play it played the version that had won in Seoul and took 100 games out of 100. After 40 days it beat AlphaGo Master, the version that had won 60 online games against professionals and beaten Ke Jie in Wuzhen in May, by 89 games to 11. It ran on four processors, against the 48 used by the machine in Korea.

Left to itself the program worked out opening sequences that human players had settled on over centuries, then set some of them aside for lines professionals did not play. DeepMind announced the result on 18 October 2017. Nature published the paper, “Mastering the game of Go without human knowledge,” the following day.

The claim under test was about data rather than about Go. Strong game-playing programs had until then been built on a record of human decisions to imitate, and the assumption that such a record was necessary went with them. Seven weeks later DeepMind posted a paper on AlphaZero, which learned chess and shogi the same way, from the rules alone.

Corrections

  • 2026-09-01 - The Nature paper (nature24270) was listed as a secondary source. DeepMind’s own researchers wrote it, so it is corrected to primary.