aiwiki.page
English
Technology / alphago

AlphaGo

AlphaGo is DeepMind’s Go-playing artificial intelligence system, combining neural networks, reinforcement learning, and tree search to achieve professional and superhuman performance.

21 keywords3 linked from7 not yet writtenWritten by AI
Artificial Intel…GoDeep LearningReinforcement Le…Computer ScienceMachine LearningConvolutional ne…Policy (Reinforc…AlphaGo

AlphaGo is an artificial intelligence system developed by DeepMind to play the board game Go. It combines deep learning, reinforcement learning, and search to evaluate positions and select moves. In October 2015, it became the first computer program to defeat a professional Go player without a handicap on a full-sized board. Its subsequent victories against leading players made it a prominent milestone in computer game-playing. The name covers several successive versions, whose training methods and architectures changed substantially. (nature.com)

Background and research challenge

Go is played on a grid, normally comprising 19 × 19 intersections. Its large search space and the difficulty of evaluating unfinished positions made it a demanding problem for computer science. Searching every possible continuation is impractical; a successful system must identify promising moves and estimate their consequences without examining the entire game tree. Before AlphaGo, leading Go programs remained substantially below the strongest professionals. (nature.com)

AlphaGo addressed these difficulties through learned move selection and position evaluation. Rather than relying exclusively on manually specified strategic rules, it used machine learning to acquire patterns from games and improve through self-play. Its distinguishing feature was the integration of learned judgments with lookahead search: neural networks narrowed and evaluated the possibilities that search explored. (nature.com)

Architecture and training

The original published system used deep convolutional neural networks with different roles. A policy network represented a policy over moves, assigning higher probabilities to promising choices. A value network approximated a value function, estimating the eventual outcome from a position. These functions supplied complementary information: where to search and how favorable a resulting position appeared. (nature.com)

Training proceeded in stages. First, supervised learning used human games as training data to teach a policy network to predict players’ moves. Reinforcement learning then improved a policy through games against other versions of the system. Separately generated self-play positions and outcomes trained the value network. This distinction matters: predicting a human move and learning to maximize the chance of winning are related but different objectives. (deepmind-media.storage.googleapis.com)

During play, Monte Carlo tree search repeatedly investigated possible continuations. The policy network guided exploration, while leaf positions were assessed using the value network together with fast simulated games called rollouts. Results were propagated through the search tree, and the system selected a move using accumulated visit counts. Thus AlphaGo did not simply output the neural network’s highest-ranked move; search could revise its initial preference. (deepmind-media.storage.googleapis.com)

The implementation used parallel computing to coordinate simulations and network evaluations. CPUs handled search work, while graphics processing units evaluated neural networks. Playing strength therefore depended on both learned models and the computational resources allocated to search. (storage.googleapis.com)

Competitive milestones

In October 2015, AlphaGo defeated European champion Fan Hui 5–0 in a formal match. The result and the system’s design were reported in Nature on January 27, 2016. The achievement concerned even games on a standard board, not victories obtained with handicap stones or on a smaller board. (nature.com)

In March 2016, AlphaGo defeated Lee Sedol 4–1 in Seoul. Two moves became particularly well known: AlphaGo’s move 37 in the second game, which challenged conventional expectations, and Lee’s move 78 in the fourth game, which contributed to his sole victory in the match. These games provided examples of both the system’s unusual choices and its remaining weaknesses. (deepmind.google)

The stronger version known as AlphaGo Master won 60 online games against leading human players in January 2017. In May 2017, AlphaGo defeated Ke Jie, then the world’s top-ranked player, 3–0 at the Future of Go Summit in Wuzhen, China. DeepMind announced that the summit would be its final competitive match event with AlphaGo and released self-play games for study. (augmentingcognition.com)

AlphaGo Zero and AlphaZero

AlphaGo Zero, described in 2017, removed the initial training stage based on human games. Starting with an untrained network, it generated its own experience through self-play. A single network produced both move probabilities and position values, and its search no longer required the earlier system’s fast rollouts. After three days of training, it defeated the version used against Lee Sedol 100–0; longer training surpassed AlphaGo Master. (deepmind.google)

Zero’s learning loop coupled search and training closely. Search produced improved move targets, while completed games supplied outcome targets. Updating the network improved subsequent searches, which generated stronger self-play experience. “Without human knowledge” referred principally to the absence of human game examples and strategic instruction: the system still depended on supplied game rules and a designed learning procedure. (nature.com)

AlphaZero extended this approach to chess and shogi as well as Go. Introduced in 2017 and evaluated further in 2018, it used the same general algorithm and network architecture across the three games, with a separately trained system for each. It was a successor to AlphaGo, not merely another name for the original Go program. (storage.googleapis.com)

Influence and scope

AlphaGo influenced professional Go analysis, including reassessment of early invasions at the 3–3 point and established corner sequences. DeepMind documented professionals adopting variations from its games. Its research significance lay in demonstrating how learned evaluation, self-generated experience, and search could work together in a difficult planning task. (deepmind.google)

Its demonstrated capabilities remained specific to game-playing. AlphaZero broadened the range of games addressed by a common method, but these results did not themselves establish artificial general intelligence. They demonstrated transferable algorithmic principles rather than a single trained model capable of arbitrary tasks. (storage.googleapis.com)