#86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning
0:001:48:28
Speed
Listening elsewhere? Jump this view to your spot.
About this episode
<p>David Silver leads the reinforcement learning research group at DeepMind and was lead researcher on AlphaGo, AlphaZero and co-lead on AlphaStar, and MuZero and lot of important work in reinforcement learning.</p>
<p>Support this podcast by signing up with these sponsors:<br />
– MasterClass: <a href="https://masterclass.com/lex">https://masterclass.com/lex</a><br />
– Cash App – use code “LexPodcast” and download:<br />
– Cash App (App Store): <a href="https://apple.co/2sPrUHe">https://apple.co/2sPrUHe</a><br />
– Cash App (Google Play): <a href="https://bit.ly/2MlvP5w">https://bit.ly/2MlvP5w</a></p>
<p>EPISODE LINKS:<br />
Reinforcement learning (book): https://amzn.to/2Jwp5zG</p>
<p><span style="font-weight: 400;">This conversation is part of the Artificial Intelligence podcast.</span> If you would like to get more information about this podcast go to <a href="https://lexfridman.com/ai">https://lexfridman.com/ai</a> or connect with @lexfridman on <a href="https://twitter.com/lexfridman">Twitter</a>, <a href="https://www.linkedin.com/in/lexfridman/">LinkedIn</a>, <a href="https://www.facebook.com/lexfridman">Facebook</a>, <a href="https://medium.com/@lexfridman">Medium</a>, or <a href="https://www.youtube.com/lexfridman">YouTube</a> where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on <a href="https://podcasts.apple.com/us/podcast/artificial-intelligence/id1434243584">Apple Podcasts</a>, follow on <a href="https://open.spotify.com/show/2MAi0BvDc6GTFvKFPXnkCL">Spotify</a>, or support it on <a href="https://www.patreon.com/lexfridman">Patreon</a>.</p>
<p>Here’s the outline of the episode. On some podcast players you should be able to click the timestamp to jump to that time.</p>
<p>OUTLINE:<br />
00:00 – Introduction<br />
04:09 – First program<br />
11:11 – AlphaGo<br />
21:42 – Rule of the game of Go<br />
25:37 – Reinforcement learning: personal journey<br />
30:15 – What is reinforcement learning?<br />
43:51 – AlphaGo (continued)<br />
53:40 – Supervised learning and self play in AlphaGo<br />
1:06:12 – Lee Sedol retirement from Go play<br />
1:08:57 – Garry Kasparov<br />
1:14:10 – Alpha Zero and self play<br />
1:31:29 – Creativity in AlphaZero<br />
1:35:21 – AlphaZero applications<br />
1:37:59 – Reward functions<br />
1:40:51 – Meaning of life</p>