Stage 2 · Intermediate · D10
Learning by Trying
AlphaGo Zero never saw a human game of Go. It beat the version that had, 100 games to 0.
15 lessons · 138 minDemanding
About this chapter
Some things have no answer key, only a score that arrives late. This chapter builds learning from that score, one small picture at a time. You will choose between four cafes, watch an agent learn a grid square by square, and poke it until it finds another way. Then the same loop grows up: networks that play Atari from pixels, programs that beat world champions by playing themselves, robots trained in simulation, and the last stage of training a chat assistant.
What you will be able to do
- 18 min
Nobody Gives It the Answer
Name the four parts of learning by trying: the agent, the world, the action and the reward.
- 28 min
Stay or Try Something New
Estimate how good each choice is from its average so far, and say why a few tries can mislead.
- 39 min
Mostly Greedy, Sometimes Curious
Set how often an agent explores, and read what exploring costs and what it buys.
- 49 min
Where You Stand Changes the Choice
Describe a problem as states, actions, rewards and moves, and check that the state holds everything needed to choose.
- 59 min
A Dollar Today
Add a string of rewards into one return, and shrink the far ones with a discount.
- 69 min
How Good Is This Square?
Read a map of values and turn it into a policy by always stepping to the most valuable neighbor.
- 711 min
Learning One Step at a Time
Update a Q-value from a single move, and watch a table of them learn a route with no map.
- 89 min
Curious Early, Careful Late
Pick an exploration schedule, and spot an agent that settled for less because it stopped looking too soon.
- 911 min
Too Many Squares for a Table
Explain how a network replaced the Q table, and what DQN needed to keep that network from falling apart.
- 1010 min
Push Up What Worked
Explain in plain words how REINFORCE makes the moves of a good episode more likely.
- 1110 min
A Coach Beside the Player
Say what a critic adds to a policy, and why PPO limits how far each update may move it.
- 128 min
It Learned the Score, Not the Task
Spot an agent that collects the reward it was given instead of doing the job it was meant for.
- 1310 min
Playing Against Yourself
Explain how self-play makes its own opponents, and how search improves a move before it is played.
- 149 min
From Games to Robots
Explain why robots learn in simulation first, and how a deliberately messy simulator helps them survive the real world.
- 158 min
The Same Loop, Pointed at Words
Map the last stage of training a chat assistant onto agent, action, reward and policy.