Skip to content

Stage 2 · Intermediate · D10

Learning by Trying

AlphaGo Zero never saw a human game of Go. It beat the version that had, 100 games to 0.

15 lessons · 138 minDemanding

About this chapter

Some things have no answer key, only a score that arrives late. This chapter builds learning from that score, one small picture at a time. You will choose between four cafes, watch an agent learn a grid square by square, and poke it until it finds another way. Then the same loop grows up: networks that play Atari from pixels, programs that beat world champions by playing themselves, robots trained in simulation, and the last stage of training a chat assistant.

What you will be able to do

  1. 1

    Nobody Gives It the Answer

    Name the four parts of learning by trying: the agent, the world, the action and the reward.

    8 min
  2. 2

    Stay or Try Something New

    Estimate how good each choice is from its average so far, and say why a few tries can mislead.

    8 min
  3. 3

    Mostly Greedy, Sometimes Curious

    Set how often an agent explores, and read what exploring costs and what it buys.

    9 min
  4. 4

    Where You Stand Changes the Choice

    Describe a problem as states, actions, rewards and moves, and check that the state holds everything needed to choose.

    9 min
  5. 5

    A Dollar Today

    Add a string of rewards into one return, and shrink the far ones with a discount.

    9 min
  6. 6

    How Good Is This Square?

    Read a map of values and turn it into a policy by always stepping to the most valuable neighbor.

    9 min
  7. 7

    Learning One Step at a Time

    Update a Q-value from a single move, and watch a table of them learn a route with no map.

    11 min
  8. 8

    Curious Early, Careful Late

    Pick an exploration schedule, and spot an agent that settled for less because it stopped looking too soon.

    9 min
  9. 9

    Too Many Squares for a Table

    Explain how a network replaced the Q table, and what DQN needed to keep that network from falling apart.

    11 min
  10. 10

    Push Up What Worked

    Explain in plain words how REINFORCE makes the moves of a good episode more likely.

    10 min
  11. 11

    A Coach Beside the Player

    Say what a critic adds to a policy, and why PPO limits how far each update may move it.

    10 min
  12. 12

    It Learned the Score, Not the Task

    Spot an agent that collects the reward it was given instead of doing the job it was meant for.

    8 min
  13. 13

    Playing Against Yourself

    Explain how self-play makes its own opponents, and how search improves a move before it is played.

    10 min
  14. 14

    From Games to Robots

    Explain why robots learn in simulation first, and how a deliberately messy simulator helps them survive the real world.

    9 min
  15. 15

    The Same Loop, Pointed at Words

    Map the last stage of training a chat assistant onto agent, action, reward and policy.

    8 min

Before you start

Keep going

All chapters