D10 · Learning by Trying
How Good Is This Square?
Read a map of values and turn it into a policy by always stepping to the most valuable neighbor.
Every square now carries a number: the return the agent can expect from there if it plays well. That map is called the value function.