D10 · Learning by Trying
Learning One Step at a Time
Update a Q-value from a single move, and watch a table of them learn a route with no map.
This agent has no map. It starts knowing nothing, every number zero, and learns only from moves it actually makes.