D10 · Learning by Trying
The Same Loop, Pointed at Words
Map the last stage of training a chat assistant onto agent, action, reward and policy.
Here is the loop you know, with new labels. The agent is a language model. The state is the conversation so far. An action is the next token. A whole reply is one episode.