Temporal-difference learning updates an estimate of a value towards a target built from the next reward and the current estimate of the next state’s value, with , without waiting for the end of the episode. Q-learning applies it to action values with the target , which learns the optimal policy’s values while acting otherwise (Watkins and Dayan, 1992); SARSA uses the action actually taken next, and learns the values of the policy it follows.
Quantitative Finance · Begrippenlijst
Wat is Temporal-difference learning, Q-learning?
Ook bekend als: temporal-difference learning · Q-learning