Temporal-difference learning updates an estimate of a value towards a target built from the next reward and the current estimate of the next state’s value, with , without waiting for the end of the episode. Q-learning applies it to action values with the target , which learns the optimal policy’s values while acting otherwise (Watkins and Dayan, 1992); SARSA uses the action actually taken next, and learns the values of the policy it follows.
Quantitative Finance · शब्दावली
Temporal-difference learning, Q-learning क्या है?
अन्य नाम: temporal-difference learning · Q-learning