सभी किताबें

पेशेवर

ऐप्स परिचय Coach लॉग इन पढ़ना शुरू करें

Quantitative Finance · शब्दावली

Temporal-difference learning, Q-learning क्या है?

अन्य नाम: temporal-difference learning · Q-learning

Definition 17.3 Machine Learning for Markets · अध्याय 17 — Reinforcement Learning Foundations

Temporal-difference learning updates an estimate of a value towards a target built from the next reward and the current estimate of the next state’s value, V(xt)←V(xt)+ηαδtV(x_t)\leftarrow V(x_t) + \eta_\alpha\delta_t with δt=gt+1+γV(xt+1)−V(xt)\delta_t = g_{t+1} + \gamma V(x_{t+1}) - V(x_t), without waiting for the end of the episode. Q-learning applies it to action values with the target gt+1+γmax⁡u′Q(xt+1,u′)g_{t+1} + \gamma\max_{u'}Q(x_{t+1}, u'), which learns the optimal policy’s values while acting otherwise (Watkins and Dayan, 1992); SARSA uses the action actually taken next, and learns the values of the policy it follows.

अध्याय में पढ़ें →