The action-value function of a policy is , the expected cumulative reward of taking in and following afterwards; the value function of Book 4 is . The Bellman equation is the recursion ; for the optimal policy, .
Quantitative Finance · المسرد
ما معنى Action-value function, Bellman equation؟
يُعرف أيضًا باسم: action-value function · Bellman equation