The action-value function of a policy is , the expected cumulative reward of taking in and following afterwards; the value function of Book 4 is . The Bellman equation is the recursion ; for the optimal policy, .
Quantitative Finance · Begrippenlijst
Wat is Action-value function, Bellman equation?
Ook bekend als: action-value function · Bellman equation