Reinforcement learning learns how to act from the consequences of actions. A Markov decision process (MDP) is a set of states , actions , transition probabilities and a reward received after each action; the cumulative reward from time is , with a discount ( on the finite horizons of this chapter). A policy gives the probability of each action in each state.
Quantitative Finance · المسرد
ما معنى Reinforcement learning, Markov decision process, policy, reward, cumulative reward؟
يُعرف أيضًا باسم: reinforcement learning · Markov decision process · reward · cumulative reward · policy