Reinforcement learning learns how to act from the consequences of actions. A Markov decision process (MDP) is a set of states , actions , transition probabilities and a reward received after each action; the cumulative reward from time is , with a discount ( on the finite horizons of this chapter). A policy gives the probability of each action in each state.
Quantitative Finance · Glossaire
Qu'est-ce que « Reinforcement learning, Markov decision process, policy, reward, cumulative reward » ?
Aussi appelé : reinforcement learning · Markov decision process · reward · cumulative reward · policy