Semua buku

Profesional

Aplikasi Tentang Pelatih Masuk Mulai membaca

Quantitative Finance · Glosarium

Apa itu Reinforcement learning, Markov decision process, policy, reward, cumulative reward?

Dikenal juga sebagai: reinforcement learning · Markov decision process · reward · cumulative reward · policy

Definition 17.1 Machine Learning for Markets · Bab 17 — Reinforcement Learning Foundations

Reinforcement learning learns how to act from the consequences of actions. A Markov decision process (MDP) is a set of states xx, actions uu, transition probabilities P(xt+1∣xt,ut)\P(x_{t+1}\mid x_t, u_t) and a reward gt+1g_{t+1} received after each action; the cumulative reward from time tt is Gt=∑k≥0γkgt+k+1G_t = \sum_{k\ge0}\gamma^kg_{t+k+1}, with a discount γ∈(0,1]\gamma\in(0,1] (γ=1\gamma = 1 on the finite horizons of this chapter). A policy π(u∣x)\pi(u\mid x) gives the probability of each action in each state.

Baca dalam konteks →