Todos os livros

Profissional

Apps Sobre Coach Entrar Começar a ler

Quantitative Finance · Glossário

O que é Action-value function, Bellman equation?

Também chamado de: action-value function · Bellman equation

Definition 17.2 Machine Learning for Markets · Capítulo 17 — Reinforcement Learning Foundations

The action-value function of a policy is Qπ(x,u)=Eπ[Gt∣xt=x,ut=u]Q^\pi(x, u) = \E_\pi[G_t\mid x_t = x, u_t = u], the expected cumulative reward of taking uu in xx and following π\pi afterwards; the value function of Book 4 is Vπ(x)=∑uπ(u∣x)Qπ(x,u)V^\pi(x) = \sum_u\pi(u\mid x)Q^\pi(x, u). The Bellman equation is the recursion Qπ(x,u)=E[gt+1+γVπ(xt+1)∣x,u]Q^\pi(x, u) = \E[g_{t+1} + \gamma V^\pi(x_{t+1})\mid x, u]; for the optimal policy, Q∗(x,u)=E[gt+1+γmax⁡u′Q∗(xt+1,u′)∣x,u]Q^*(x, u) = \E[g_{t+1} + \gamma\max_{u'}Q^*(x_{t+1}, u')\mid x, u].

Ler no capítulo →