Todos os livros

Profissional

Apps Sobre Coach Entrar Começar a ler

Quantitative Finance · Glossário

O que é Exploration–exploitation trade-off, multi-armed bandit, contextual bandit?

Também chamado de: exploration--exploitation trade-off · multi-armed bandit · contextual bandit

Definition 17.5 Machine Learning for Markets · Capítulo 17 — Reinforcement Learning Foundations

The exploration–exploitation trade-off is the choice between the action that looks best now and an action that teaches more. A multi-armed bandit is the MDP with one state: each round an arm is chosen and its reward drawn; the loss against always playing the best arm is the regret. A contextual bandit observes a context (the order, the market) before choosing, and learns a reward model per arm; its actions do not change future contexts.

Cumulative regret of four exploration rules choosing among five execution algorithms, mean of 20 runs. Data: ml_rl.bandits.
Figure 17.2. Cumulative regret of four exploration rules choosing among five execution algorithms, mean of 20 runs. Data: ml_rl.bandits.
Ler no capítulo →