Todos los libros

Profesional

Apps Acerca de Coach Iniciar sesión Empezar a leer

Quantitative Finance · Glosario

¿Qué es Adam, weight decay?

También llamado: Adam · weight decay

Definition 7.3 Machine Learning for Markets · Capítulo 7 — Neural Networks for Noisy Tabular Data

Adam is stochastic gradient descent (Book 4, chapter 24) with a per-parameter step: with gradient gkg_k at step kk, it keeps averages mk=β1mk−1+(1−β1)gkm_k = \beta_1m_{k-1} + (1-\beta_1)g_k and vk=β2vk−1+(1−β2)gk2v_k = \beta_2v_{k-1} + (1-\beta_2)g_k^2 and moves θ\theta by −η m^k/(v^k+ϵ)-\eta\,\hat m_k/(\sqrt{\hat v_k} + \epsilon), where hats undo the averages’ bias towards zero (Kingma and Ba). Weight decay shrinks every weight by ηλθ\eta\lambda\theta at each step, apart from the gradient step (AdamW, Loshchilov and Hutter); for plain gradient descent it is a ridge penalty.

Leer en el capítulo →