Alle boeken

Professioneel

Apps Over Coach Inloggen Begin met lezen

Quantitative Finance · Begrippenlijst

Wat is Adam, weight decay?

Ook bekend als: Adam · weight decay

Definition 7.3 Machine Learning for Markets · Hoofdstuk 7 — Neural Networks for Noisy Tabular Data

Adam is stochastic gradient descent (Book 4, chapter 24) with a per-parameter step: with gradient gkg_k at step kk, it keeps averages mk=β1mk−1+(1−β1)gkm_k = \beta_1m_{k-1} + (1-\beta_1)g_k and vk=β2vk−1+(1−β2)gk2v_k = \beta_2v_{k-1} + (1-\beta_2)g_k^2 and moves θ\theta by −η m^k/(v^k+ϵ)-\eta\,\hat m_k/(\sqrt{\hat v_k} + \epsilon), where hats undo the averages’ bias towards zero (Kingma and Ba). Weight decay shrinks every weight by ηλθ\eta\lambda\theta at each step, apart from the gradient step (AdamW, Loshchilov and Hutter); for plain gradient descent it is a ridge penalty.

Lees in het hoofdstuk →