Adam is stochastic gradient descent (Book 4, chapter 24) with a per-parameter step: with gradient at step , it keeps averages and and moves by , where hats undo the averages’ bias towards zero (Kingma and Ba). Weight decay shrinks every weight by at each step, apart from the gradient step (AdamW, Loshchilov and Hutter); for plain gradient descent it is a ridge penalty.
Quantitative Finance · Glosarium
Apa itu Adam, weight decay?
Dikenal juga sebagai: Adam · weight decay