A deep Q-network (DQN) is Q-learning (chapter 17) with the action-value function represented by a neural network, trained by stochastic gradient steps on the temporal-difference error (Mnih and co-authors, 2015). Two stabilisers make it work: experience replay stores past transitions and trains on random batches of them, breaking the correlation of consecutive steps; a target network is a copy of the network, updated only every few hundred steps, used to compute the learning targets so that they do not move with every update.
Quantitative Finance · Glossário
O que é Deep Q-network, experience replay, target network?
Também chamado de: deep Q-network · experience replay · target network