Gradient accumulation sums the gradients of several micro-batches before one optimiser step, which equals one step on their union. Data parallelism runs a copy of the model on each device, each on its own share of the batch, and averages their gradients before every step (Li and co-authors, 2020); model parallelism splits one model’s layers or matrices across devices when it does not fit on one (Shoeybi and co-authors, 2019). An all-reduce is the collective operation that leaves every device with the sum (or average) of all devices’ arrays.
Quantitative Finance · Glossário
O que é Gradient accumulation, data parallelism, model parallelism, all-reduce?
Também chamado de: gradient accumulation · data parallelism · model parallelism · all-reduce