Todos los libros

Profesional

Apps Acerca de Coach Iniciar sesión Empezar a leer

Quantitative Finance · Glosario

¿Qué es Gradient accumulation, data parallelism, model parallelism, all-reduce?

También llamado: gradient accumulation · data parallelism · model parallelism · all-reduce

Definition 23.4 Machine Learning for Markets · Capítulo 23 — Training Infrastructure

Gradient accumulation sums the gradients of several micro-batches before one optimiser step, which equals one step on their union. Data parallelism runs a copy of the model on each device, each on its own share of the batch, and averages their gradients before every step (Li and co-authors, 2020); model parallelism splits one model’s layers or matrices across devices when it does not fit on one (Shoeybi and co-authors, 2019). An all-reduce is the collective operation that leaves every device with the sum (or average) of all devices’ arrays.

Leer en el capítulo →