सभी किताबें

पेशेवर

ऐप्स परिचय Coach लॉग इन पढ़ना शुरू करें

Quantitative Finance · शब्दावली

Gradient accumulation, data parallelism, model parallelism, all-reduce क्या है?

अन्य नाम: gradient accumulation · data parallelism · model parallelism · all-reduce

Definition 23.4 Machine Learning for Markets · अध्याय 23 — Training Infrastructure

Gradient accumulation sums the gradients of several micro-batches before one optimiser step, which equals one step on their union. Data parallelism runs a copy of the model on each device, each on its own share of the batch, and averages their gradients before every step (Li and co-authors, 2020); model parallelism splits one model’s layers or matrices across devices when it does not fit on one (Shoeybi and co-authors, 2019). An all-reduce is the collective operation that leaves every device with the sum (or average) of all devices’ arrays.

अध्याय में पढ़ें →