A data loader is the code that turns stored data into training batches: it reads, decodes, samples, cuts and stacks examples, ideally while the device computes on the previous batch. A memory-mapped dataset stores the data in one binary file that the operating system maps into the process’s address space: nothing is read until a slice is touched, the operating system caches what is read, and several processes can share one copy.
bench_train.py, measured_loader.csv.