---
title: "Training Infrastructure"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 23
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/23-training-infrastructure
---

# Chapter 23 — Training Infrastructure

A team’s order-book model trains for a day, and the team prices a second [accelerator](#def-ml-training-infrastructure-accelerator). Before it buys one it should know where the time goes. On the chapter’s small version of the problem, measured on a laptop, a third of every training step is spent assembling the batch, not computing on it; changing the storage format barely matters, and gathering the batch in one indexed read instead of window by window cuts the data’s share to a fifth. This chapter is about the machinery under the models of Parts I to IV: [accelerators](#def-ml-training-infrastructure-accelerator) and where training time goes, [data loaders](#def-ml-training-infrastructure-loader) for tick data that never let a window cross a session, lower-precision arithmetic, training on several devices, and checkpoints that let a killed run resume exactly. Everything runs on one CPU thread, with no worker processes; the multi-device parts are simulated in one process, which is enough to check their arithmetic, and the timings are the author’s laptop’s, labelled as such.

## 23.1 Accelerators and where the time goes

**Definition 23.1 (Accelerator).**

An *accelerator* is a processor built for the dense linear algebra of [neural networks](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-mlp) (a graphics processor or a dedicated tensor processor): thousands of arithmetic units, high-bandwidth memory beside them, and low-precision matrix units that do most of the work.

**As of September 2026 — A data-centre accelerator.**

NVIDIA’s specification page for the H100 lists 80 GB of memory with 3.35 TB/s of bandwidth for the SXM version (94 GB and 3.9 TB/s for the NVL version), and 1 979 teraflops of [bfloat16](#def-ml-training-infrastructure-mixed) tensor-core arithmetic for the SXM version, a figure the page marks as with sparsity.

Figures like these explain the usual bottleneck. An [accelerator](#def-ml-training-infrastructure-accelerator) can consume data far faster than a storage system, a decompressor or a Python loop can feed it, so the first question about a slow training job is what fraction of each step the device waits. The step divides into loading (reading the data from storage), batching (cutting windows and stacking them) and computing (forward pass, backward pass, optimiser update). The chapter measures the three on its own job: eight ten-minute `firm.tape` sessions of the book’s synthetic market, each a table of 1 140 half-second rows and 41 columns (the mid price and the ten-level book of chapter 8), and a two-layer network that reads the last twenty rows and predicts the mid’s change over the next ten.

## 23.2 Data loaders for tick data

**Definition 23.2 (Data loader, memory-mapped dataset).**

A *data loader* is the code that turns stored data into training batches: it reads, decodes, samples, cuts and stacks examples, ideally while the device computes on the previous batch. A *memory-mapped dataset* stores the data in one binary file that the operating system maps into the process’s address space: nothing is read until a slice is touched, the operating system caches what is read, and several processes can share one copy.

Tick data impose one rule on the loader: a training window must not cross a session boundary, or a window will contain the close of one day and the open of the next as if they were consecutive (and, with overlapping sessions for several instruments, leak across them). The sampler of [Listing 23.1](#lst-ml-train-sampler) draws (session, start) pairs uniformly among the windows that fit, and its random state is part of what a checkpoint saves.

The sessions are stored three ways: one CSV file per session (2.21 MB in all), one NumPy file per session (1.50 MB), and one memory-mapped float32 file with a small JSON index (1.50 MB). [Figure 23.1](#fig-ml-train-time) shows the measured cost of a training step with each, loading being charged once per [epoch](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop) of 100 steps. Parsing the CSV files takes 12.5 ms per [epoch](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop) (176 MB/s), reading the NumPy files 0.83 ms (1.8 GB/s), and mapping the single file 0.39 ms whatever its size. Per step these are tiny against the half millisecond of cutting 256 windows one by one in Python and the 1.1 to 1.2 ms of computing: the data’s share of the step is 0.34, 0.31 and 0.39. Gathering the whole batch from the mapped file in one indexed read ([Listing 23.2](#lst-ml-train-gather)) takes 0.27 ms and brings the share to 0.20.

![Where a training step’s time goes, by storage format and batch assembly: medians of five epochs of 100 steps of 256 windows. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, one thread, no isolated cores; other jobs may have run. Data: bench_train.py, measured_loader.csv.](https://one-course.com/images/onecourse/chapters/quant-12/ml-training-infrastructure/fig-1aad67d61b41.svg)

***Figure 23.1.** Where a training step’s time goes, by storage format and batch assembly: medians of five [epochs](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop) of 100 steps of 256 windows. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, one thread, no isolated cores; other jobs may have run. Data: `bench_train.py`, `measured_loader.csv`.*

At this scale the lesson is about code, not storage: the Python loop over windows costs more than any format. At the scale of a year of order-book data the formats diverge. A year of the same half-second table for one instrument, 250 sessions of six and a half hours, is about 1.9 GB of float32 and 2.8 GB of text: at the measured rates, parsing the text costs about 16 seconds per [epoch](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop), reading the binary files about one second, and mapping costs nothing until the windows are read, from the page cache after the first [epoch](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop). And beyond the memory of the machine only a memory-mapped or streamed format works at all.

## 23.3 Mixed precision

**Definition 23.3 (Mixed-precision training, bfloat16).**

*Mixed-precision training* runs the expensive operations (matrix products, convolutions) in a 16-bit format while keeping the weights, the loss and the optimiser’s state in 32 bits (Micikevicius and co-authors, 2018). *bfloat16* is a 16-bit floating-point format with float32’s eight exponent bits and seven fraction bits: the same range as float32, and a machine epsilon (Book 4, chapter 25) of $2^{-7}\approx0.0078$ instead of $2^{-23}$, so no loss scaling is needed against underflow (Kalamkar and co-authors, 2019).

Under the CPU’s [bfloat16](#def-ml-training-infrastructure-mixed) autocast the chapter’s network trains to the same place ([Figure 23.2](#fig-ml-train-bf16)): over the last 50 of 300 steps the mean loss is 0.0525 against 0.0527 in float32. Step by step the two runs differ: the largest difference between their losses over the run is 0.0132, 22% of that step’s float32 loss, because rounding changes each update slightly and the trajectories drift apart. Reproducing a run bit for bit across precisions is not to be expected; equivalence is judged on the validation metric, over the run.

![Training loss of the same network, same data order and same initialisation, in float32 and under bfloat16 autocast. Data: ml_train.run (fig_train.py).](https://one-course.com/images/onecourse/chapters/quant-12/ml-training-infrastructure/fig-37784fee8b98.svg)

***Figure 23.2.** Training loss of the same network, same data order and same initialisation, in float32 and under [bfloat16](#def-ml-training-infrastructure-mixed) autocast. Data: `ml_train.run` (`fig_train.py`).*

## 23.4 Distributed training

**Definition 23.4 (Gradient accumulation, data parallelism, model parallelism, all-reduce).**

*Gradient accumulation* sums the gradients of several micro-batches before one optimiser step, which equals one step on their union. *Data parallelism* runs a copy of the model on each device, each on its own share of the batch, and averages their gradients before every step (Li and co-authors, 2020); *model parallelism* splits one model’s layers or matrices across devices when it does not fit on one (Shoeybi and co-authors, 2019). An *all-reduce* is the collective operation that leaves every device with the sum (or average) of all devices’ arrays.

Both identities are checked in the chapter’s code ([Listing 23.3](#lst-ml-train-loop)). Four accumulated micro-batches of 64 and one batch of 256 of the same windows leave parameters that differ by at most $1.4\times10^{-7}$ after 20 steps, rounding in a different order of summation. The average of two workers’ gradients on the halves of a batch differs from the whole batch’s gradient by at most $1.5\times10^{-8}$, against gradients of size 0.12: an [all-reduce](#def-ml-training-infrastructure-dist) computes exactly the large-batch gradient, and [data parallelism](#def-ml-training-infrastructure-dist) is [gradient accumulation](#def-ml-training-infrastructure-dist) across devices. What changes with a larger effective batch is the optimisation: the [learning rate](https://one-course.com/books/quant/12/en/chapter/5-trees-and-boosting#def-ml-trees-and-boosting-boosting) and its warm-up must be rescaled (Goyal and co-authors, 2017), and for the small, noisy models of trading, larger batches rarely help.

## 23.5 Checkpoints and deterministic resumption

**Definition 23.5 (Training checkpoint).**

A *training checkpoint* is a saved state from which a run can continue as if it had not stopped: the model’s parameters, the optimiser’s state (moment estimates, step counts), the data sampler’s position and random state, the random generators’ states and the step count.

A run of 200 steps and a run killed at step 100, checkpointed ([Listing 23.4](#lst-ml-train-ckpt)), and resumed into newly constructed objects with a different initialisation end with bitwise identical parameters, and identical losses step for step; the checkpoint takes 693 kB. Saving only the model’s weights, the usual shortcut, restarts [Adam](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-adam)’s moment estimates and replays the sampler’s first windows; the resumed run ends with parameters up to 0.039 away from the uninterrupted one’s. For research that must be reproduced (chapter 25) and for jobs on machines that can be pre-empted, the full checkpoint is the difference between resuming and restarting with a different run.

**Method 23.6 (Before buying an accelerator).**

1. Measure load, batch and compute time per step; if data take more than a small share, fix the loader first (binary or memory-mapped storage, batch gathering in one read, prefetching).
2. Try [bfloat16](#def-ml-training-infrastructure-mixed) autocast and judge it on the validation metric over the run.
3. Grow the batch by accumulation before growing the hardware; rescale the [learning rate](https://one-course.com/books/quant/12/en/chapter/5-trees-and-boosting#def-ml-trees-and-boosting-boosting) and check that the model improves.
4. Checkpoint everything that makes the run’s future, and test resumption bitwise.

## 23.6 Tutorial: where the time goes

**Goal.** Store tick sessions three ways, measure a step’s time budget, train in [bfloat16](#def-ml-training-infrastructure-mixed), verify accumulation and [data parallelism](#def-ml-training-infrastructure-dist), and resume a run bitwise. **End state:** Figures [23.1](#fig-ml-train-time) and [23.2](#fig-ml-train-bf16).

1. **A session-aware window sampler with a saveable state.** `class WindowSampler : """Uniform over (session, start) with the window and its target horizon inside one session.""" def __init__(self , lengths, T=20 , horizon=10 , batch=256 , seed=0 ): self .lengths, self .T, self .h, self .batch = np.asarray(lengths), T, horizon, batch valid = np.maximum(self .lengths - T - horizon + 1 , 0 ) self .cum = np.r_[0 , np.cumsum(valid)] self .rng = np.random.default_rng(seed) def next (self ): k = self .rng.integers(0 , self .cum[-1 ], self .batch) s = np.searchsorted(self .cum, k, side=" right " ) - 1 return s, k - self .cum[s] def state (self ): return self .rng.bit_generator.state def set_state (self , st): self .rng.bit_generator.state = st` **Listing 23.1.** Windows that never cross a session boundary. code/firm/mltrain/firm_mltrain.py
2. **One gather for a whole batch.** `def make_batch_vec (shards, sessions, starts, T=20 , horizon=10 , target_col=0 ): """The same batch as make_batch from a memory-mapped store, in one gather: row indices for every window at once.""" rows = (np.asarray(shards.offsets)[sessions] + starts)[:, None ] + np.arange(T + horizon)[None , :] w = np.asarray(shards.mm[rows.ravel()]).reshape(len (starts), T + horizon, -1 ) X = w[:, :T].reshape(len (starts), -1 ) y = w[:, T + horizon - 1 , target_col] - w[:, T - 1 , target_col] return torch.as_tensor(X, dtype=torch.float32), torch.as_tensor(y, dtype=torch.float32)` **Listing 23.2.** A batch from the memory-mapped store in one indexed read. code/firm/mltrain/firm_mltrain.py
3. **Accumulation and [bfloat16](#def-ml-training-infrastructure-mixed) autocast.** `def train (model, shards, sampler, steps, lr=1e-3 , accumulate=1 , bf16=False , T=20 , horizon=10 , target_col=0 , opt=None , start_step=0 , scale=None ): """Adam on mean squared error. With accumulate = k, each optimiser step sums the gradients of k micro-batches (each scaled by 1/k): the same step as one batch k times larger. With bf16, the forward pass runs under CPU bfloat16 autocast (matrix products in bfloat16, the loss and the update in float32).""" torch.set_num_threads(1 ) opt = opt or torch.optim.Adam(model.parameters(), lr=lr) losses = [] for _ in range (start_step, steps): opt.zero_grad() tot = 0.0 for _ in range (accumulate): X, y = make_batch(shards, *sampler.next(), T, horizon, target_col) if scale is not None : X = (X - scale[0 ]) / scale[1 ] with torch.autocast(" cpu " , dtype=torch.bfloat16, enabled=bf16): p = model(X)[:, 0 ] loss = ((p.float() - y) ** 2 ).mean() / accumulate loss.backward() tot += float (loss.detach()) opt.step() losses.append(tot) return losses` **Listing 23.3.** The training loop. code/firm/mltrain/firm_mltrain.py
4. **Checkpoints.** `def save_checkpoint (path, model, opt, sampler, step): torch.save({" model " : model.state_dict(), " opt " : opt.state_dict(), " sampler " : sampler.state(), " step " : step, " torch_rng " : torch.get_rng_state()}, path) def load_checkpoint (path, model, opt, sampler): ck = torch.load(path, weights_only=False ) model.load_state_dict(ck[" model " ]) opt.load_state_dict(ck[" opt " ]) sampler.set_state(ck[" sampler " ]) torch.set_rng_state(ck[" torch_rng " ]) return ck[" step " ]` **Listing 23.4.** Saving and restoring everything a run’s future depends on. code/firm/mltrain/firm_mltrain.py
5. **Run** `ml_train.precision()` , `accumulation()` , `data_parallel()` , `resume()` , `naive_resume()` , `fig_train.py` ; and, once, `bench_train.py` under `nice` .

**What to change next.** Prefetch the next batch in a background thread and measure the overlap; add a second instrument and check that windows never mix instruments.

## 23.7 Build: the training toolkit

**Purpose.** Training runs that are fast where it matters, correct at session boundaries, and exactly resumable.

**Interface.** `write_shards(sessions, root, fmt)`, `Shards(root, fmt)`, `WindowSampler(lengths, T, horizon, batch, seed)` with `state`/`set_state`, `make_batch`, `make_batch_vec`, `train(model, shards, sampler, steps, lr, accumulate, bf16)`, `data_parallel_grads`, `save_checkpoint`, `load_checkpoint`.

**Rules.** Windows never cross sessions; timings are measured, labelled with the machine, and never asserted; everything that determines the rest of a run goes into its checkpoint.

**Acceptance tests.** `code/firm/mltrain/tests/`: the three formats return identical windows; no sampled window crosses a boundary; the gathered batch equals the looped one; accumulation equals a larger batch; the averaged shard gradients equal the full gradient; resumption is bitwise.

**Stretch.** Real two-process [data parallelism](#def-ml-training-infrastructure-dist) with the gloo backend when the machine allows; streaming from compressed shards; prefetching.

Sources and further reading

- P. Micikevicius and co-authors, “Mixed precision training”, arXiv:1710.03740, 2018.
- D. Kalamkar and co-authors, “A study of BFLOAT16 for deep learning training”, arXiv:1905.12322, 2019.
- S. Li and co-authors, “PyTorch distributed: experiences on accelerating data parallel training”, arXiv:2006.15704, 2020.
- P. Goyal and co-authors, “Accurate, large minibatch SGD: training ImageNet in 1 hour”, arXiv:1706.02677, 2017.
- M. Shoeybi and co-authors, “Megatron-LM: training multi-billion parameter language models using model parallelism”, arXiv:1909.08053, 2019.
- NVIDIA, H100 Tensor Core GPU specifications (web page, accessed September 2026).

## 23.8 Exercises

**Exercise 23.1 ★.**

What is the machine epsilon of [bfloat16](#def-ml-training-infrastructure-mixed), float16 and float32? Which of them can represent $10^{-30}$?

**Solution of Exercise 23.1.**

[bfloat16](#def-ml-training-infrastructure-mixed) $2^{-7}\approx7.8\times10^{-3}$, float16 $2^{-10}\approx9.8\times10^{-4}$, float32 $2^{-23}\approx1.2\times10^{-7}$. $10^{-30}$ is representable in [bfloat16](#def-ml-training-infrastructure-mixed) and float32 (smallest normal numbers about $10^{-38}$) but not in float16, whose smallest positive number is about $6\times10^{-8}$: it underflows to zero, which is why float16 training needs loss scaling.

**Exercise 23.2 ★.**

A session has 1 140 rows; windows are 20 rows with a 10-row target horizon. How many windows does it hold? What goes wrong if a sampler draws starts uniformly over the concatenated sessions?

**Solution of Exercise 23.2.**

$1\,140 - 20 - 10 + 1 = 1\,111$ windows. Uniform starts over the concatenation would produce windows (or targets) that begin in one session and end in the next, splicing a close to an open: a price jump the model would learn to predict from nothing.

**Exercise 23.3 ★.**

From [Figure 23.1](#fig-ml-train-time), what speed-up would halving the compute time give with the looped loader, and with the gather?

**Solution of Exercise 23.3.**

With the memory map and the loop, a step is $0.74 + 1.17 = 1.91$ ms; halving compute gives 1.32 ms, a speed-up of 1.44. With the gather, $0.27 + 1.10 = 1.37$ ms becomes 0.82 ms, a speed-up of 1.67. The faster the loader, the more a faster device is worth.

**Exercise 23.4 ★★.**

Why do float32 and [bfloat16](#def-ml-training-infrastructure-mixed) losses differ by up to 22% at some steps while their averages agree?

**Solution of Exercise 23.4.**

Each step’s rounding changes the update slightly; the two trajectories diverge in parameter space and see the same batches with different parameters, so their per-step losses differ by as much as the loss varies from batch to batch. Both converge to models of the same quality; the comparison must be of the metric over the run, not step by step.

**Exercise 23.5 ★★.**

Show that with a mean loss, accumulating $k$ micro-batches each scaled by $1/k$ gives the gradient of the mean over their union. What changes with [batch normalisation](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-dropout)?

**Solution of Exercise 23.5.**

The loss on the union is $\frac1{kb}\sum_{j=1}^k\sum_{i\in B_j}\ell_i = \frac1k\sum_jL_j$ with $L_j$ each micro-batch’s mean, and the gradient is linear: summing the gradients of $L_j/k$ gives it. [Batch normalisation](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-dropout) computes statistics per micro-batch, so the forward pass itself differs and the equivalence breaks; so does anything else that couples examples within a batch.

**Exercise 23.6 ★★.**

*Find the flaw.* “We resume pre-empted jobs from the last saved weights, so our runs are reproducible.”

**Solution of Exercise 23.6.**

Weights alone do not determine the rest of the run: the optimiser’s moments restart, the sampler replays or skips data, and the random generators restart. In the chapter’s experiment the resumed run ended up to 0.039 away from the uninterrupted one. Save and restore every state, and test resumption bitwise.

**Exercise 23.7 ★★★.**

Estimate, from the measured throughputs, the [per-epoch](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop) loading time of a year of the chapter’s table in each format (250 sessions of six and a half hours of half-second rows, 41 float32 columns, text about 1.5 times larger than binary).

**Solution of Exercise 23.7.**

Rows: $250\times6.5\times3\,600/0.5 = 11.7$ million; binary $11.7\times10^6\times41\times4 = 1.92$ GB; text about 2.8 GB. At 176 MB/s the text takes about 16 s per [epoch](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop); the NumPy files at 1.8 GB/s about 1.1 s; the memory map reads only what the batches touch, from the page cache once warm.

**Exercise 23.8 ★★★.**

In a ring [all-reduce](#def-ml-training-infrastructure-dist) over $p$ devices each holding $n$ numbers, each device sends and receives about $2n(p-1)/p$ numbers. Explain why, and what it implies for how data-parallel training scales with the number of devices.

**Solution of Exercise 23.8.**

In a ring, the array is cut into $p$ chunks; a reduce-scatter passes chunks round the ring $p - 1$ times, each device sending $n/p$ numbers each time, and an all-gather does the same again: $2n(p-1)/p$ in all, nearly $2n$ whatever $p$. The communication per device does not grow with the number of devices, but it does not shrink either, while each device’s compute shrinks as $1/p$: beyond some $p$, communication dominates.

## 23.9 Problem: Where the Time Goes

**Problem 23.1.**

Weekend problem — measure before you buy

The chapter’s sessions, formats, network and runs.

**Part I — The budget.**

1. How large is the data in each format?
2. What are the load, batch and compute times per step?
3. What share of the step goes to data for each storage format, and with the gather?
4. What would you change first, and why?

**Part II — Precision.**

5. What are [bfloat16](#def-ml-training-infrastructure-mixed) ’s bits, range and epsilon?
6. How do the float32 and [bfloat16](#def-ml-training-infrastructure-mixed) runs compare?
7. Why is no loss scaling needed with [bfloat16](#def-ml-training-infrastructure-mixed) ?
8. Where would you keep float32 in a trading model’s pipeline?

**Part III — Parallelism and checkpoints.**

9. What do the accumulation and data-parallel checks show?
10. What must a checkpoint contain, and what does leaving part of it out cost?
11. How large is the chapter’s checkpoint?
12. What does a larger effective batch require of the [learning rate](https://one-course.com/books/quant/12/en/chapter/5-trees-and-boosting#def-ml-trees-and-boosting-boosting) ?

**Part IV — The verdict.**

13. State the *named result* : the data-loading share of step time for each storage format, and the largest loss difference between float32 and [bfloat16](#def-ml-training-infrastructure-mixed) training over the run.
14. Would the team need a second [accelerator](#def-ml-training-infrastructure-accelerator) ?
15. When does [model parallelism](#def-ml-training-infrastructure-dist) become necessary for a trading model?
16. How would you make a training run reproducible across machines?
17. What is the risk of a sampler that crosses sessions?
18. How would you measure a loader on shared infrastructure?
19. What would you log about every training run?
20. In one sentence: what is the first thing to measure about a slow training job?

**Solution of Problem 23.1.**

**Part I.**

1. CSV 2.21 MB, NumPy 1.50 MB, memory map 1.50 MB.
2. Load per [epoch](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop) 12.5 ms (CSV), 0.83 ms (NumPy), 0.39 ms (map); batch 0.48 to 0.74 ms looped and 0.27 ms gathered; compute 1.1 to 1.2 ms (measured on the author’s laptop).
3. 0.34 (CSV), 0.31 (NumPy), 0.39 (map, looped) and 0.20 (map, gathered).
4. The batch assembly: a vectorised gather halves the data’s share before any storage change matters.

**Part II.**

1. Eight exponent bits, seven fraction bits; float32’s range; epsilon $2^{-7}$ .
2. Last-50-step mean losses 0.0525 ( [bfloat16](#def-ml-training-infrastructure-mixed) ) and 0.0527 (float32); largest step difference 0.0132 (22%).
3. Its exponent range equals float32’s: small gradients do not underflow.
4. Loss, optimiser state and weights; accumulations of many terms (Book 4’s compensated sums); risk and P&L computations downstream of the model.

**Part III.**

1. Accumulation equals a larger batch to $1.4\times10^{-7}$ ; averaged shard gradients equal the full gradient to $1.5\times10^{-8}$ .
2. Model, optimiser, sampler and generator states and the step; with weights alone the run ended up to 0.039 away.
3. 693 kB.
4. Rescaling with the batch, usually with a warm-up; and a check that the larger batch improves the model at all.

**Part IV.**

1. *Where the time goes.* The data take 0.34, 0.31 and 0.39 of a training step with CSV, NumPy and memory-mapped storage and window-by-window batching, and 0.20 with the memory map and one gather (laptop measurements); the largest loss difference between float32 and [bfloat16](#def-ml-training-infrastructure-mixed) over the 300-step run is 0.0132, while their final losses agree to 0.0002.
2. Not before the loader: a faster device would wait for data a third of the time.
3. When the model’s parameters and activations do not fit on one device: rarely for trading models, often for [language models](https://one-course.com/books/quant/12/en/chapter/14-large-language-models-in-finance#def-ml-large-language-models-in-finance-lm) .
4. Fix seeds, deterministic algorithms, library versions and the data order; accept that different hardware can round differently and compare on metrics.
5. Windows spanning a close and an open teach the model a jump it cannot predict and leak information across days.
6. Many repetitions, medians, the machine’s load recorded, and ratios rather than absolute times.
7. Code version, data version, configuration, seeds, the machine, the checkpoints and the metrics (chapter 25).
8. The share of each step spent waiting for data.

## 23.10 Interview questions

**Interview question 23.1 ★ mle.**

What is [bfloat16](#def-ml-training-infrastructure-mixed) and why is it preferred to float16 for training?

**Solution of Interview question 23.1.**

A 16-bit format with float32’s exponent and seven fraction bits: less precision than float16 but the full range, so gradients do not underflow and no loss scaling is needed; matrix units run it at full speed.

*What the interviewer is looking for: the bit layout and the range argument.*

**Interview question 23.2 ★★ mle, developer.**

Your GPU utilisation is 30%. How do you find out why?

**Solution of Interview question 23.2.**

Profile the step: time data loading, host-to-device transfer and compute separately; check loader parallelism and prefetching, batch size, synchronisation points (logging, .item() calls), and small kernels; fix the largest.

*What the interviewer is looking for: measurement first, then the usual culprits.*

**Interview question 23.3 ★★ mle.**

Explain [data parallelism](#def-ml-training-infrastructure-dist) and what an [all-reduce](#def-ml-training-infrastructure-dist) does.

**Solution of Interview question 23.3.**

Each device holds a copy of the model and processes its share of the batch; after the backward pass an [all-reduce](#def-ml-training-infrastructure-dist) leaves every device with the average gradient, and all apply the same update.

*What the interviewer is looking for: replicas, sharded batches, averaged gradients.*

**Interview question 23.4 ★★ mle, researcher.**

How would you store and sample a year of order-book data for training sequence models?

**Solution of Interview question 23.4.**

Per-instrument, per-session binary columns (memory-mapped or chunked), an index of session boundaries, samplers that draw windows within sessions and respect the train/validation split in time, batch gathering in one read, and normalisation fitted on training data only.

*What the interviewer is looking for: binary storage, session-aware sampling and time-respecting splits.*

**Interview question 23.5 ★★ mle.**

What does a [training checkpoint](#def-ml-training-infrastructure-ckpt) need to contain for bitwise resumption?

**Solution of Interview question 23.5.**

Model parameters and buffers, optimiser state, learning-rate schedule state, data sampler position and random state, all random generator states, the step or [epoch](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop) count, and the code and data versions.

*What the interviewer is looking for: the optimiser and the sampler, not only the weights.*

**Interview question 23.6 ★★★ mle, developer.**

[Gradient accumulation](#def-ml-training-infrastructure-dist) and [data parallelism](#def-ml-training-infrastructure-dist) are said to be equivalent. When are they not?

**Solution of Interview question 23.6.**

When the model has per-batch statistics ([batch normalisation](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-dropout)), when the micro-batches are not equal in size, when the order of floating-point summation matters for bitwise results, and when devices see different data distributions.

*What the interviewer is looking for: batch-coupled layers, unequal shards and summation order.*
