---
title: "Low-Latency Inference"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 26
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/26-low-latency-inference
---

# Chapter 26 — Low-Latency Inference

A gradient-boosted model of 300 trees answers in 265 microseconds through its library’s Python interface and in 2.6 microseconds as generated C++, measured on the same laptop; the model did not change, and the two answers are equal to the last bit. Most of the first figure is not the model: it is the interpreter, the wrapper, the conversion of one row into an array and back. This chapter is about taking models out of that environment: compiling trees into code, shrinking small networks to eight-bit integers, deciding when to batch, and doing it in C++20 and Rust with the exact results of the Python reference. It rests on Book 13’s method for measuring latency, and its timings are one laptop’s, labelled as such; what the tests check is parity, and what the chapter concludes is about orders of magnitude.

## 26.1 The inference budget

**Definition 26.1 (Inference latency).**

*Inference latency* is the time from the moment a model’s inputs are available to the moment its output is, measured per call and reported as a distribution (median and tail), within the strategy’s latency budget (Book 13, chapter 1).

A market-making or execution model decides inside a budget of a few microseconds to a few hundred, shared with feature computation (chapter 24), risk checks and the network. Inference must fit its share at the tail, not the median: the decision that arrives late is the one taken when the market moved. The chapter’s two models are those of `firm.mlinfer`’s fixture, fitted to a synthetic problem with 16 features: a LightGBM forest (chapter 5) of 300 trees with 15 leaves each, 4 200 splits in all, and a [multilayer perceptron](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-mlp) (chapter 7) of $16\to32\to16\to1$ units. On 20 000 new observations the forest’s rank information coefficient is 0.224 and the network’s 0.182.

## 26.2 Compiling trees

**Definition 26.2 (Model compilation).**

*Model compilation* translates a trained model into code or data structures specialised for inference on a target (generated source, a flat array layout, a hardware description), removing the training framework from the prediction path (Asadi, Lin and de Vries, 2014; Lucchese and co-authors, 2015, for ranking forests).

The forest is exported two ways. As flat arrays of feature indices, thresholds and child indices, one entry per node, which a loop walks ([Listing 26.2](#lst-ml-inf-loop)): the same loop runs in Python, C++20 and Rust. And as generated C++ with one nested `if` per split and the thresholds as literals ([Listing 26.1](#lst-ml-inf-gen)), 17 107 lines that the compiler turns into straight branches. Both reproduce LightGBM’s predictions exactly on the 200 shared test vectors, because they visit the same leaves and add the leaf values in the same order in double precision; the thresholds survive the trip through text because they are written with the shortest representation that reads back to the same double.

[Figure 26.1](#fig-ml-inf-latency) shows the measured latencies. LightGBM’s scikit-learn interface takes 265 µs per single-row call at the median and 544 µs at the 99th percentile; the same flat loop in Python is no faster (300 µs), because the interpreter’s cost per node dominates. In C++ the loop takes 6.2 µs and the generated branches 2.6 µs (5.7 µs at the 99th percentile), a hundred times less than the library call at both the median and the tail; the Rust loop takes 7.8 µs. The branches win over the loop because the thresholds and child addresses are in the instruction stream rather than in memory, and the compiler can lay out the likely path; with 300 trees and 4 200 splits the whole forest still fits in the caches (Book 13, chapter 3).

![Single-call latency of the same two models by execution path; 200 000 calls in C++ and Rust, 5 000 in Python (300 for the pure-Python forest), timed one by one. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2 at low priority, no isolated cores; other jobs may have run. Data: bench_infer.py, measured_latency.csv.](https://one-course.com/images/onecourse/chapters/quant-12/ml-low-latency-inference/fig-5c707857fa7e.svg)

***Figure 26.1.** Single-call latency of the same two models by execution path; 200 000 calls in C++ and Rust, 5 000 in Python (300 for the pure-Python forest), timed one by one. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2 at low priority, no isolated cores; other jobs may have run. Data: `bench_infer.py`, `measured_latency.csv`.*

## 26.3 Small networks: layout, fusion and quantisation

**Definition 26.3 (Operator fusion, quantisation, post-training quantisation, quantisation-aware training).**

*Operator fusion* computes several consecutive operations (a matrix product, a bias, an activation, a rescaling) in one pass over the data. *Quantisation* represents weights and activations as small integers with scale factors, so that inference runs in integer arithmetic (Jacob and co-authors, 2018). *Post-training quantisation* derives the integers and scales from a trained float model and a calibration sample; *quantisation-aware training* fine-tunes the model with the rounding simulated in the forward pass, so that it learns weights that survive it (Nagel and co-authors, 2021).

The chapter’s network is quantised symmetrically, one scale per tensor: weights and activations to integers in $[-127, 127]$, biases to 32-bit integers at the product of the input and weight scales, and between layers a fixed-point multiplier and shift that bring the 32-bit accumulator back to eight bits, rounding half away from zero ([Listing 26.3](#lst-ml-inf-mlp)). Each layer is one fused pass: products, bias, rescaling and ReLU. The integer arithmetic makes parity exact by construction: the NumPy reference, the C++ kernel and the Rust kernel agree on every one of the 200 test vectors to the bit. The eight-bit weights take 1 236 bytes against 4 356 in float32, and one call takes 0.47 µs in C++ and 0.58 µs in Rust.

|  | bits per weight and activation |
| --- | --- |
| rank IC on 20 000 new observations | 8 | 6 | 4 |
| quantised after training | 0.182 | 0.181 | 0.143 |
| quantisation-aware [fine-tuning](https://one-course.com/books/quant/12/en/chapter/10-representation-learning#def-ml-representation-learning-ssl), then quantised | 0.179 | 0.176 | 0.164 |

***Table 26.1.** The network’s information coefficient after [quantisation](#def-ml-low-latency-inference-quant) (float: 0.182; the forest: 0.224). Data: `ml_infer.accuracy`.*

[Table 26.1](#tab-ml-inf-quant) holds the second half of the chapter’s named result. At eight bits, [quantisation](#def-ml-low-latency-inference-quant) after training costs nothing measurable (0.1817 against 0.1822; the integer model’s outputs differ from the float model’s by 4.5% of their standard deviation, and the ranking barely moves); quantisation-aware [fine-tuning](https://one-course.com/books/quant/12/en/chapter/10-representation-learning#def-ml-representation-learning-ssl), which spends ten [epochs](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop) adapting to the rounding, loses a little (0.179), because the [fine-tuning](https://one-course.com/books/quant/12/en/chapter/10-representation-learning#def-ml-representation-learning-ssl) itself moves the model. At six bits the story is the same. At four, [post-training quantisation](#def-ml-low-latency-inference-quant) loses a fifth of the signal (0.143) and [quantisation-aware training](#def-ml-low-latency-inference-quant) recovers more than half of the loss (0.164). Eight bits is free for a network this small; fewer bits are where training for the rounding pays.

## 26.4 Batching against latency

**Definition 26.4 (Dynamic batching).**

*Dynamic batching* groups requests that arrive close together into one call of the model, trading the waiting time of the first request for a lower cost per request.

For a library whose per-call overhead dwarfs its per-row cost, batching is a large lever ([Figure 26.2](#fig-ml-inf-batch)): LightGBM’s call takes 273 µs for one row and 3.1 ms for 1 024, so the cost per row falls from 273 µs to 3.0 µs. A research job scoring a universe each second should batch. A trading decision on one instrument cannot wait for others’ requests to arrive, and a compiled model already costs a few microseconds per row: there, batching buys throughput the strategy does not need at the price of latency it cannot afford. When batching is used, latency is measured from each request’s arrival, not from the batch’s start, or the waiting time vanishes from the statistics (coordinated omission, Book 13, chapter 5).

![LightGBM’s prediction latency (median) against batch size through its Python interface, per call and per row. Same machine and conditions as . Data: bench_infer.py, measured_batch.csv.](https://one-course.com/images/onecourse/chapters/quant-12/ml-low-latency-inference/fig-bb350f4c32b3.svg)

***Figure 26.2.** LightGBM’s prediction latency (median) against batch size through its Python interface, per call and per row. Same machine and conditions as [Figure 26.1](#fig-ml-inf-latency). Data: `bench_infer.py`, `measured_batch.csv`.*

## 26.5 Inference in C++, in Rust and on programmable hardware

The kernels load their model once, allocating, and then predict without touching the heap: fixed-size buffers on the stack, no exceptions, no virtual calls, the model’s arrays read-only and shared. The Rust version ([Listing 26.4](#lst-ml-inf-rust)) is the same algorithm with bounds checks, which cost it about a quarter against C++ here; both could remove them with unchecked indexing once the model’s arrays are validated at load time. Beyond the CPU, programmable hardware runs trees and small networks in a fixed number of clock cycles, laid out as a circuit: Duarte and co-authors (2018) compiled small networks for FPGAs in the trigger systems of particle physics, where decisions have fixed latency budgets of the same kind. The same parity discipline applies: a hardware model is validated against the Python reference on shared vectors, bit for bit where the arithmetic is integer.

**Method 26.5 (Putting a model in the fast path).**

1. Measure the library call’s median and tail on the target machine, one request at a time, before changing anything.
2. Export the model to a portable form and compile it (trees to code or flat arrays; networks to fused integer kernels), with no allocation on the hot path.
3. Test parity on shared vectors: exact for trees and integer networks, within stated tolerances for floating-point networks.
4. Quantise after training and measure the loss of signal; use [quantisation-aware training](#def-ml-low-latency-inference-quant) only when the loss matters.
5. Register the compiled artefact with its parity report next to the model it came from (chapter 25).

## 26.6 Tutorial: two microseconds for a forest

**Goal.** Export a forest and a network, compile them to C++20 and Rust, check parity, quantise, and measure. **End state:** Figures [26.1](#fig-ml-inf-latency) and [26.2](#fig-ml-inf-batch), [Table 26.1](#tab-ml-inf-quant).

1. **Trees as generated code.** `def cpp_branches (f, name=" forest_branches " ): """One C++ function per forest: each tree an if/else nest with the thresholds as literals.""" lines = [f " inline double { name} (const double* x) {{ " , " double s = 0.0; " ] def emit (n, ind): pad = " " * ind if n < 0 : lines.append(f " { pad} s += { float (f[' value ' ][-n - 1 ])!r} ; " ) return lines.append(f " { pad} if (x[ { int (f[' feature ' ][n])} ] <= { float (f[' threshold ' ][n])!r} ) {{ " ) emit(int (f[" left " ][n]), ind + 1 ) lines.append(f " { pad} }} else {{ " ) emit(int (f[" right " ][n]), ind + 1 ) lines.append(f " { pad} }} " ) for root in f[" roots " ]: emit(int (root), 1 ) lines += [" return s; " , " } " ] return " \n " .join(lines) + " \n "` **Listing 26.1.** Generating nested C++ branches from the exported forest. code/firm/mlinfer/firm_mlinfer.py
2. **The flat forest in C++20.** `double predict (const double * x) const { double s = 0.0 ; for (int32_t n : roots) { while (n >= 0 ) n = x[feature[n]] <= threshold[n] ? left[n] : right[n]; s += value[static_cast <size_t >(-n - 1 )]; } return s; }` **Listing 26.2.** Walking the flat node arrays. code/firm/mlinfer/cpp/mlinfer.hpp
3. **The int8 network in C++20.** `// round(acc * mult / 2^shift), halves away from zero, clamped to int8. inline int64_t requant (int64_t acc, int64_t mult, int shift) { const int64_t p = acc * mult; const int64_t r = ((p < 0 ? -p : p) + (int64_t {1 } << (shift - 1 ))) >> shift; return std::clamp<int64_t >(p < 0 ? -r : r, -127 , 127 ); } struct Mlp { double s_in = 1.0 ; std::vector<Layer> layers; double predict (const double * x) const { std::array<int64_t , kMaxWidth> a{}, b{}; const Layer& first = layers.front(); for (int i = 0 ; i < first.in; ++i) a[static_cast <size_t >(i)] = std::clamp<int64_t >(static_cast <int64_t >(std::nearbyint(x[i] / s_in)), -127 , 127 ); for (size_t l = 0 ; l < layers.size(); ++l) { const Layer& L = layers[l]; for (int o = 0 ; o < L.out; ++o) { int64_t acc = L.b[static_cast <size_t >(o)]; const int32_t * row = &L.w[static_cast <size_t >(o * L.in)]; for (int i = 0 ; i < L.in; ++i) acc += static_cast <int64_t >(row[i]) * a[static_cast <size_t >(i)]; if (l + 1 == layers.size()) return static_cast <double >(acc) * L.prod; b[static_cast <size_t >(o)] = std::max<int64_t >(requant(acc, L.mult, L.shift), 0 ); } a = b; } return 0.0 ; }` **Listing 26.3.** Integer inference with fixed-point requantisation. code/firm/mlinfer/cpp/mlinfer.hpp
4. **The flat forest in Rust.** `pub fn predict (&self , x: & [f64 ]) -> f64 { let mut s = 0.0 ; for &root in &self .roots { let mut n = root; while n >= 0 { let i = n as usize ; n = if x[self .feature[i] as usize ] <= self .threshold[i] { self .left[i] } else { self .right[i] }; } s += self .value[(-n - 1 ) as usize ]; } s }` **Listing 26.4.** The same loop in Rust. code/firm/mlinfer/rust/src/lib.rs
5. **Run** `make_mlinfer_fixture.py` (the fixture), the C++ and Rust parity tests ( `make test-code` ), `ml_infer.accuracy()` , and, once, `bench_infer.py` under `nice` .

**What to change next.** Store the forest’s nodes in breadth-first order to improve locality, or evaluate several trees at once with vector instructions (Book 13, chapter 14), and measure.

## 26.7 Build: the inference kernels

**Purpose.** Models in the fast path with the Python reference’s exact results.

**Interface.** Python: `export_forest`, `forest_predict`, `write_forest`, `read_forest`, `cpp_branches`, `MLP`, `fit_mlp(model, X, y, epochs, seed, qat, lr, batch, bits)`, `quantise(model, X_calib, bits)`, `int8_forward`, `write_mlp`; `make_mlinfer_fixture.py`. C++20: `mlinfer::Forest`, `read_forest`, `Mlp`, `read_mlp`, the generated `forest_branches`. Rust: `Forest`, `Mlp`, `requant`.

**Rules.** No allocation, exception or virtual call on the prediction path; parity on shared vectors is exact; timings are measured, labelled with the machine, and never asserted.

**Acceptance tests.** `code/firm/mlinfer/tests/` (Python: export equals LightGBM, fixture reproducible, integer reference against the float model), `cpp/mlinfer_test.cpp` and `rust` (`cargo test`): exact parity of the flat forest, the generated branches and the int8 network on the 200 vectors.

**Stretch.** SIMD evaluation of several trees; per-channel [quantisation](#def-ml-low-latency-inference-quant); a hardware-description export of the int8 network.

Sources and further reading

- B. Jacob and co-authors, “Quantization and training of neural networks for efficient integer-arithmetic-only inference”, arXiv:1712.05877, 2018.
- M. Nagel and co-authors, “A white paper on neural network quantization”, arXiv:2106.08295, 2021.
- N. Asadi, J. Lin and A. P. de Vries, “Runtime optimizations for tree-based machine learning models”, *IEEE Transactions on Knowledge and Data Engineering* , 2014.
- C. Lucchese and co-authors, “QuickScorer: a fast algorithm to rank documents with additive ensembles of regression trees”, *SIGIR* , 2015.
- J. Duarte and co-authors, “Fast inference of deep neural networks in FPGAs for particle physics”, arXiv:1804.06913, 2018.

## 26.8 Exercises

**Exercise 26.1 ★.**

A symmetric int8 tensor has scale $s = \max|w|/127$. What is the largest rounding error of one weight, as a share of $\max|w|$?

**Solution of Exercise 26.1.**

Rounding to the nearest multiple of $s$ errs by at most $s/2 = \max|w|/254$: 0.39% of the largest weight, and a far larger share of a small weight.

**Exercise 26.2 ★.**

From [Figure 26.1](#fig-ml-inf-latency), what is the ratio of the generated C++ forest’s 99th-percentile latency to the library call’s?

**Solution of Exercise 26.2.**

$5\,738/544\,349 = 0.011$: the compiled forest’s 99th percentile is about a ninety-fifth of the library call’s.

**Exercise 26.3 ★.**

Why can the compiled forest reproduce LightGBM’s prediction exactly while a compiled float network usually cannot reproduce PyTorch’s?

**Solution of Exercise 26.3.**

A forest’s prediction is a sum of leaf values chosen by comparisons: the compiled version makes the same comparisons and adds the same doubles in the same order. A network’s float arithmetic depends on the order of the sums in each product, on fused multiply-adds and on vectorisation, which differ between libraries and compilers; parity is then a tolerance, unless the arithmetic is integer, as in the int8 network.

**Exercise 26.4 ★★.**

Write the fixed-point multiplier and shift for a real multiplier of 0.0123. Why use integers for it at all?

**Solution of Exercise 26.4.**

$0.0123 = 0.7872\times2^{-6}$; $M = \mathrm{round}(0.7872\times2^{31}) = 1\,690\,499\,128$ and a total shift of $6 + 31 = 37$: the product is $\mathrm{acc}\times M$ shifted right by 37 bits. Integers make the result identical on every platform and avoid floating point on hardware that lacks it or where it is slower.

**Exercise 26.5 ★★.**

Why did [quantisation-aware training](#def-ml-low-latency-inference-quant) lose a little at eight bits and win at four?

**Solution of Exercise 26.5.**

At eight bits the rounding is too fine to hurt, so [fine-tuning](https://one-course.com/books/quant/12/en/chapter/10-representation-learning#def-ml-representation-learning-ssl) only moves the model away from its trained optimum; at four bits the rounding is coarse enough to damage the post-training model, and a model trained to expect it recovers more than half of the loss (0.164 against 0.143, float 0.182).

**Exercise 26.6 ★★.**

*Find the flaw.* “Our compiled model’s median latency is 800 nanoseconds, measured by timing 10 000 predictions in a batch and dividing by 10 000.”

**Solution of Exercise 26.6.**

Dividing a batch’s time by its size gives the throughput’s inverse, not the latency of a request: it hides the first request’s wait, the tail, and anything that happens once per call. Time single calls, as they arrive, and report the distribution.

**Exercise 26.7 ★★★.**

*Coding.* Quantise the network to four bits after training with one weight scale per output unit instead of one per tensor. How much of the loss against the float model does it recover?

**Solution of Exercise 26.7.**

Per-output scales give each row of weights its own grid, so small rows are not flattened by one large weight elsewhere: at four bits the information coefficient rises from 0.143 to 0.153 (float 0.182), about a third of the loss recovered without any retraining. At eight bits it is 0.182, as before.

**Exercise 26.8 ★★★.**

With symmetric scales $s_x$, $s_w$ and $s_y$ for a layer’s input, weights and output, write the layer’s float output in terms of the integers $X_i$, $W_i$ and an integer bias $B$, and derive the scale at which $B$ must be stored and the multiplier that maps the accumulator to the output’s integers.

**Solution of Exercise 26.8.**

With $x_i = s_xX_i$ and $w_i = s_wW_i$, $\sum_iw_ix_i + b = s_xs_w\big(\sum_iW_iX_i + b/(s_xs_w)\big)$: storing $B =
\mathrm{round}(b/(s_xs_w))$ in the accumulator’s scale lets the bias be added to the integer sum. The output’s integer is $Y = \mathrm{round}\big((s_xs_w/s_y)(\sum_iW_iX_i + B)\big)$: the multiplier is $s_xs_w/s_y$, applied as a fixed-point product and a shift.

## 26.9 Problem: Two Microseconds for a Forest

**Problem 26.1.**

Weekend problem — the model is not the latency

The chapter’s forest, network, kernels and measurements.

**Part I — The models.**

1. What are the two models, and how good are they?
2. How is the forest exported, and in how many forms?
3. Why are the kernels exact?
4. How large is each model in memory?

**Part II — Latency.**

5. Give the median and 99th-percentile latencies of each path.
6. Where does the library call’s time go?
7. Why do the generated branches beat the loop?
8. What does the Rust version cost against C++, and why?

**Part III — [Quantisation](#def-ml-low-latency-inference-quant) and batching.**

9. What does int8 cost in information coefficient, after training and with [quantisation-aware training](#def-ml-low-latency-inference-quant) ?
10. What changes at four bits?
11. How does LightGBM’s latency scale with batch size?
12. When would you batch?

**Part IV — The verdict.**

13. State the *named result* : the ratio of the compiled model’s tail latency to the library call’s, and the IC lost to int8 [quantisation](#def-ml-low-latency-inference-quant) with and without [quantisation-aware training](#def-ml-low-latency-inference-quant) .
14. Which model would you deploy in a five-microsecond budget, and in what form?
15. What must the registry hold for a compiled model?
16. How would you make these measurements trustworthy on a production machine?
17. What would change with 5 000 trees?
18. Where would programmable hardware be worth its cost?
19. How would a feature-store skew (chapter 24) show up at this layer?
20. In one sentence: where does inference time go?

**Solution of Problem 26.1.**

**Part I.**

1. A 300-tree LightGBM forest (IC 0.224) and a $16\to32\to16\to1$ network (IC 0.182).
2. As flat node arrays (walked by the same loop in Python, C++ and Rust) and as generated C++ branches.
3. Trees: the same comparisons and the same order of double additions; the network: integer arithmetic.
4. The forest’s 4 200 splits and 4 500 leaves in arrays of a few hundred kilobytes of text and far less in binary; the network 1 236 bytes in int8 against 4 356 in float32.

**Part II.**

1. Median and 99th percentile (µs): LightGBM 265 and 544; Python forest 300 and 397; PyTorch network 21 and 44; NumPy int8 network 24 and 66; C++ forest loop 6.2 and 13.2; C++ branches 2.6 and 5.7; C++ int8 network 0.47 and 1.1; Rust forest 7.8 and 17.8; Rust int8 network 0.58 and 1.9.
2. The interpreter, argument checks, conversions of the row to and from arrays, and thread-pool machinery; the trees themselves are a few microseconds.
3. Thresholds and branch targets are in the code, not loaded from memory, and the compiler lays out the paths.
4. About a quarter, for bounds checks on every array access; removable after validating the model at load.

**Part III.**

1. Nothing measurable after training (0.182); quantisation-aware [fine-tuning](https://one-course.com/books/quant/12/en/chapter/10-representation-learning#def-ml-representation-learning-ssl) 0.179.
2. [Post-training quantisation](#def-ml-low-latency-inference-quant) falls to 0.143, [quantisation-aware training](#def-ml-low-latency-inference-quant) holds 0.164, per-channel scales 0.153.
3. From 273 µs per call for one row to 3.1 ms for 1 024: 3.0 µs per row.
4. For scoring many requests with no latency constraint (a universe each second); not for a decision on one instrument.

**Part IV.**

1. *Two microseconds for a forest.* The generated C++ forest’s 99th percentile is 5.7 µs against 544 µs for the library call, a ratio of about 1 to 95 (2.6 against 265 µs at the median); int8 [quantisation](#def-ml-low-latency-inference-quant) costs the network nothing measurable after training (0.182) and 0.003 with quantisation-aware [fine-tuning](https://one-course.com/books/quant/12/en/chapter/10-representation-learning#def-ml-representation-learning-ssl) (0.179); at four bits, 0.040 and 0.018.
2. The forest as generated C++ (2.6 µs median, 5.7 at the tail) if its extra signal is worth it, or the int8 network at half a microsecond.
3. The source model, the compiled artefact and its hash, the compiler and flags, and the parity report.
4. Isolated cores, no other jobs, warm caches as in production, single requests at the production rate, and tails over long runs, with the measurement’s own overhead measured.
5. Ten times the code: the generated source may no longer fit the instruction cache, and flat arrays or vectorised evaluation may win.
6. Where the budget is below a microsecond and fixed, and the model is small and stable enough to be worth a hardware cycle.
7. As a difference between the features the compiled model receives and those the research model saw, invisible to the kernel’s parity tests: parity must also be checked on live inputs.
8. Mostly around the model, not in it.

## 26.10 Interview questions

**Interview question 26.1 ★ developer, mle.**

How would you make a gradient-boosted model predict in a few microseconds?

**Solution of Interview question 26.1.**

Export the trees and compile them (generated branches or flat arrays in C++), keep the process hot, avoid allocation and conversions on the path, check parity against the library, and measure tails one request at a time.

*What the interviewer is looking for: compilation, a clean hot path, parity and tail measurement.*

**Interview question 26.2 ★★ mle.**

Explain int8 [quantisation](#def-ml-low-latency-inference-quant): scales, zero points, accumulators and requantisation.

**Solution of Interview question 26.2.**

Real values are integers times a scale (plus a zero point in asymmetric schemes); products accumulate in int32; each layer’s accumulator is rescaled to the next layer’s integers by a fixed-point multiplier; biases are stored at the accumulator’s scale.

*What the interviewer is looking for: scales, accumulators and requantisation.*

**Interview question 26.3 ★★ developer.**

What rules would you impose on code in the inference hot path?

**Solution of Interview question 26.3.**

No heap allocation, locks, system calls, exceptions or virtual dispatch; bounded loops; data laid out for the cache; everything loaded and validated at start-up.

*What the interviewer is looking for: the usual hot-path rules with reasons.*

**Interview question 26.4 ★★ mle, developer.**

How do you test that a compiled model is the model you trained?

**Solution of Interview question 26.4.**

Shared test vectors with reference outputs from the training framework; exact comparison for trees and integer networks, stated tolerances for float ones; the compiled artefact’s hash registered with the source model; checks on live inputs too.

*What the interviewer is looking for: shared vectors, exactness where possible, registration.*

**Interview question 26.5 ★★ developer.**

When does batching help [inference latency](#def-ml-low-latency-inference-latency), and when does it hurt?

**Solution of Interview question 26.5.**

It helps when per-call overhead dominates and requests can wait; it hurts when each request has a deadline, because the first request in a batch waits for the last.

*What the interviewer is looking for: overhead amortisation against waiting time.*

**Interview question 26.6 ★★★ mle.**

[Post-training quantisation](#def-ml-low-latency-inference-quant) or [quantisation-aware training](#def-ml-low-latency-inference-quant)? How would you decide?

**Solution of Interview question 26.6.**

Quantise after training and measure the metric that matters; if the loss is negligible, stop; if not, try finer granularity (per channel) and, if needed, quantisation-aware [fine-tuning](https://one-course.com/books/quant/12/en/chapter/10-representation-learning#def-ml-representation-learning-ssl), and compare on held-out data.

*What the interviewer is looking for: measure first, then escalate.*
