---
title: "Sequence Models on Order Books"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 8
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/8-sequence-models-on-order-books
---

# Chapter 8 — Sequence Models on Order Books

A widely cited paper reports that a convolutional-recurrent network reading the raw states of a ten-level order book outperforms all existing algorithms on a benchmark order-book dataset, and delivers stable out-of-sample accuracy on a year of London Stock Exchange quotes. On a laptop, on a few hours of a simulated book, with labels built honestly and every test session kept out of training, the best sequence model reaches an information coefficient of 0.43, and a [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) on ten hand-made features (the order-flow imbalance, the queue imbalance, recent flow and moves) reaches 0.54. The paper is not wrong about its data; the result is that raw books are a hard input, sequence models need a lot of them, and the claim “deep learning finds the features by itself” has a data cost that must be paid in sessions. This chapter builds the four standard sequence architectures, trains them on order-book windows, and measures what they buy.

## 8.1 Order-book streams as inputs

The chapter’s data are ten-minute `firm.tape` sessions (Book 7, chapter 2), each on its own seed: eight for training, three for [early stopping](https://one-course.com/books/quant/12/en/chapter/5-trees-and-boosting#def-ml-trees-and-boosting-early), eight for testing, and eight more to double the [training set](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) later. Every half second the book is rebuilt from the messages and summarised level by level: for each of the ten best levels on each side, the distance of the price from the mid, in ticks, and the logarithm of the size in lots, forty numbers per snapshot. An input is a window of the last 20 snapshots (ten seconds); its label is the direction of the mean mid over the next ten seconds against the current mid, down, flat or up, with a dead band of a quarter tick ([Figure 8.1](#fig-ml-lobseq-window)). Windows never cross sessions, and the scaling of the forty columns is fitted on training sessions only.

![An input window and its label. The label compares the average mid over the next ten seconds with the mid at the decision time and is flat within a quarter tick; the smoothing follows the DeepLOB literature and makes the label easier to predict than the next move itself.](https://one-course.com/images/onecourse/chapters/quant-12/ml-sequence-models-on-order-books/fig-0695ac0da798.svg)

***Figure 8.1.** An input window and its label. The label compares the average mid over the next ten seconds with the mid at the decision time and is flat within a quarter tick; the smoothing follows the DeepLOB literature and makes the label easier to predict than the next move itself.*

The classes are unbalanced: 73% of the test windows are flat, 15% down and 13% up. Always predicting “flat” is right 73% of the time, so accuracy is a poor score; the chapter reports it for comparison with the literature and ranks the models by the information coefficient of $\P(\text{up}) - \P(\text{down})$ with the forward change, session by session.

## 8.2 Convolutional and recurrent models

**Definition 8.1 (Convolutional neural network).**

A *convolutional neural network* applies the same small filter at every position of its input: a one-dimensional convolutional layer maps a sequence $z_1,\dots,z_T$ of $d$-vectors to $u_t = \mathrm{act}(\sum_{j=0}^{k-1}B_jz_{t+j-\lfloor k/2\rfloor} + c)$, with $k$ filter matrices $B_j$ shared across $t$.

**Definition 8.2 (Recurrent neural network, long short-term memory).**

A *recurrent neural network* reads a sequence one step at a time and carries a hidden state, $h_t = \mathrm{act}(A x_t + C h_{t-1} + c)$. A *long short-term memory* network (LSTM) adds a cell state updated through gates (input, forget and output, each a sigmoid of $x_t$ and $h_{t-1}$), which lets gradients flow over long sequences (Hochreiter and Schmidhuber, 1997).

The CNN of the chapter has two convolutional layers of 32 filters of width five, averages over the window and classifies; the LSTM has 32 units and classifies from its last state. Both read the forty columns of every snapshot of the window.

## 8.3 Temporal convolutions and attention

**Definition 8.3 (Causal convolution, temporal convolutional network).**

A *causal convolution* uses only the current and earlier positions: $u_t = \sum_{j=0}^{k-1}B_jz_{t - jr}$, with dilation $r$. A *temporal convolutional network* (TCN) stacks causal convolutions with dilations $1, 2, 4,\dots$, so that the receptive field grows exponentially with depth, usually with residual connections (Bai, Kolter and Koltun, 2018).

**Definition 8.4 (Attention mechanism, transformer).**

An *attention mechanism* lets each position of a sequence form its output as a weighted average of all positions: with queries $\mathsf Q = ZA_{\mathsf Q}$, keys $\mathsf K = ZA_{\mathsf K}$ and values $\mathsf V = ZA_{\mathsf V}$ computed from the sequence $Z$ ($T\times h$), the output is $\mathrm{softmax}(\mathsf Q\mathsf K^\top/\sqrt h)\mathsf V$. A *transformer* layer combines several attention heads, a position-wise network, residual connections and [layer normalisation](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-dropout); positions are encoded by adding learned or fixed vectors to the inputs (Vaswani et al., 2017).

The TCN has four causal blocks (dilations 1 to 8, filters of width three, receptive field 31 snapshots, more than the window) and classifies from the last position; the attention model is one [transformer](#def-ml-sequence-models-on-order-books-attention) layer with two heads of width 32 and learned positions, also read at the last position.

![Information coefficient of ( up) - ( down) with the forward change, averaged over eight test sessions (and two seeds for the sequence models); bars show the standard deviation across sessions and seeds. The logistic regression reads ten hand-made features at the decision time; the sequence models read the raw ten-level book over ten seconds. Data: ml_lobseq.table.](https://one-course.com/images/onecourse/chapters/quant-12/ml-sequence-models-on-order-books/fig-33b2d1d20ab4.svg)

***Figure 8.2.** Information coefficient of $\P(\text{up}) - \P(\text{down})$ with the forward change, averaged over eight test sessions (and two seeds for the sequence models); bars show the standard deviation across sessions and seeds. The [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) reads ten hand-made features at the decision time; the sequence models read the raw ten-level book over ten seconds. Data: `ml_lobseq.table`.*

## 8.4 What the literature replicates

On the chapter’s data ([Figure 8.2](#fig-ml-lobseq-models)), the multinomial [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) on the hand-built features of Book 8 (chapter 14; order-flow imbalance over three windows, queue imbalance, flow and past moves) reaches an accuracy of 0.781 and an IC of 0.543. Among the sequence models the [transformer](#def-ml-sequence-models-on-order-books-attention) does best (accuracy 0.775, IC 0.426), then the TCN (0.766, 0.366), the LSTM (0.729, 0.292) and the CNN (0.727, 0.060): the CNN averages its filters over the window and so cannot tell the last snapshot from the first, and it barely beats always predicting “flat”. The spread across sessions and seeds (0.15 to 0.20 of IC for the sequence models, 0.09 for the [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic)) is as large as the gaps.

The models were early-stopped after one to four [epochs](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop): eight sessions are 8 808 windows, and a model with thousands of parameters that reads 800 numbers per window fits their noise quickly. Doubling the data helps ([Figure 8.3](#fig-ml-lobseq-scaling)): the TCN’s IC goes from 0.13 on four sessions to 0.30 on eight and 0.44 on sixteen, while the [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) stays between 0.54 and 0.57 throughout. The sequence models are learning from the raw book what the hand-made features give the linear model directly, and they are still learning it at sixteen sessions.

![Test IC as the training set grows from four to sixteen sessions (one seed; the same eight test sessions). Data: ml_lobseq.scaling.](https://one-course.com/images/onecourse/chapters/quant-12/ml-sequence-models-on-order-books/fig-731cf3cf4e7f.svg)

***Figure 8.3.** Test IC as the [training set](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) grows from four to sixteen sessions (one seed; the same eight test sessions). Data: `ml_lobseq.scaling`.*

The published evidence points the same way once the evaluation is strict. Zhang, Zohren and Roberts’s DeepLOB outperformed the earlier algorithms on the FI-2010 benchmark (Ntakaris et al.) and kept a stable out-of-sample accuracy on a year of London Stock Exchange quotes, including instruments it was not trained on. Sirignano and Cont, training on billions of US quotes, found a stable, universal relation between order-flow history and the direction of the next move, and that a model pooled across stocks beats stock-specific ones: scale is what made the raw input work. Kolm, Turiel and Westray, on 115 Nasdaq stocks, found that simpler networks trained on order flow significantly outperform most models trained directly on order books. The LOBCAST benchmark of fifteen published deep models (Prata et al., 2024) found that all of them lose a large part of their performance on data they were not tuned on. Two lessons carry over to any order-book model: build the stationary features the microstructure literature already knows (order flow, imbalance, Book 7, chapter 8), and judge a model on sessions, days and instruments it has never seen.

**Method 8.5 (Training a sequence model on order books).**

1. Split by session (day, instrument), never by window; fit every normalisation on training sessions.
2. Fix the label’s horizon and dead band from the trade the model serves; report the class shares and a score that does not reward predicting the majority class (the IC, or a cost-aware P&L).
3. Start from a linear model on order-flow and imbalance features; give the sequence model the same features as extra channels before asking it to find them.
4. Report the spread across sessions and seeds, and a learning curve in sessions.

## 8.5 Tutorial: DeepLOB on a laptop

**Goal.** Turn simulated sessions into order-book windows, train four sequence models and a logistic baseline, and measure the learning curve in sessions. **End state:** Figures [8.2](#fig-ml-lobseq-models) and [8.3](#fig-ml-lobseq-scaling).

1. **Windows and labels**, inside each session. `def windows (S, mid, T: int = 40 , horizon: int = 10 , threshold: float = 0.25 ): n = len (S) idx = np.arange(T - 1 , n - horizon) fut = np.array([mid[i + 1 :i + 1 + horizon].mean() for i in idx]) - mid[idx] y = np.where(fut > threshold, 2 , np.where(fut < -threshold, 0 , 1 )) X = np.stack([S[i - T + 1 :i + 1 ] for i in idx]) return X, y, fut, idx class Scaler : def fit (self , X): flat = X.reshape(-1 , X.shape[-1 ]) self .mu, self .sd = flat.mean(axis=0 ), flat.std(axis=0 ) + 1e-6 return self def transform (self , X): return ((X - self .mu) / self .sd).astype(np.float32)` **Listing 8.1.** Windows of snapshots, smoothed direction labels, and a scaler fitted on training windows. code/firm/lobseq/firm_lobseq.py
2. **The TCN and the attention model.** `class _CausalConv (nn.Module): def __init__(self , cin, cout, k, dil): super ().__init__() self .pad = (k - 1 ) * dil self .conv = nn.Conv1d(cin, cout, k, dilation=dil) def forward (self , x): return self .conv(nn.functional.pad(x, (self .pad, 0 ))) class TCN (nn.Module): """Dilated causal convolutions (dilations 1, 2, 4, 8) with residual connections; the last step feeds the head.""" def __init__(self , d, h=32 , k=3 ): super ().__init__() self .inp = nn.Conv1d(d, h, 1 ) self .blocks = nn.ModuleList([_CausalConv(h, h, k, 2 **i) for i in range (4 )]) self .out = nn.Linear(h, 3 ) def forward (self , x): z = self .inp(x.transpose(1 , 2 )) for b in self .blocks: z = z + torch.relu(b(z)) return self .out(z[:, :, -1 ]) class Attention (nn.Module): """One transformer-encoder layer (two heads) over the window, with learned positions; the last step feeds the head.""" def __init__(self , d, T, h=32 ): super ().__init__() self .inp = nn.Linear(d, h) self .pos = nn.Parameter(torch.zeros(1 , T, h)) self .enc = nn.TransformerEncoderLayer(h, nhead=2 , dim_feedforward=64 , dropout=0.0 , batch_first=True ) self .out = nn.Linear(h, 3 ) def forward (self , x): z = self .enc(self .inp(x) + self .pos) return self .out(z[:, -1 ])` **Listing 8.2.** Causal dilated convolutions and a one-layer transformer. code/firm/lobseq/firm_lobseq.py
3. **Run** `ml_lobseq.table()` , `scaling(4)` , `scaling(8)` , `scaling(16)` and `fig_lobseq.py` .

**What to change next.** Feed the TCN the hand-made features as ten extra channels and see how much of the gap to the [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) closes; replace the smoothed label by the sign of the next mid change.

## 8.6 Build: order-book sequence datasets and models

**Purpose.** Order-book sequence models trained and scored the same way on simulated sessions today and on Book 10’s exchange simulator’s recorded tapes when they are available.

**Interface.** `snapshots(tape, step, levels, start)`, `windows(S, mid, T, horizon, threshold)`, `Scaler`, `CNN`, `LSTMNet`, `TCN`, `Attention`, `train(model, X, y, Xv, yv, epochs, lr, batch, patience, seed)`, `evaluate(model, X, y, fwd)`.

**Rules.** A session’s windows stay in one split; the book is rebuilt from messages with nothing after the snapshot’s time; one CPU thread, deterministic algorithms.

**Acceptance tests.** `code/firm/lobseq/tests/`: snapshots match a book rebuilt by hand at a given time and use no later message; windows and labels by hand on a short series; the TCN’s output at a step does not depend on later inputs (causality); each model trains on a planted pattern and beats chance.

**Stretch.** Market-by-order inputs (each order’s age and queue position); a model pooled across several simulated instruments, in the spirit of Sirignano and Cont.

Sources and further reading

- Z. Zhang, S. Zohren and S. Roberts, “DeepLOB: deep convolutional neural networks for limit order books”, *IEEE Transactions on Signal Processing* 67(11), 2019.
- J. Sirignano and R. Cont, “Universal features of price formation in financial markets: perspectives from deep learning”, *Quantitative Finance* 19(9), 2019.
- P. N. Kolm, J. Turiel and N. Westray, “Deep order flow imbalance: extracting alpha at multiple horizons from the limit order book”, *Mathematical Finance* 33(4), 2023.
- M. Prata et al., “LOB-based deep learning models for stock price trend prediction: a benchmark study”, *Artificial Intelligence Review* 57, 2024.
- A. Ntakaris, M. Magris, J. Kanniainen, M. Gabbouj and A. Iosifidis, “Benchmark dataset for mid-price forecasting of limit order book data with machine learning methods”, *Journal of Forecasting* 37(8), 2018.
- S. Hochreiter and J. Schmidhuber, “Long short-term memory”, *Neural Computation* 9(8), 1997; S. Bai, J. Z. Kolter and V. Koltun, “An empirical evaluation of generic convolutional and recurrent networks for sequence modeling”, arXiv, 2018; A. Vaswani et al., “Attention is all you need”, *NeurIPS* , 2017.

## 8.7 Exercises

**Exercise 8.1 ★.**

A TCN has filters of width three and dilations 1, 2, 4 and 8. What is its receptive field in snapshots? With width five?

**Solution of Exercise 8.1.**

$1 + (3 - 1)(1 + 2 + 4 + 8) = 31$ snapshots; with width five, $1 + 4\times15 = 61$.

**Exercise 8.2 ★.**

The test windows are 72.7% flat, 14.7% down and 12.5% up. What accuracy does always predicting “flat” reach, and what does a random guess in the class proportions reach?

**Solution of Exercise 8.2.**

Always “flat”: 72.7%. Guessing in the class proportions: $0.727^2 + 0.147^2 + 0.125^2 = 56.6\%$.

**Exercise 8.3 ★.**

One attention head over $T = 20$ positions of width $h = 32$: how many multiplications does $\mathsf Q\mathsf K^\top$ take, and how does it scale with $T$?

**Solution of Exercise 8.3.**

$T\times T\times h = 20\times20\times32 = 12\,800$ multiply-adds; quadratic in the sequence length (and linear in the width).

**Exercise 8.4 ★★.**

Why does the CNN, which averages over the window, do so badly, and what small change would fix it?

**Solution of Exercise 8.4.**

Averaging the convolution’s outputs over the window weights the snapshot of twenty seconds ago as much as the current one, although only the recent state predicts the next seconds. Read the last position instead (as the TCN and the [transformer](#def-ml-sequence-models-on-order-books-attention) do), or flatten the time axis into the classifier.

**Exercise 8.5 ★★.**

The label averages the mid over the next ten seconds. Why does that make it easier to predict than the next mid change, and what does it mean for trading?

**Solution of Exercise 8.5.**

The mean of the next ten seconds’ mids moves less and more smoothly than the next move, and its first seconds continue the current flow, which the inputs see; much of its predictability comes from that overlap. A trade earns the price change to its exit, not an average of future mids, so a model scored on the smoothed label overstates the tradable edge; score it on the P&L of the trade it would drive (chapter 29).

**Exercise 8.6 ★★.**

*Find the flaw.* “We cut our 50 days of order-book windows into train and [test sets](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) at random and our network is 88% accurate.”

**Solution of Exercise 8.6.**

Random splits put windows from the same day, overlapping in time, in both sets: the network is tested on near-copies of its training windows (chapter 3). And 88% must be compared with always predicting the majority class. Split by day, purge the edges, and report an IC or a P&L.

**Exercise 8.7 ★★★.**

*Coding.* Train the LSTM on sixteen sessions (`scaling(16, name=’LSTM’)`). Report its IC and compare with the TCN’s at sixteen sessions.

**Solution of Exercise 8.7.**

`scaling(16, name=’LSTM’)`: IC 0.381, against the TCN’s 0.441 and the [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic)’s 0.568 on the same sixteen sessions.

**Exercise 8.8 ★★★.**

Show that a stack of [causal convolutions](#def-ml-sequence-models-on-order-books-tcn) of width $k$ with dilations $1, 2,\dots,2^{L-1}$ has a receptive field of $1 + (k - 1)(2^L - 1)$ positions and that its output at $t$ depends on no input after $t$.

**Solution of Exercise 8.8.**

A causal layer with dilation $r$ at position $t$ reads $t, t - r,\dots, t - (k-1)r$: it extends the reach back by $(k-1)r$ and never forward. Stacking layers adds the reaches: $(k-1)(1 + 2 + \dots + 2^{L-1}) = (k-1)(2^L - 1)$, plus the position itself. By induction on the layers, every output at $t$ is a function of inputs at $t$ or earlier; the residual connections add the same position only.

## 8.8 Problem: DeepLOB on a Laptop

**Problem 8.1.**

Weekend problem — the raw book against hand-made features

The chapter’s sessions: eight to train, three to stop, eight to test, eight more to grow the [training set](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets).

**Part I — The data.**

1. What is in one snapshot and one window?
2. How is the label built, and what are the class shares?
3. Why is accuracy a poor score here, and what does the chapter use instead?
4. Why must windows never cross the split between sessions?

**Part II — The models.**

5. What are the four sequence models’ receptive fields over the window?
6. What does each score (accuracy and IC)?
7. What does the [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) on hand-made features score?
8. How large is the spread across sessions and seeds?

**Part III — Data.**

9. After how many [epochs](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop) do the models stop, and why?
10. What does the TCN score with four, eight and sixteen training sessions?
11. Does the [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) gain from more sessions? Why not?
12. What would the curve need to do for the TCN to overtake the baseline?

**Part IV — The verdict.**

13. What did Kolm, Turiel and Westray find about raw books against order flow?
14. What did the LOBCAST benchmark find about published models on new data?
15. What made Sirignano and Cont’s raw-input model work?
16. State the *named result* : the best sequence model’s IC and accuracy against the [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) on order-flow imbalance, with the spread across seeds and sessions.
17. What would you put in production for a ten-second forecast, and what would you try next?
18. How would you turn the IC into a trading decision, given chapter 2’s labels?
19. What changes on Book 10’s simulator with several instruments?
20. In one sentence: what does a sequence model need to beat a good feature?

**Solution of Problem 8.1.**

1. A snapshot: for each of ten levels per side, the distance from the mid in ticks and the log size in lots (40 numbers); a window: the last 20 snapshots, ten seconds.
2. Mean mid over the next ten seconds minus the current mid, flat within a quarter tick. Test shares: 72.7% flat, 14.7% down, 12.5% up.
3. Predicting “flat” always is right 72.7% of the time; the chapter ranks models by the IC of $\P(\text{up}) - \P(\text{down})$ with the forward change.
4. Windows of one session overlap and share its state: across a split they leak.
5. CNN nine snapshots per filter stack, then an average over all 20; LSTM and [transformer](#def-ml-sequence-models-on-order-books-attention) all 20; TCN 31, more than the window.
6. CNN 0.727 and 0.060; LSTM 0.729 and 0.292; TCN 0.766 and 0.366; attention 0.775 and 0.426.
7. Accuracy 0.781, IC 0.543.
8. Standard deviations of 0.15 to 0.20 of IC for the sequence models, 0.09 for the [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) .
9. After one to four [epochs](https://one-course.com/books/quant/12/en/chapter/7-neural-networks-for-noisy-tabular-data#def-ml-neural-networks-for-noisy-tabular-data-backprop) : 8 808 windows are few for models reading 800 numbers each.
10. 0.13, 0.30 and 0.44.
11. Hardly (0.54 to 0.57): its ten features already summarise what the book says about the next seconds; it has few parameters and little variance to reduce.
12. Keep rising past 0.57 as sessions are added, which the doubling from eight to sixteen does not yet show.
13. That simpler networks on order-flow inputs outperform most models trained directly on order books.
14. That all fifteen published models lose a large part of their performance on new data.
15. Billions of quotes pooled across many stocks.
16. *Named result* : the best sequence model (attention) reaches IC 0.426 and accuracy 0.775 against the [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) ’s 0.543 and 0.781, with standard deviations across sessions and seeds of 0.15 against 0.09: a gain of $-0.12$ of IC, the gap narrowing as sessions are added.
17. The [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) (or boosted trees) on order-flow features; next, the TCN with the same features as extra channels, trained on many more sessions and instruments.
18. Through chapter 2: a meta-labelled threshold on $\P(\text{up}) - \P(\text{down})$ that pays the spread only when the expected move exceeds it, evaluated on trade P&L.
19. More sessions and instruments to pool, and market-by-order inputs; the same split rules.
20. Much more data than the feature needs, or the feature itself as an input.

## 8.9 Interview questions

**Interview question 8.1 ★ mle, researcher.**

What is a [causal convolution](#def-ml-sequence-models-on-order-books-tcn), and why does it matter for a trading model?

**Solution of Interview question 8.1.**

A convolution whose output at $t$ uses only inputs at $t$ and before. A trading model must not see the future of its own window; a centred convolution trained offline leaks the next steps into each prediction.

*What the interviewer is looking for: the definition and the look-ahead it prevents.*

**Interview question 8.2 ★★ mle.**

Explain self-attention in a few equations. What does it cost in the length of the sequence?

**Solution of Interview question 8.2.**

$\mathsf Q = ZA_{\mathsf Q}$, $\mathsf K = ZA_{\mathsf K}$, $\mathsf V = ZA_{\mathsf V}$, output $\mathrm{softmax}(\mathsf
Q\mathsf K^\top/\sqrt h)\mathsf V$: each position averages all positions’ values with weights from query–key similarity. Cost $O(T^2h)$ in time and $O(T^2)$ in memory for the weights.

*What the interviewer is looking for: the formula, the scaling, and the quadratic cost.*

**Interview question 8.3 ★★ researcher.**

Your order-book model is 80% accurate on a three-class mid-price label. What questions do you ask before you are impressed?

**Solution of Interview question 8.3.**

What the majority class share is; how the label is built (smoothing, horizon, dead band); how data were split (by day, purged); results by day and instrument, with spreads; an IC or P&L after costs; comparison with a linear model on order-flow features.

*What the interviewer is looking for: class balance, label construction, split, and a baseline.*

**Interview question 8.4 ★★ researcher, trader.**

Raw order-book levels or order-flow features as inputs? Argue from the evidence.

**Solution of Interview question 8.4.**

The evidence (Kolm, Turiel and Westray; this chapter’s learning curve) favours order-flow features at realistic data sizes: they are stationary and already summarise the book. Raw levels can win with very large pooled datasets (Sirignano and Cont). Start with the features; add raw inputs as extra channels.

*What the interviewer is looking for: data size as the deciding variable, with the published evidence.*

**Interview question 8.5 ★★ mle.**

How do you split order-book data for training and testing, and why?

**Solution of Interview question 8.5.**

By whole sessions (days), in time order, with the test period after training and validation, purged at the edges; never window by window at random; normalisation fitted on training days.

*What the interviewer is looking for: session-level splits and why windows leak.*

**Interview question 8.6 ★★★ researcher, mle.**

Why might an LSTM underperform a TCN on short windows of order-book data, and when would you expect the opposite?

**Solution of Interview question 8.6.**

On short windows the most recent snapshots carry the signal; a TCN reads them directly and trains in parallel, while an LSTM must carry them through its state and is slower to fit. An LSTM can do better when the relevant history is long and irregular, with state to accumulate (a slowly building queue, a metaorder).

*What the interviewer is looking for: recency, training dynamics, and when recurrence pays.*
