---
title: "Build: An Order-Book Model, End to End"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 29
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/29-build-an-order-book-model-end-to-end
---

# Chapter 29 — Build: An Order-Book Model, End to End

The book’s last model is small on purpose: a forecast of the next second of the mid price, trained on the simulator’s sessions, validated without leaks, compiled into a few hundred nanoseconds, and traded in the same simulator, where every microsecond it takes is a microsecond of stale decisions. Every step reuses a component built earlier in the book, and the whole run is one reproducible graph. The chapter ends on the question the title of its problem asks: how much is a microsecond worth to this model? The answer is not what the title suggests. The model earns $21.73 per twenty-minute session over six test sessions, four times what the same quoting earns without it; its edge is unchanged up to 50 milliseconds of decision time and falls to zero at about 106 milliseconds. In a market whose top of book changes four times a second, microseconds are worth nothing and a tenth of a second is worth everything; what the latency costs depends on how fast the market moves, not on how fast the code runs.

## 29.1 The specification

The model, written as the [handoff specification](https://one-course.com/books/quant/12/en/chapter/28-the-machine-learning-team#def-ml-the-machine-learning-team-handoff) of chapter 28 would have it:

- *Market.* Book 10’s exchange simulator, one venue and one instrument at $100 with a one-cent tick, default fees (a rebate of $0.002 a share for providing liquidity, $0.003 for taking it), Book 7’s `firm_tape` flow as background, the agent behind Book 10’s latency model (20 microseconds each way).
- *Decision.* At every change of the top of book the agent sees: compute ten features, forecast the mid change over the next second (in ticks), and act.
- *Action.* Quote one lot (100 shares) at the touch on each side its position allows (at most one lot long or short), except the side the forecast says is about to be run over: the bid is withdrawn when the forecast is below $-\theta$ , the ask when it is above $\theta$ . A quote whose price is still the touch is left alone, since its place in the queue is worth keeping. A position held 30 seconds is closed at the market.
- *Sessions.* Twenty minutes each, on disjoint seeds: eight for training, three for validation, six for testing.
- *Budget and monitoring.* The decision time is the compiled model’s measured latency; chapter 27’s input monitors run on the test sessions.

The strategy is a market maker that uses the model to avoid being picked off, the classic use of a short-horizon forecast by a liquidity provider (Cartea, Jaimungal and Penalva treat its optimal form). A taker version was tried first and abandoned: in this market the mid moves by a tick or more within a second about five times in a hundred, and crossing a one-tick spread and paying the taker fee twice costs 1.6 ticks a round trip, more than any forecast here is worth.

The whole run is one `firm.workflow` graph ([Listing 29.1](#lst-ml-lob-graph)): two data stages, purged cross-validation, the fit and its export, the TCN challenger, the choice of threshold on the validation sessions, and the test sessions at every latency. Each stage is keyed by its code, parameters, seed and inputs, so a change anywhere reruns exactly what depends on it.

## 29.2 Data, labels and validation

The research sessions use the same agent class as production, with trading switched off: it records its feed as the events of chapter 24’s [feature store](https://one-course.com/books/quant/12/en/chapter/24-data-and-feature-stores#def-ml-data-and-feature-stores-store) (each change of the top of book, with its order-flow imbalance increment, and each trade with its aggressor’s sign) and computes its features online at every book change ([Listing 29.2](#lst-ml-lob-onbook)). The ten features are the imbalance at the touch, the spread, the order-flow imbalance of Cont, Kukanov and Stoikov over one and five seconds, signed traded volume over one and five seconds, the number of trades over five seconds, the mid change over one and five seconds, and the number of book updates over one second. The same definitions computed offline over the recorded events equal the online values exactly at every decision (the acceptance test of `firm.lobmodel`); training and serving cannot drift apart because they are the same code.

The label is chapter 2’s [fixed-horizon label](https://one-course.com/books/quant/12/en/chapter/2-targets-labels-and-sample-weights#def-ml-targets-labels-and-sample-weights-fixed) in event time ([Listing 29.3](#lst-ml-lob-data)): the mid change, in ticks, from the last event known at the decision to the last one a second later, computed by `firm.labeling` on the per-event mid changes. Most labels are zero (89.4%): in a second the mid usually does not move. The eight training sessions give 51 924 decisions and the three validation sessions 33 457; activity differs a lot from session to session.

Validation respects time. Each label spans one second, so neighbouring rows share their future; chapter 3’s purged five-fold cross-validation (`firm.cvsplit`, with a one-second embargo) removes from each training fold the rows whose labels overlap the test fold. Three LightGBM configurations of 150 trees with 7, 15 and 31 leaves score mean fold information coefficients of 0.558, 0.555 and 0.548; the smallest wins. Refitted on all training sessions, it has an information coefficient of 0.625 in sample and 0.522 on the validation sessions.

The challenger is chapter 8’s [temporal convolutional network](https://one-course.com/books/quant/12/en/chapter/8-sequence-models-on-order-books#def-ml-sequence-models-on-order-books-tcn) (`firm.lobseq`, 12 867 parameters) on windows of the last 16 decisions’ features, trained on seven training sessions and stopped early on the eighth. Its validation information coefficient is 0.452. The forest wins on accuracy, and, as the next section measures, it is also the model that can be made fast; Kolm, Turiel and Westray found similarly that off-the-shelf networks on stationary order-flow inputs beat networks on raw order-book states.

## 29.3 The model and its compiled form

The chosen forest has 150 trees of six splits each. Exported to flat arrays by chapter 26’s `firm.mlinfer`, it reproduces LightGBM’s predictions bit for bit on the validation rows; so do the generated C++20 branches, the C++ loop over the arrays and the Rust kernel, on 200 shared test vectors ([Listing 29.5](#lst-ml-lob-cpp), [Listing 29.6](#lst-ml-lob-rust)). The agent itself calls a NumPy walk of the same arrays that adds the leaves in the same order ([Listing 29.4](#lst-ml-lob-predict)).

| serving path | median | 99th percentile | 99.9th percentile |
| --- | --- | --- | --- |
| LightGBM’s Python call (one row) | 263 $\mu$s | 560 $\mu$s | 869 $\mu$s |
| TCN forward pass in PyTorch (one window) | 166 $\mu$s | 460 $\mu$s | 718 $\mu$s |
| flat forest in Python and NumPy | 37 $\mu$s | 98 $\mu$s | 172 $\mu$s |
| flat forest, C++20 loop | 1.1 $\mu$s | 2.5 $\mu$s | 9.1 $\mu$s |
| flat forest, generated C++20 branches | 153 ns | 383 ns | 1.1 $\mu$s |

***Table 29.1.** Latency of one prediction, measured once on a laptop (Intel Core Ultra 7 155H, WSL2, a shared machine without isolated cores; the clock’s own overhead is about 15 nanoseconds). Data: `bench_lobmodel.py`, `measured_latency.csv`.*

[Table 29.1](#tab-ml-lob-latency) gives the budget’s other side. The generated branches answer in 153 nanoseconds at the median, about 1 700 times faster than LightGBM’s Python call; the TCN, less accurate, would also be a thousand times slower than the compiled forest. The trading runs use these measurements as the agent’s declared decision time: 150 nanoseconds for the compiled model, 263 microseconds for the library call.

## 29.4 Trading it in the simulator

**Definition 29.1 (Decision staleness).**

The *decision staleness* of an order is the time from the market event its decision is based on to the moment the order leaves the trader: the market-data latency, any wait while earlier decisions are still being computed, and the decision time itself. With decisions triggered at a rate $\lambda$ and each taking $L$, the wait stays bounded only while $\lambda L<1$.

The agent’s decision ([Listing 29.7](#lst-ml-lob-act)) runs at every book change it sees. Its threshold is chosen on the validation sessions at the compiled model’s latency: mean net P&L per session of $43.4, $53.9, $48.0 and $50.1 for $\theta$ of 0.05, 0.1, 0.15 and 0.2 ticks, so $\theta=0.1$. On the six test sessions, at 150 nanoseconds, the model earns $21.73 per session (standard error $3.90), positive in all six, from 34.3 entries per session; its net exchange fees are a credit of $6.73. Its entries’ mark-out one second later is 0.405 ticks in its favour. The same quoting without a model (never withdrawing a side) earns $5.47 (standard error $2.22) from 58.7 entries, with a one-second mark-out of 0.078 ticks: it is filled more often and mostly by orders that know better. The model’s value is in the fills it declines.

The test sessions also run chapter 27’s input monitor, each session against the training sessions. The largest [population stability index](https://one-course.com/books/quant/12/en/chapter/27-monitoring-and-retraining#def-ml-monitoring-and-retraining-psi) of a session ranges from 0.14 to 1.60, far above any rule of thumb, and in every session it comes from a feature that counts events (updates in the last second, trades in the last five), which moves with the session’s activity: the test sessions see between 2.8 and 6.4 book changes a second. The four price features stay between 0.02 and 0.13. The model is profitable in all six sessions. A monitor on count features needs its reference to span the range of activity the model will meet, or features normalised by activity, or it pages every quiet afternoon.

## 29.5 What the latency costs

The test sessions are traded again with the decision time set to each point of a grid from zero to 150 milliseconds ([Table 29.2](#tab-ml-lob-results), [Figure 29.1](#fig-ml-lob-pnl)). Everything else is identical: the sessions, the model, the threshold.

| decision time | median staleness | net P&L per session ($) | 1-s mark-out (ticks) | sessions positive |
| --- | --- | --- | --- | --- |
| 0 | 0.02 ms | 21.73 | 0.405 | 6 of 6 |
| 150 ns | 0.02 ms | 21.73 | 0.405 | 6 of 6 |
| 263 $\mu$s | 0.28 ms | 21.73 | 0.405 | 6 of 6 |
| 1 ms | 1.02 ms | 21.23 | 0.399 | 6 of 6 |
| 10 ms | 10.0 ms | 21.72 | 0.405 | 6 of 6 |
| 30 ms | 30.0 ms | 20.63 | 0.394 | 6 of 6 |
| 50 ms | 50.0 ms | 20.52 | 0.417 | 6 of 6 |
| 70 ms | 70.0 ms | 12.40 | 0.353 | 6 of 6 |
| 100 ms | 154 ms | 1.60 | 0.245 | 3 of 6 |
| 150 ms | 802 ms | $-12.27$ | 0.148 | 2 of 6 |
| no model | 0.02 ms | 5.47 | 0.078 |  |

***Table 29.2.** The model on six test sessions at each decision time: median [decision staleness](#def-ml-build-an-order-book-model-end-to-end-staleness), mean net P&L per twenty-minute session, mean one-second mark-out of entries, and the sessions with a positive P&L; the last row quotes without a model. Data: `ml_lobmodel.latency_table`.*

![Mean net P&L per twenty-minute test session against the decision time, from the compiled model’s 150 nanoseconds to 150 milliseconds (log scale; zero gives the same as 150 nanoseconds), with a band of one standard error over the six sessions; dashed: the same quoting without a model. Data: ml_lobmodel.latency_table.](https://one-course.com/images/onecourse/chapters/quant-12/ml-build-an-order-book-model-end-to-end/fig-b4ba7d9eafb2.svg)

***Figure 29.1.** Mean net P&L per twenty-minute test session against the decision time, from the compiled model’s 150 nanoseconds to 150 milliseconds (log scale; zero gives the same as 150 nanoseconds), with a band of one standard error over the six sessions; dashed: the same quoting without a model. Data: `ml_lobmodel.latency_table`.*

Up to 50 milliseconds nothing changes: the P&L stays within its standard error of $21.73, and the mark-out within 0.02 ticks of 0.405. Beyond, the edge goes quickly. At 70 milliseconds the P&L is $12.40; at 100 milliseconds it is $1.60, positive in only three sessions; at 150 milliseconds it is $-\$12.27$. By linear interpolation on the grid, the P&L falls to zero at about 106 milliseconds, and to what quoting without a model earns at about 89 milliseconds: past that, the model is worse than no model.

Two mechanisms produce the cliff, and the staleness column separates them. Up to 70 milliseconds the median staleness equals the decision time: each decision is late by exactly that much, and a quote withdrawn 50 milliseconds late is rarely hit in those 50 milliseconds, because the book changes only about four times a second. At 100 milliseconds the median staleness is 154 milliseconds: in the busiest bursts, decisions arrive faster than one every 100 milliseconds, $\lambda L$ passes one, and they queue behind each other. At 150 milliseconds the median staleness is 802 milliseconds, and the agent quotes on a book that is almost a second old. A model that is slow on average is slower still exactly when the market is busy, which is when its forecasts matter.

The one-second mark-out tells the same story from the fills’ side, falling from 0.405 to 0.148 ticks; it stays positive, because even a late withdrawal avoids some of the worst fills, but not by enough to pay for the rest.

Why do microseconds not matter here? Because the time that counts is measured in the market’s events, not in seconds. Kolm, Turiel and Westray found that the effective horizon of their stock-specific forecasts was about two average price changes. In this market the book changes about four times a second and no one else is racing for the same quotes: the background flow does not react to the agent. On a real exchange the top of a liquid stock changes far more often than four times a second, and the agent’s quotes are picked off by other fast traders: Aquilina, Budish and O’Neill, from exchange message data, find latency-arbitrage races about once a minute per FTSE 100 stock, with the modal race lasting 5 to 10 millionths of a second. The same cliff would sit correspondingly closer to zero, and the microseconds that bought nothing here would be the whole business. The simulator answers the question it was built to answer, the cost of staleness in its own event time; carrying the answer to a real market means rescaling it by that market’s event rate and adding its competitors.

**Method 29.2 (Building a short-horizon model that trades).**

1. Specify the market, the decision, the action and the budget before the model.
2. Record research data with the production agent, so features are the same code in both; test that offline equals online.
3. Validate with purged folds; choose the threshold on sessions the model was not fitted on; test on others.
4. Compile the chosen model and check parity on shared vectors; measure its latency and use the measurement.
5. Trade it at a grid of decision times, and report P&L, mark-outs and staleness against the time, next to a no-model baseline.
6. Read latency in the market’s event time: the edge’s horizon is a number of price changes, not of microseconds.

## 29.6 Tutorial: microseconds are money

**Goal.** Run the whole pipeline, compile the model, trade it at every decision time, and find where its edge ends. **End state:** [Table 29.1](#tab-ml-lob-latency), [Table 29.2](#tab-ml-lob-results), [Figure 29.1](#fig-ml-lob-pnl).

1. **The graph.** `def make_pipeline (cache_dir, cfg): """The whole chapter as one firm.workflow graph (cfg: seeds, seconds, grids, latencies).""" from firm_workflow import Pipeline, Stage return Pipeline([ Stage(" train " , stage_data, (), {" seeds " : cfg[" train " ], " seconds " : cfg[" seconds " ]}), Stage(" valid " , stage_data, (), {" seeds " : cfg[" valid " ], " seconds " : cfg[" seconds " ]}), Stage(" cv " , stage_cv, (" train " ,), {" grid " : cfg[" grid " ]}, 1 ), Stage(" fit " , stage_fit, (" train " , " valid " , " cv " ), {}, 1 ), Stage(" tcn " , stage_tcn, (" train " , " valid " ), {" T " : cfg[" T " ], " epochs " : cfg[" epochs " ]}, 1 ), Stage(" threshold " , stage_threshold, (" fit " ,), {" seeds " : cfg[" valid " ], " seconds " : cfg[" seconds " ], " grid " : cfg[" thresholds " ], " decision_ns " : cfg[" decision_ns " ]}), Stage(" test " , stage_test, (" fit " , " threshold " ), {" seeds " : cfg[" test " ], " seconds " : cfg[" seconds " ], " latencies " : cfg[" latencies " ]}), ], cache_dir)` **Listing 29.1.** The chapter as one firm.workflow graph. code/firm/lobmodel/firm_lobmodel.py
2. **Features at every book change.** `def on_book (self , ctx, locate, top): b, bq, a, aq = top if b is None or a is None or a <= b: return t = (ctx.now_ns - OPEN_NS) / SEC ofi = 0.0 if self .prev is not None : pb, pbq, pa, paq = self .prev ofi = ((bq if b >= pb else 0 ) - (pbq if b <= pb else 0 ) - (aq if a <= pa else 0 ) + (paq if a >= pa else 0 )) / 100.0 self .prev = top self ._event(t, (a + b) / (2 * TICK), (a - b) / TICK, (bq - aq) / (bq + aq), ofi, 0.0 , 0 ) if not self .trade: return x = self .engine.values(t) f = float (self .predict(x)) ctx.compute(self .decision_ns) self .decisions.append((t, f)) self .staleness.append(ctx.now_ns + self .decision_ns - self .last_ts) # exchange event to order departure self ._act(ctx, f, b, a)` **Listing 29.2.** The agent records an event, computes its features online, forecasts and declares its decision time. code/firm/lobmodel/firm_lobmodel.py
3. **Labels.** `def dataset (ev, times, horizon=HORIZON): """Features at the decision times (firm.featstore offline) and the fixed-horizon label of firm.labeling in event time: the mid change (ticks) from the last event known at the decision to the last one `horizon` seconds later, with each label's span (t0, t1) for purged validation.""" X = offline(ev, FEATURES, times) t = ev[" t " ] k0 = np.searchsorted(t, times, side=" right " ) - 1 k1 = np.searchsorted(t, times + horizon, side=" right " ) - 1 r = np.r_[0.0 , np.diff(ev[" mid " ])] # per-event mid changes y = fixed_horizon(r, k0, k1 - k0)[" ret " ] return X, y, times, times + horizon` **Listing 29.3.** Features offline and the fixed-horizon label in event time. code/firm/lobmodel/firm_lobmodel.py
4. **Serving in Python.** `class ForestPredictor : """firm.mlinfer's flat forest walked for all trees at once (NumPy over trees, a loop over depth); the leaf values are then added tree by tree in order, as forest_predict and the C++ and Rust kernels do, so that all four agree bit for bit.""" def __init__(self , f): self .f = {k: np.asarray(v) for k, v in f.items()} def __call__(self , x): f, n = self .f, self .f[" roots " ].astype(np.int64) while True : inner = n >= 0 if not inner.any(): break m = n[inner] go = np.asarray(x)[f[" feature " ][m]] <= f[" threshold " ][m] n[inner] = np.where(go, f[" left " ][m], f[" right " ][m]) s = 0.0 for v in f[" value " ][-n - 1 ].tolist(): s += v return s` **Listing 29.4.** The flat forest walked for all trees at once, leaves added in order. code/firm/lobmodel/firm_lobmodel.py
5. **Serving in C++20.** `// firm.lobmodel -- the C++20 serving path of the chapter 29 order-book model (Book 12): the flat forest of // firm.mlinfer and the generated branches, behind one call that takes the ten features in FEATURES order. # pragma once # include <array> # include <string> # include "../../mlinfer/cpp/mlinfer.hpp" # include "lobmodel_branches.hpp" namespace lobmodel { inline constexpr int kFeatures = 10 ; using Features = std::array<double , kFeatures>; struct Model { mlinfer::Forest forest; explicit Model(const std::string& path) : forest(mlinfer::read_forest(path)) {} double loop(const Features& x) const { return forest.predict(x.data()); } static double branches(const Features& x) { return lobmodel_branches(x.data()); } }; } // namespace lobmodel` **Listing 29.5.** The C++20 serving path: firm.mlinfer’s loop and the generated branches. code/firm/lobmodel/cpp/lobmodel.hpp
6. **Serving in Rust.** `use firm_mlinfer::Forest; pub const FEATURES: usize = 10 ; pub const FOREST: & str = include_str!(" ../../data/forest.txt " ); pub const VECTORS: & str = include_str!(" ../../data/vectors.csv " ); pub struct Model { forest: Forest , } impl Model { pub fn load () -> Model { Model { forest: Forest ::parse(FOREST) } } pub fn predict (&self , x: & [f64 ; FEATURES]) -> f64 { self .forest.predict(x) } }` **Listing 29.6.** The Rust serving path on firm.mlinfer’s kernel. code/firm/lobmodel/rust/src/lib.rs
7. **The decision.** `def _act (self , ctx, f, b, a): """Quote one lot at the touch on each side the position allows, except the side the forecast says is about to be run over (bid withdrawn when f <= -threshold, ask when f >= threshold); keep a quote whose price is still the touch (its queue place is worth keeping), move one that is not.""" pos = ctx.position(1 ) want = {" B " : b if pos < self .qty and f > -self .threshold else None , " S " : a if pos > -self .qty and f < self .threshold else None } have = {" B " : False , " S " : False } for cl, (side, px) in list (self .live.items()): if want[side] == px and not have[side]: have[side] = True continue ctx.cancel(cl) del self .live[cl] for side in (" B " , " S " ): if want[side] is not None and not have[side] and a - b == TICK: cl = ctx.send(Order(1 , side, self .qty, want[side], post_only=True )) self .live[cl] = (side, want[side])` **Listing 29.7.** Quote each allowed side at the touch unless the forecast says it is about to be run over. code/firm/lobmodel/firm_lobmodel.py
8. **Run** `ml_lobmodel.summary()` , `latency_table()` , `zero_crossing()` , `monitors()` , `fig_lobmodel.py` (about three minutes on one core), and `bench_lobmodel.py` once for the measured latencies.

**What to change next.** Add a second, faster quoting agent to the simulator that withdraws on the same signal: the cliff should move towards the difference between the two agents’ decision times.

## 29.7 Build: the order-book model

**Purpose.** A short-horizon model from simulated sessions to trading, with every step reproducible and its serving path in three languages.

**Interface.** `FEATURES`, `HORIZON`, `DEFAULT`; `ModelAgent(predict, threshold, hold_s, decision_ns, qty, trade)` (a `firm.exchsim` agent); `run_session`, `record`, `events_array`, `dataset`, `report`; `ForestPredictor`; the stages `stage_data`, `stage_cv`, `stage_fit`, `stage_tcn`, `stage_threshold`, `stage_test`; `make_pipeline(cache_dir, cfg)`; `cpp/lobmodel.hpp` and the Rust crate `firm_lobmodel`; `make_lobmodel_fixture.py`.

**Rules.** Research and production share the agent and the feature code; sessions of different roles never share a seed; the decision time is a measured latency; the serving paths agree bit for bit.

**Acceptance tests.** `code/firm/lobmodel/tests/`: online features equal offline ones at every decision and labels equal the mid change; the flat forest, the NumPy walk and LightGBM agree on the fixture, and the pipeline rebuilds the fixture’s forest; a longer decision time makes decisions staler. `cpp/lobmodel_test.cpp` and `cargo test`: exact parity on 200 vectors.

**Stretch.** A competing fast agent; features normalised by activity; the TCN compiled with chapter 26’s int8 path.

Sources and further reading

- P. N. Kolm, J. Turiel and N. Westray, “Deep order flow imbalance: extracting alpha at multiple horizons from the limit order book”, *Mathematical Finance* , 2023.
- J. Sirignano and R. Cont, “Universal features of price formation in financial markets: perspectives from deep learning”, *Quantitative Finance* , 2019.
- R. Cont, A. Kukanov and S. Stoikov, “The price impact of order book events”, *Journal of Financial Econometrics* , 2014.
- Á. Cartea, S. Jaimungal and J. Penalva, *Algorithmic and High-Frequency Trading* , Cambridge University Press, 2015.
- M. Aquilina, E. Budish and P. O’Neill, “Quantifying the high-frequency trading ‘arms race”’, *Quarterly Journal of Economics* , 2022.

## 29.8 Exercises

**Exercise 29.1 ★.**

Why does the agent record its research data itself, rather than the research using the simulator’s full tape?

**Solution of Exercise 29.1.**

Because production sees the market through the agent’s own feed, with its latency, and computes features from what it has received; research data recorded by the same agent with the same feature code are what production will see, and the offline and online values can be tested equal. A full tape seen from nowhere would train the model on information, and timing, that the agent never has.

**Exercise 29.2 ★.**

Why was a taker strategy abandoned? Compute its round-trip cost in ticks.

**Solution of Exercise 29.2.**

Buying at the ask and selling at the bid a second later costs the spread, one tick, plus the taker fee twice, $2\times\$0.003=\$0.006$ a share, 0.6 ticks at one cent: 1.6 ticks a round trip. The mid moves by a tick or more within a second in about 5% of cases, and the forecasts are much smaller than a tick; no threshold leaves trades that cover the cost.

**Exercise 29.3 ★.**

The model earns more than quoting without it, yet enters less often. Explain with the mark-outs.

**Solution of Exercise 29.3.**

Without a model the quotes are filled 58.7 times a session with a one-second mark-out of 0.078 ticks: most fills come from orders that move the price against the quote. The model withdraws the side about to be run over, so it is filled less (34.3 times) but its fills are worth 0.405 ticks a second later. P&L per session: $21.73 against $5.47.

**Exercise 29.4 ★★.**

At 100 milliseconds the median staleness is 154 milliseconds. Explain, with $\lambda L$, and say what a production system would do about it.

**Solution of Exercise 29.4.**

Decisions are triggered at every book change, about four a second on average but many more in bursts. With $L=100$ milliseconds, $\lambda L$ passes one in the bursts and decisions wait behind one another, so their median staleness (154 milliseconds) exceeds the decision time. A production system conflates: when it finishes a decision it computes the next from the latest book, skipping the updates it missed, so staleness is bounded by about twice the decision time; better still, it makes the decision fast enough that $\lambda L$ stays well below one in the busiest bursts.

**Exercise 29.5 ★★.**

The test sessions’ input monitor reads up to 1.60. Should the model have been stopped? What would you change in the monitor?

**Solution of Exercise 29.5.**

No: the model was profitable in all six sessions, and the monitor’s large values come from the features that count events, which move with the session’s activity (2.8 to 6.4 book changes a second), while the price features stay between 0.02 and 0.13. The monitor should use a reference spanning the range of activity (more training sessions, or history), normalise count features by the session’s activity, or watch counts with a separate activity monitor that pages the desk, not the [model owner](https://one-course.com/books/quant/12/en/chapter/28-the-machine-learning-team#def-ml-the-machine-learning-team-roles).

**Exercise 29.6 ★★.**

*Find the flaw.* “The compiled model is 1 700 times faster than LightGBM’s Python call, so it will make 1 700 times more money.”

**Solution of Exercise 29.6.**

Latency is worth money only where decisions compete with events or other traders. Here the P&L is identical at 150 nanoseconds and at 263 microseconds ($21.73), and unchanged within its standard error up to 50 milliseconds: the speed-up buys nothing in this market. It buys margin (a model that stays fast in bursts) and it buys everything in a market where competitors race in microseconds.

**Exercise 29.7 ★★★.**

Where would the zero-crossing sit for a stock whose top of book changes 2 000 times a second, if the edge’s horizon is a fixed number of book changes? What does the estimate leave out?

**Solution of Exercise 29.7.**

The test sessions change their top of book about 4.3 times a second, and the P&L crosses zero at about 106 milliseconds, some 0.45 book changes. At 2 000 changes a second the same number of changes takes $106\times4.3/2000\approx0.23$ milliseconds, about 230 microseconds. The estimate leaves out competition (other fast traders take the stale quotes first, which moves the cliff to the difference between their speed and the agent’s), queue position (faster markets have longer queues), the different model a busier market would need, and the burstiness that makes the decision queue grow.

**Exercise 29.8 ★★★.**

Write the monitoring specification for this model in production: the monitors, their references, thresholds and owners.

**Solution of Exercise 29.8.**

Inputs: the four price features against a reference spanning many sessions, with a threshold calibrated to one false page a month (chapter 27); the count features against the session’s activity, owned by the desk. Outputs: the share of decisions with a side withdrawn, the quoting uptime, the [decision staleness](#def-ml-build-an-order-book-model-end-to-end-staleness) distribution against a budget (median and 99th percentile, with an alarm when $\lambda L$ in bursts approaches one). Outcomes: the one-second mark-out of fills and the session P&L, with a CUSUM on the daily mark-out, owned by the [model owner](https://one-course.com/books/quant/12/en/chapter/28-the-machine-learning-team#def-ml-the-machine-learning-team-roles). Serving: parity of the C++ path with the reference on a daily sample of live inputs, owned by the [machine-learning engineer](https://one-course.com/books/quant/12/en/chapter/28-the-machine-learning-team#def-ml-the-machine-learning-team-roles).

## 29.9 Problem: Microseconds Are Money

**Problem 29.1.**

Weekend problem — microseconds are money

The chapter’s pipeline, model and latency grid.

**Part I — Specification and data.**

1. State the market, the decision, the action and the sessions.
2. What are the features, and why do offline and online values agree?
3. How is the label defined, and why are most labels zero?
4. How is the model validated, and what does the cross-validation choose?

**Part II — The model.**

5. How do the forest and the TCN compare?
6. What does the compiled forest cost per prediction, and against what?
7. How is parity checked?
8. How is the threshold chosen?

**Part III — Trading.**

9. What does the model earn at 150 nanoseconds, and against what baseline?
10. What do the mark-outs show?
11. What does the input monitor report, and why?
12. Define [decision staleness](#def-ml-build-an-order-book-model-end-to-end-staleness) and explain its two parts.

**Part IV — The verdict.**

13. State the *named result* : net P&L per session and the one-second mark-out as functions of decision latency, and the latency at which the model’s edge falls to zero.
14. Why is the curve flat up to 50 milliseconds?
15. Why does it fall so fast after?
16. What would change on a real exchange?
17. Should the firm spend on the C++ serving path for this market?
18. What would you add to the simulator before trusting the curve?
19. How would you register and monitor this model?
20. In one sentence: what is a microsecond worth?

**Solution of Problem 29.1.**

**Part I.**

1. One venue of Book 10’s simulator, one $100 stock with a one-cent tick, rebate $0.002 and taker fee $0.003 a share, `firm_tape` background, 20-microsecond latencies; at every book change, ten features and a one-second forecast; quote one lot at each allowed touch, withdrawing the side the forecast says will be run over ( $\theta=0.1$ ticks), closing positions after 30 seconds; eight training, three validation and six test sessions of twenty minutes.
2. Touch imbalance, spread, order-flow imbalance over 1 and 5 seconds, signed volume over 1 and 5 seconds, trades over 5 seconds, mid change over 1 and 5 seconds, updates over 1 second; offline and online agree because the same agent and the same [feature definitions](https://one-course.com/books/quant/12/en/chapter/24-data-and-feature-stores#def-ml-data-and-feature-stores-store) compute both.
3. The mid change in ticks over the next second, from the last event known at the decision; the mid usually does not move in a second (89.4% zeros).
4. Purged five-fold cross-validation with a one-second embargo; 7 leaves (0.558 against 0.555 and 0.548).

**Part II.**

1. Validation information coefficient 0.522 for the forest, 0.452 for the TCN.
2. 153 nanoseconds at the median for the generated branches, against 263 microseconds for LightGBM’s Python call and 166 for the TCN.
3. The flat arrays, the C++ loop and branches and the Rust kernel reproduce LightGBM’s predictions exactly on 200 shared vectors.
4. On the validation sessions at the compiled model’s latency: $53.9 per session at 0.1 ticks, the best of four.

**Part III.**

1. $21.73 per session (standard error $3.90), positive in all six; quoting without a model earns $5.47.
2. Model entries are worth 0.405 ticks one second later, against 0.078 without a model: the model declines the fills that lose.
3. Largest indices from 0.14 to 1.60, all from count features that follow the session’s activity; price features at most 0.13.
4. The time from the market event to the order’s departure: the decision time, plus the wait behind earlier decisions when they arrive faster than they can be made.

**Part IV.**

1. *Microseconds are money.* Per twenty-minute session, $21.73 from zero to 263 microseconds, $21.23 at 1 millisecond, $21.72 at 10, $20.63 at 30, $20.52 at 50, $12.40 at 70, $1.60 at 100 and $-\$12.27$ at 150; the one-second mark-out from 0.405 ticks to 0.148; the edge falls to zero at about 106 milliseconds, and below the no-model baseline at about 89.
2. Because the book changes about four times a second and nothing else races for the quotes: a withdrawal 50 milliseconds late is rarely too late.
3. Staleness grows by a decision’s own time and, in bursts, by the queue of decisions ( $\lambda L>1$ ), exactly when the market is busiest.
4. Far more book changes a second and fast competitors: the cliff would sit closer to zero by the ratio of event rates, and closer still where competitors race.
5. Not for this market’s sake; for robustness in bursts, and for the markets where it matters.
6. Competing fast agents, a reactive background ( `firm.agentmkt` ), realistic event rates, and conflation.
7. As in chapter 25 (the graph’s stage keys as the code version, the forest as the artefact) and chapter 27 (the specification of [Exercise 29.8](#exo-ml-build-an-order-book-model-end-to-end-8) ).
8. What the market’s event rate and its competitors say it is: here nothing, elsewhere everything.

## 29.10 Interview questions

**Interview question 29.1 ★ researcher, mle.**

What features would you use to forecast the next second of a stock’s mid price?

**Solution of Interview question 29.1.**

Order-flow imbalance at the touch over short windows, the touch’s size imbalance, the spread, signed traded volume, recent mid changes and activity; built as stationary quantities over windows in event or clock time, as Cont, Kukanov and Stoikov and Kolm, Turiel and Westray use.

*What the interviewer is looking for: order flow and imbalance, stationarity.*

**Interview question 29.2 ★★ researcher.**

How do you validate a model whose labels overlap in time?

**Solution of Interview question 29.2.**

Purge from the training folds every row whose label span overlaps the test fold, add an embargo after it, and keep sessions or days of different roles on disjoint data; choose thresholds on data the model was not fitted on.

*What the interviewer is looking for: purging, embargo and disjoint test data.*

**Interview question 29.3 ★★ mle, developer.**

How do you make sure a model computes the same features in production as in research?

**Solution of Interview question 29.3.**

One [feature definition](https://one-course.com/books/quant/12/en/chapter/24-data-and-feature-stores#def-ml-data-and-feature-stores-store) executed by both paths (a [feature store](https://one-course.com/books/quant/12/en/chapter/24-data-and-feature-stores#def-ml-data-and-feature-stores-store)), research data recorded through the production code path, and a test that offline values equal online ones at the same decision times, run in continuous integration and on samples in production.

*What the interviewer is looking for: shared code and a parity test.*

**Interview question 29.4 ★★ researcher, trader.**

A market maker adds a short-horizon forecast. Where does its value show up, and how do you measure it?

**Solution of Interview question 29.4.**

In the fills it declines: fewer fills with better mark-outs. Measure the mark-out curve of fills with and without the model and the P&L per session, on the same sessions.

*What the interviewer is looking for: adverse selection and mark-outs.*

**Interview question 29.5 ★★ mle, developer.**

Your model takes 200 microseconds in Python. How do you get it under a microsecond, and how do you know it still gives the same answers?

**Solution of Interview question 29.5.**

Export the model to a flat form (tree arrays, quantised weights), generate or write a native kernel with no allocation on the hot path, and check bit-for-bit or tolerance parity on shared test vectors; measure median and tail latency on the target machine.

*What the interviewer is looking for: compilation, parity vectors, measured tails.*

**Interview question 29.6 ★★★ researcher, developer.**

How would you decide how much to spend on latency for a given strategy?

**Solution of Interview question 29.6.**

Trade the strategy at a grid of decision times in a simulator with realistic event rates and competitors, find where its P&L bends and where it crosses zero, and spend until the strategy’s latency sits comfortably on the flat part, including in bursts.

*What the interviewer is looking for: a measured P&L curve against latency.*
