Research Craft: Predictors, Backtests, Measurement, Portfolios · Research
18Order-Book Replay Simulation
A strategy that keeps one lot on the best bid and one on the best ask earns $1 106 an hour in a replay of the simulated order book, filling 61% of all the volume that trades. The same strategy on the same messages, placed where a real order would stand, behind the orders already resting at its price, fills 10% of the volume and loses $76 an hour. The first replay assumed that the strategy’s quotes stood first in every queue they joined. Level 3 of the firm’s backtests replays the market’s messages one by one, puts the strategy’s own orders in the queue, and makes their place in it, and the time it takes them to get there, explicit models. This chapter builds it (firm.lobreplay, with a C++20 core and a Rust twin that reproduce its fills), measures how much of a passive strategy’s result is queue position and latency, and bounds what a replay cannot see: the strategy’s own impact.
18.1 Replaying a market-by-order feed
Definition 18.1 (Order-book replay)
An order-book replay rebuilds the book from a recorded message feed and runs a strategy against it, inserting the strategy’s own orders as shadow orders: they can be filled by the recorded executions but do not change the messages that follow.
A market-by-order feed (Book 1, chapter 19) carries every order’s arrival, cancellation and execution with its identifier, so the replay knows the queue at each price in time priority. firm.lobreplay replays the messages of firm.tape: one simulated hour holds 129 303 of them. Its book is the queue of (order, size) pairs at each price; its clock is the messages’ timestamps, merged with the strategy’s own events (its orders’ arrivals and cancellations at the exchange, its view of the market) in one priority queue, as in chapter 17.
18.2 Where am I in the queue?
Definition 18.2 (Queue position, queue-position model)
The queue position of a resting order is the quantity ahead of it at its price, which must execute or cancel before it can fill. A queue-position model says how a shadow order’s queue position is set when it arrives and how it changes as the book changes.
firm.lobreplay implements three models:
- front: the order is first at its price as soon as it arrives, and every execution at that price fills it (the touch fill of chapter 17);
- fifo: the order joins behind the orders resting at its price when it arrives, known by identifier from the market-by-order feed; it moves up as those orders execute or cancel, and fills when an execution reaches an order that arrived after it;
- prob: only the level’s total size is known (a market-by-price feed): executions consume the quantity ahead first, and a cancellation removes size ahead of the order in proportion to the share of the level that is ahead of it.
The open-source backtester hftbacktest, which replays market-by-price data, offers the same two families: a risk-averse model in which cancellations happen only at the tail of the queue, so the order advances only on trades, and probabilistic models in which decreases happen both before and after the order’s position, with the probability given by a function of its place in the queue.
18.3 Latency
Definition 18.3 (Latency model, order-entry latency, market-data latency)
A latency model gives the delays between the market and the strategy: the market-data latency, from an event at the exchange to the strategy seeing it, and the order-entry latency, from the strategy sending an order or cancellation to the exchange acting on it.
The chapter’s touch quoter keeps one lot on the best bid and one on the best ask, moving an order when its price is no longer the best, within five lots of inventory. Replayed on two simulated hours under each model, with order-entry latencies from zero to five seconds and a market-data latency of half of it (Figure 18.2):
| order-entry latency | 0 | 50 ms | 200 ms | 1 s | 5 s |
|---|---|---|---|---|---|
| front: P&L an hour ($) | 1 106 | 778 | 672 | 498 | 192 |
| front: share of traded volume | 61% | 48% | 43% | 28% | 11% |
| fifo: P&L an hour ($) | |||||
| fifo: lots filled an hour | 522 | 505 | 470 | 350 | 182 |
| fifo: mark-out after 10 s (ticks) | |||||
| prob: P&L an hour ($) |
The front model’s profits are an artefact: a quote that is first in every queue takes the benign executions that a real order behind the queue never sees, and its mark-out stays near ticks at every latency. With queue positions, the quoter has no edge: its P&L is within its noise of zero, and its 10-second mark-out, a hundredth of a tick with no latency, turns negative by 50 milliseconds and reaches ticks at one second: the orders that still fill when the quoter is slow are the ones the fast traders let through. Lehalle and Mounjid (2017) found the same erosion in a model of limit orders: the value of knowing the liquidity imbalance is eroded by latency when there is not enough time to cancel and reinsert. The probabilistic model, which knows only the level’s size, lands close to the exact one (499 lots against 522 with no latency).
firm.tape under three queue-position models, against the order-entry latency (the market-data latency is half of it). Only the front-of-queue model, which no real order enjoys, makes money. Data: rs_lobreplay.touch_grid.18.4 Fill logic
Definition 18.4 (Passive fill probability)
The passive fill probability of a resting order over a horizon is the probability that the executions at its price consume the queue ahead of it and reach it within the horizon.
Under the fifo model a fill has one rule: an execution against an order that arrived after the shadow means, in time priority, that the aggressor reached the shadow first; the shadow fills by the smaller of its remaining size and the execution’s. Every other message at its price only moves it up (an execution or cancellation of an order ahead) or does nothing (an order joining behind, a cancellation behind). The rule is simple enough to write three times: in Python for research, in C++20 for the firm’s engines (Book 13) and in Rust, and the three must agree. The shared fixture is five simulated minutes (8 513 messages) and the 75 shadow orders a touch quoter sent with a 50-millisecond latency; the Python engine, the Python reference track_fifo, the C++20 core and the Rust twin all produce the same 29 fills, to the last bit of the timestamps.
The same engine gives chapter 17 its answer. The bar quoter of that chapter, replayed on its four days under the fifo model, fills 344 lots a day (47% of its orders) and loses $239 a day, with a five-minute mark-out of ticks: between the touch fill’s 523 lots and and the penetration fill’s 290 lots and , and nearer the pessimistic end. The penetration fill was the better level-2 guess; neither was the answer.
18.5 The missing market impact and how to bound it
Definition 18.5 (Market impact, counterfactual impact)
The market impact of a strategy is the change its own orders and trades cause in the prices it later trades at. In a replay it is a counterfactual impact: the recorded market did not contain the strategy’s orders, so what it would have done with them can only be modelled.
A replay leaves the recorded messages unchanged: when the shadow fills, the real order behind it still executes in the data, and the aggressor who in reality would have taken the shadow’s lot and then stopped (or moved the price further) is replayed as if the shadow were not there. hftbacktest’s documentation states the assumption plainly: the replayed order cannot change the simulated market, and must be small enough not to make any impact. The assumption can be bounded. Kyle’s lambda on this simulated tape is 0.10 ticks per 100 shares (chapter 9); if each lot the quoter trades moved the price against it by half that on average, the fifo quoter’s hour would cost about $26 more with no latency and $9 more at five seconds, small against its noise here and decisive for a strategy with an edge of a few tens of dollars. Where the bound matters, the answer is not a better replay but a reactive simulator, in which the other traders respond to the strategy’s orders (Book 10), or live trading at small size (chapter 21).
18.6 The fourth level
The four levels of chapter 16 now have their machinery: firm.vecbt for weights and costs, firm.evbt for orders on bars, firm.lobreplay for orders in the queue with latency. Each level has answered a question the one below could not: level 2 showed that the passive quoter’s result was the fill model’s; level 3 showed which fill model, and that its edge was an artefact of queue position. What no replay can supply is the market’s reaction to the strategy and the strategy’s real latency and fills: that is level 4, trading live at small size or on paper against live data, and chapters 19 and 21 measure how far the three simulations are from it.
18.7 Tutorial: where did the queue go?
Goal. Replay a touch quoter under three queue-position models and five latencies, replay chapter 17’s bar quoter at level 3, and check the C++20 and Rust cores against the Python fills. End state: Figure 18.2; the table; three agreeing implementations.
The queue models: what each message at the shadow’s price does to it.
def on_message(self, book: Book, m, shadows) -> list: """Update the working shadows at the message's price and side, before the book applies it. Returns fills [(vid, qty)] caused by an execution that, in time priority, would have reached the shadow.""" kind, oid, side, px, qty = m["kind"], int(m["oid"]), int(m["side"]), int(m["price"]), int(m["qty"]) out = [] for s in shadows: if s.status != "working" or s.side != side or s.price != px: continue if kind == b"A": if self.model == "prob": s.behind += qty continue if self.model == "fifo": if oid in s.ahead_ids: take = min(qty, s.ahead_ids[oid]) s.ahead_ids[oid] -= take s.ahead -= take if s.ahead_ids[oid] <= 0: del s.ahead_ids[oid] elif kind == b"E": out.append((s, min(s.remaining, qty))) continue if kind == b"E": if self.model == "front": out.append((s, min(s.remaining, qty))) continue take = min(qty, s.ahead) s.ahead -= take if qty - take > 0: out.append((s, min(s.remaining, qty - take))) s.behind = max(0.0, s.behind - max(0.0, qty - take - s.remaining)) continue if self.model == "prob": # a cancellation at our price tot = s.ahead + s.behind share = s.ahead / tot if tot > 0 else 0.0 s.ahead = max(0.0, s.ahead - qty * share) s.behind = max(0.0, s.behind - qty * (1.0 - share)) return outListing 18.1. Queue positions under three models. code/firm/lobreplay/firm_lobreplay.py The C++20 core of the fifo model, reproducing the Python fills.
std::vector<int> now; for (int v : order_seen) if (active.count(v)) now.push_back(v); for (int v : now) { Shadow& s = sh[v]; if (s.o.side != m.side || s.o.price != m.price || m.kind == 'A') continue; auto it = s.ahead.find(m.oid); if (it != s.ahead.end()) { it->second -= std::min(m.qty, it->second); if (it->second <= 0) s.ahead.erase(it); } else if (m.kind == 'E') { long q = std::min(s.o.qty - s.filled, m.qty); if (q > 0) { s.filled += q; out.push_back({v, m.t, q}); if (s.filled >= s.o.qty) { s.status = 2; active.erase(v); } } } } book.apply(m);Listing 18.2. The fifo model’s per-message step in C++20. code/firm/lobreplay/cpp/firm_lobreplay.hpp - Run
touch_grid(),bar_quoter_level3(),impact_bound(),fig_lobreplay.pyandmake_lobreplay_fixture.py; build the C++ test andcargo testthe Rust crate.
What to change next. Quote only when the queue at the best is short (a queue-imbalance filter, chapter 8) and see whether an edge appears; replace the proportional cancellation rule of the prob model with a risk-averse one and compare its fills with the exact model’s.
18.8 Build: the level-3 replay engine
Purpose. The firm’s backtester for strategies whose fills depend on the queue and on latency: market making, passive execution, queue-imbalance signals; the reference the production engines’ fill logic is checked against.
Interface. Book, Shadow, QueueTracker(model), Replay(msgs, strategy, model, entry_latency, data_latency).run() returning shadows, fills, position, cash and the mid path; Strategy.on_market(ctx, t, snapshot), on_fill; track_fifo(msgs, orders); C++20 firm::track_fifo in cpp/firm_lobreplay.hpp; Rust firm_lobreplay::track_fifo.
Rules. Shadow orders never change the replayed messages; our events at a message’s time act before it; every queue assumption lives in the model; the three implementations agree on the fixture.
Acceptance tests. code/firm/lobreplay/tests/: the book by hand; the fifo queue moving up on executions and cancellations ahead and filling on an execution behind, with and without a cancellation; the three models and both latencies on a hand stream; the fixture’s 29 fills from the reference; the C++20 and Rust tests on the same fixture.
Stretch. Pro-rata matching; iceberg orders; a reactive book in which other traders respond (Book 10).
Sources and further reading
- C.-A. Lehalle and O. Mounjid, “Limit order strategic placement with adverse selection risk and the role of latency”, Market Microstructure and Liquidity 3(1), 2017.
- R. Cont and A. de Larrard, “Price dynamics in a Markovian limit order market”, SIAM Journal on Financial Mathematics 4(1), 2013.
- hftbacktest (open-source backtester), documentation, “Order fill”: exchange models and queue models.
18.9 Exercises
Exercise 18.1 ★
A shadow buy joins a bid queue of 300, 200 and 250 shares. Then 300 execute, 100 of the 200-share order cancel, an order of 400 joins, and 450 execute. How much of the shadow’s 100 fills under the fifo model?
Solution
Solution of Exercise 18.1.
The shadow joins behind 750. The 300 executed leave 450 ahead; the cancellation of 100 leaves 350; the order of 400 joins behind. Of the 450 executed, 350 consume the two orders still ahead and the last 100 execute against the order that joined behind: the shadow fills all 100.
Exercise 18.2 ★
Under the prob model a shadow has 600 ahead and 200 behind when 400 shares cancel at its price. What is its new position?
Solution
Solution of Exercise 18.2.
The share ahead is : 300 of the cancelled 400 come from ahead and 100 from behind. New position: 300 ahead, 100 behind.
Exercise 18.3 ★
The strategy sees an event 3 milliseconds after it happens and its orders reach the exchange 5 milliseconds after it sends them. How old is the book its order meets?
Solution
Solution of Exercise 18.3.
The order was decided on a book 3 milliseconds old and reaches the exchange 5 milliseconds later: it meets a book 8 milliseconds newer than the one it was priced on.
Exercise 18.4 ★★
Why does the front-of-queue model’s mark-out stay positive at every latency while the fifo model’s turns negative?
Solution
Solution of Exercise 18.4.
The front model fills the shadow on every execution at its price, including the many small executions after which the price turns back: benign fills. The fifo model fills it only after the queue ahead is consumed, which happens mostly when the flow is heavy and the price is about to move through: adverse fills. Latency makes it worse, because the orders that still fill late are the ones faster traders did not want.
Exercise 18.5 ★★
The fifo quoter trades 522 lots an hour. With Kyle’s lambda at 0.10 ticks per 100 shares, what impact bound does the chapter’s rule give, in dollars an hour?
Solution
Solution of Exercise 18.5.
ticks of a cent: $26 an hour.
Exercise 18.6 ★★
Why does a shadow order placed at a price with no resting orders (inside the spread) never fill in a replay, and what does that say about replaying strategies that improve the price?
Solution
Solution of Exercise 18.6.
Shadows fill only by recorded executions at their price, and a price with no resting orders has none: an aggressor that would have hit the better-priced shadow is recorded executing elsewhere. Strategies that improve the price, or take liquidity in size, need a model of the aggressors’ response (a reactive simulator) or live evidence; the replay understates their fills.
Exercise 18.7 ★★★
Coding. Run the touch quoter under the fifo model with the market-data latency set to zero and only the order-entry latency varying. Does the edge depend more on seeing late or on acting late?
Solution
Solution of Exercise 18.7.
rs_lobreplay.see_or_act(1.0): with only a one-second order-entry latency the quoter fills 376 lots an hour with a mark-out of ticks and loses $133; with only a one-second market-data latency it fills 416 lots with ticks and loses $173. Seeing late is worse here: a stale view keeps orders at prices the market has already left, and the fills that come are the ones that have become bad.
Exercise 18.8 ★★★
Find the flaw. “Our market-making replay fills our quotes whenever a trade prints at our price, and it shows a Sharpe ratio of 9.”
Solution
Solution of Exercise 18.8.
Filling on every print at the quote’s price is the front-of-queue model: it assumes the quote was first in every queue and ignores the orders ahead of it, taking every benign fill a real order would miss. On the simulated tape that model turned a strategy with no edge into $1 106 an hour. Replay with queue positions and latency, look at the mark-outs, and bound the impact.
18.10 Problem: Where Did the Queue Go?
Problem 18.1
Weekend problem — a passive strategy, level 3
The chapter’s touch quoter on two simulated hours, and chapter 17’s bar quoter on four simulated days, replayed through firm.lobreplay.
Part I — The replay.
- How many messages does a simulated hour hold, and what does a market-by-order feed give the queue models that a market-by-price feed does not?
- What does a shadow order change in the replay, and what does it not?
- In what order do the strategy’s events and a market message at the same time act?
- How were the three implementations of the fifo model checked?
Part II — Queue position.
- What do the front and fifo models give the touch quoter with no latency: P&L, share of traded volume, lots?
- What does the prob model give, and why is it close to fifo?
- Why is the front model’s profit an artefact?
- What are the mark-outs of the front and fifo models?
Part III — Latency.
- How do the fifo model’s lots and mark-out change from zero to five seconds?
- At what latency does the quoter’s edge vanish?
- What did Lehalle and Mounjid find about latency?
- What does the front model do as latency grows, and why?
Part IV — The verdict.
- State the named result: the P&L of the passive strategy under each queue model and latency, and the latency at which its edge vanishes.
- What is the level-3 answer for chapter 17’s bar quoter, and which level-2 fill model was closer?
- What does the impact bound add, and when would it decide?
- What does a replay fundamentally miss, and where does the firm go to find it?
- What would you change in the strategy before replaying it again?
- What would you log live to calibrate the queue model?
- Why write the fill rule in three languages?
- In one sentence: what does a queue-position model assume?
Solution
Solution of Problem 18.1.
- 129 303. The identifier of every resting order, so the orders ahead of a shadow are known exactly; market-by-price gives only each level’s total size, and the shadow’s place must be modelled.
- It can be filled by recorded executions and changes the strategy’s position and cash; it does not change any recorded message: no impact.
- The strategy’s events (its orders’ arrivals and cancellations, its view of the market) act before a market message at the same time.
- On a shared fixture (8 513 messages, 75 shadow orders): the Python engine, the Python reference, the C++20 core and the Rust twin all give the same 29 fills.
- Front: $1 106 an hour, 61% of traded volume, 3 195 lots. Fifo: , 10%, 522 lots.
- an hour and 499 lots: executions consume the quantity ahead first in both, and proportional cancellation is a fair guess when cancellations fall anywhere in the queue, as they do in the simulator.
- No real order is first in every queue: the model takes the benign executions a real order behind the queue never gets.
- Front: ticks after 10 seconds; fifo: .
- Lots from 522 to 182 an hour; mark-out from to ticks.
- The mark-out turns negative at 50 milliseconds (); the P&L is within its noise of zero at every latency.
- That the value of predicting liquidity-consuming flows is eroded by latency when there is not enough time to cancel and reinsert the order.
- Its profit falls from $1 106 to $192 an hour and its share of volume from 61% to 11%, but its mark-out stays near : it still takes only benign fills, just fewer.
- Named result. Front of queue: $1 106, 778, 672, 498 and 192 an hour at 0, 50 ms, 200 ms, 1 s and 5 s. Fifo: , , , and . Prob: , , , and . With queue positions the quoter has no edge at any latency, and its mark-out turns negative at 50 milliseconds.
- 344 lots a day (47% of its orders), a day, a five-minute mark-out of ticks; the penetration fill was closer than the touch fill.
- About $26 an hour with no latency ($9 at five seconds): small against this strategy’s noise, but decisive for one whose edge is tens of dollars.
- The market’s reaction to the strategy (impact, and other traders’ responses) and its real latency and fills: a reactive simulator (Book 10) and live trading at small size (chapters 19 and 21).
- Quote only when the queue is short or the imbalance favourable (chapter 8), and cancel faster when the flow turns toxic (chapter 9).
- Each order’s send, acknowledgement and fill times, the exchange’s queue position reports where available, and the market’s executions at the order’s price, to compare fills with the model’s.
- So that research and production fill orders by the same rule, checked on one fixture, and a discrepancy is a bug rather than a debate.
- Where the strategy’s order stands among the orders at its price, and how that changes as the book does.
18.11 Interview questions
Interview question 18.1 ★ researcher, trader
Why do market-making backtests usually overstate profits?
Solution
Solution of Interview question 18.1.
They assume fills a real order would not get: first place in the queue (fills on every touch or print), no latency, no impact. The fills the model adds are the benign ones (the price turns back), so the backtest understates adverse selection; on the chapter’s tape, front-of-queue fills made $1 106 an hour of a strategy that loses with queue positions.
Interview question 18.2 ★★ developer, researcher
How would you estimate your order’s position in the queue from a market-by-price feed?
Solution
Solution of Interview question 18.2.
Start at the level’s total size when the order arrives; move it up by every execution at the price; on each size decrease without an execution, attribute part of it to the orders ahead with a probability that depends on the position (proportional to the share ahead, or a power or log function of it, as hftbacktest’s probabilistic models do); calibrate against live fills or a market-by-order feed.
Interview question 18.3 ★★ developer
Design an order-book replay engine that supports latency and shadow orders. What are its invariants?
Solution
Solution of Interview question 18.3.
A single event queue merging market messages and the strategy’s delayed events; a book rebuilt by order identifier; shadow orders with queue positions under a pluggable model; latency queues for data and orders. Invariants: recorded messages are never altered; nothing acts before its timestamp; the strategy’s events at a time act before the market message at that time; runs are deterministic; the fill rule is the same as production’s.
Interview question 18.4 ★★ researcher, trader
How does latency affect a passive strategy’s fills and their quality?
Solution
Solution of Interview question 18.4.
Fewer fills (orders arrive later in the queue and are cancelled later), and worse ones: the fills that remain are those faster traders let through, so mark-outs fall. On the chapter’s quoter the mark-out goes from ticks with no latency to at one second; seeing late costs more than acting late.
Interview question 18.5 ★★ researcher
What does a replay backtest miss about market impact, and how would you bound it?
Solution
Solution of Interview question 18.5.
The strategy’s orders and trades would have moved prices and changed other traders’ behaviour; a replay keeps the recorded messages. Bound it with an impact coefficient (Kyle’s lambda times the strategy’s traded size), check the strategy’s share of the volume at its prices, and confirm at small size live.
Interview question 18.6 ★★★ developer
How would you make sure a research replay in Python and a production engine in C++ fill orders identically?
Solution
Solution of Interview question 18.6.
Write the fill rule once as a specification, implement it in each language, and check them on shared fixtures (messages, orders with their times, expected fills) bit for bit in continuous integration, as the chapter’s Python, C++20 and Rust implementations are; add live fills to the fixtures as they arrive.