Quantitative Finance · Book 7 · Research

Research Craft: Predictors, Backtests, Measurement, Portfolios

Research Craft: Predictors, Backtests, Measurement, Portfolios · Research

18Order-Book Replay Simulation

A strategy that keeps one lot on the best bid and one on the best ask earns $1 106 an hour in a replay of the simulated order book, filling 61% of all the volume that trades. The same strategy on the same messages, placed where a real order would stand, behind the orders already resting at its price, fills 10% of the volume and loses $76 an hour. The first replay assumed that the strategy’s quotes stood first in every queue they joined. Level 3 of the firm’s backtests replays the market’s messages one by one, puts the strategy’s own orders in the queue, and makes their place in it, and the time it takes them to get there, explicit models. This chapter builds it (firm.lobreplay, with a C++20 core and a Rust twin that reproduce its fills), measures how much of a passive strategy’s result is queue position and latency, and bounds what a replay cannot see: the strategy’s own impact.

18.1 Replaying a market-by-order feed

Definition 18.1 (Order-book replay)

An order-book replay rebuilds the book from a recorded message feed and runs a strategy against it, inserting the strategy’s own orders as shadow orders: they can be filled by the recorded executions but do not change the messages that follow.

A market-by-order feed (Book 1, chapter 19) carries every order’s arrival, cancellation and execution with its identifier, so the replay knows the queue at each price in time priority. firm.lobreplay replays the messages of firm.tape: one simulated hour holds 129 303 of them. Its book is the queue of (order, size) pairs at each price; its clock is the messages’ timestamps, merged with the strategy’s own events (its orders’ arrivals and cancellations at the exchange, its view of the market) in one priority queue, as in chapter 17.

18.2 Where am I in the queue?

Definition 18.2 (Queue position, queue-position model)

The queue position of a resting order is the quantity ahead of it at its price, which must execute or cancel before it can fill. A queue-position model says how a shadow order’s queue position is set when it arrives and how it changes as the book changes.

A shadow order in the queue at its price. Executions consume the queue from the front; cancellations ahead of the order move it up, cancellations behind do not; it fills when an execution reaches it.
Figure 18.1. A shadow order in the queue at its price. Executions consume the queue from the front; cancellations ahead of the order move it up, cancellations behind do not; it fills when an execution reaches it.

firm.lobreplay implements three models:

  • front: the order is first at its price as soon as it arrives, and every execution at that price fills it (the touch fill of chapter 17);
  • fifo: the order joins behind the orders resting at its price when it arrives, known by identifier from the market-by-order feed; it moves up as those orders execute or cancel, and fills when an execution reaches an order that arrived after it;
  • prob: only the level’s total size is known (a market-by-price feed): executions consume the quantity ahead first, and a cancellation removes size ahead of the order in proportion to the share of the level that is ahead of it.

The open-source backtester hftbacktest, which replays market-by-price data, offers the same two families: a risk-averse model in which cancellations happen only at the tail of the queue, so the order advances only on trades, and probabilistic models in which decreases happen both before and after the order’s position, with the probability given by a function of its place in the queue.

18.3 Latency

Definition 18.3 (Latency model, order-entry latency, market-data latency)

A latency model gives the delays between the market and the strategy: the market-data latency, from an event at the exchange to the strategy seeing it, and the order-entry latency, from the strategy sending an order or cancellation to the exchange acting on it.

The chapter’s touch quoter keeps one lot on the best bid and one on the best ask, moving an order when its price is no longer the best, within five lots of inventory. Replayed on two simulated hours under each model, with order-entry latencies from zero to five seconds and a market-data latency of half of it (Figure 18.2):

order-entry latency050 ms200 ms1 s5 s
front: P&L an hour ($)1 106778672498192
front: share of traded volume61%48%43%28%11%
fifo: P&L an hour ($)−76-76−94-94−89-89−146-146−8-8
fifo: lots filled an hour522505470350182
fifo: mark-out after 10 s (ticks)+0.008+0.008−0.001-0.001−0.025-0.025−0.100-0.100−0.115-0.115
prob: P&L an hour ($)−49-49−42-42−41-41−26-26+33+33

The front model’s profits are an artefact: a quote that is first in every queue takes the benign executions that a real order behind the queue never sees, and its mark-out stays near +0.4+0.4 ticks at every latency. With queue positions, the quoter has no edge: its P&L is within its noise of zero, and its 10-second mark-out, a hundredth of a tick with no latency, turns negative by 50 milliseconds and reaches −0.1-0.1 ticks at one second: the orders that still fill when the quoter is slow are the ones the fast traders let through. Lehalle and Mounjid (2017) found the same erosion in a model of limit orders: the value of knowing the liquidity imbalance is eroded by latency when there is not enough time to cancel and reinsert. The probabilistic model, which knows only the level’s size, lands close to the exact one (499 lots against 522 with no latency).

The touch quoter replayed on two simulated hours of firm.tape under three queue-position models, against the order-entry latency (the market-data latency is half of it). Only the front-of-queue model, which no real order enjoys, makes money. Data: rs_lobreplay.touch_grid.
Figure 18.2. The touch quoter replayed on two simulated hours of firm.tape under three queue-position models, against the order-entry latency (the market-data latency is half of it). Only the front-of-queue model, which no real order enjoys, makes money. Data: rs_lobreplay.touch_grid.

18.4 Fill logic

Definition 18.4 (Passive fill probability)

The passive fill probability of a resting order over a horizon is the probability that the executions at its price consume the queue ahead of it and reach it within the horizon.

Under the fifo model a fill has one rule: an execution against an order that arrived after the shadow means, in time priority, that the aggressor reached the shadow first; the shadow fills by the smaller of its remaining size and the execution’s. Every other message at its price only moves it up (an execution or cancellation of an order ahead) or does nothing (an order joining behind, a cancellation behind). The rule is simple enough to write three times: in Python for research, in C++20 for the firm’s engines (Book 13) and in Rust, and the three must agree. The shared fixture is five simulated minutes (8 513 messages) and the 75 shadow orders a touch quoter sent with a 50-millisecond latency; the Python engine, the Python reference track_fifo, the C++20 core and the Rust twin all produce the same 29 fills, to the last bit of the timestamps.

The same engine gives chapter 17 its answer. The bar quoter of that chapter, replayed on its four days under the fifo model, fills 344 lots a day (47% of its orders) and loses $239 a day, with a five-minute mark-out of −0.77-0.77 ticks: between the touch fill’s 523 lots and +$17+\$17 and the penetration fill’s 290 lots and −$584-\$584, and nearer the pessimistic end. The penetration fill was the better level-2 guess; neither was the answer.

18.5 The missing market impact and how to bound it

Definition 18.5 (Market impact, counterfactual impact)

The market impact of a strategy is the change its own orders and trades cause in the prices it later trades at. In a replay it is a counterfactual impact: the recorded market did not contain the strategy’s orders, so what it would have done with them can only be modelled.

A replay leaves the recorded messages unchanged: when the shadow fills, the real order behind it still executes in the data, and the aggressor who in reality would have taken the shadow’s lot and then stopped (or moved the price further) is replayed as if the shadow were not there. hftbacktest’s documentation states the assumption plainly: the replayed order cannot change the simulated market, and must be small enough not to make any impact. The assumption can be bounded. Kyle’s lambda on this simulated tape is 0.10 ticks per 100 shares (chapter 9); if each lot the quoter trades moved the price against it by half that on average, the fifo quoter’s hour would cost about $26 more with no latency and $9 more at five seconds, small against its noise here and decisive for a strategy with an edge of a few tens of dollars. Where the bound matters, the answer is not a better replay but a reactive simulator, in which the other traders respond to the strategy’s orders (Book 10), or live trading at small size (chapter 21).

18.6 The fourth level

The four levels of chapter 16 now have their machinery: firm.vecbt for weights and costs, firm.evbt for orders on bars, firm.lobreplay for orders in the queue with latency. Each level has answered a question the one below could not: level 2 showed that the passive quoter’s result was the fill model’s; level 3 showed which fill model, and that its edge was an artefact of queue position. What no replay can supply is the market’s reaction to the strategy and the strategy’s real latency and fills: that is level 4, trading live at small size or on paper against live data, and chapters 19 and 21 measure how far the three simulations are from it.

18.7 Tutorial: where did the queue go?

Goal. Replay a touch quoter under three queue-position models and five latencies, replay chapter 17’s bar quoter at level 3, and check the C++20 and Rust cores against the Python fills. End state: Figure 18.2; the table; three agreeing implementations.

  1. The queue models: what each message at the shadow’s price does to it.

        def on_message(self, book: Book, m, shadows) -> list:
            """Update the working shadows at the message's price and side, before the book applies it. Returns fills
            [(vid, qty)] caused by an execution that, in time priority, would have reached the shadow."""
            kind, oid, side, px, qty = m["kind"], int(m["oid"]), int(m["side"]), int(m["price"]), int(m["qty"])
            out = []
            for s in shadows:
                if s.status != "working" or s.side != side or s.price != px:
                    continue
                if kind == b"A":
                    if self.model == "prob":
                        s.behind += qty
                    continue
                if self.model == "fifo":
                    if oid in s.ahead_ids:
                        take = min(qty, s.ahead_ids[oid])
                        s.ahead_ids[oid] -= take
                        s.ahead -= take
                        if s.ahead_ids[oid] <= 0:
                            del s.ahead_ids[oid]
                    elif kind == b"E":
                        out.append((s, min(s.remaining, qty)))
                    continue
                if kind == b"E":
                    if self.model == "front":
                        out.append((s, min(s.remaining, qty)))
                        continue
                    take = min(qty, s.ahead)
                    s.ahead -= take
                    if qty - take > 0:
                        out.append((s, min(s.remaining, qty - take)))
                        s.behind = max(0.0, s.behind - max(0.0, qty - take - s.remaining))
                    continue
                if self.model == "prob":                          # a cancellation at our price
                    tot = s.ahead + s.behind
                    share = s.ahead / tot if tot > 0 else 0.0
                    s.ahead = max(0.0, s.ahead - qty * share)
                    s.behind = max(0.0, s.behind - qty * (1.0 - share))
            return out
    Listing 18.1. Queue positions under three models. code/firm/lobreplay/firm_lobreplay.py
  2. The C++20 core of the fifo model, reproducing the Python fills.

            std::vector<int> now;
            for (int v : order_seen) if (active.count(v)) now.push_back(v);
            for (int v : now) {
                Shadow& s = sh[v];
                if (s.o.side != m.side || s.o.price != m.price || m.kind == 'A') continue;
                auto it = s.ahead.find(m.oid);
                if (it != s.ahead.end()) {
                    it->second -= std::min(m.qty, it->second);
                    if (it->second <= 0) s.ahead.erase(it);
                } else if (m.kind == 'E') {
                    long q = std::min(s.o.qty - s.filled, m.qty);
                    if (q > 0) {
                        s.filled += q;
                        out.push_back({v, m.t, q});
                        if (s.filled >= s.o.qty) { s.status = 2; active.erase(v); }
                    }
                }
            }
            book.apply(m);
    Listing 18.2. The fifo model’s per-message step in C++20. code/firm/lobreplay/cpp/firm_lobreplay.hpp
  3. Run touch_grid(), bar_quoter_level3(), impact_bound(), fig_lobreplay.py and make_lobreplay_fixture.py; build the C++ test and cargo test the Rust crate.

What to change next. Quote only when the queue at the best is short (a queue-imbalance filter, chapter 8) and see whether an edge appears; replace the proportional cancellation rule of the prob model with a risk-averse one and compare its fills with the exact model’s.

18.8 Build: the level-3 replay engine

Purpose. The firm’s backtester for strategies whose fills depend on the queue and on latency: market making, passive execution, queue-imbalance signals; the reference the production engines’ fill logic is checked against.

Interface. Book, Shadow, QueueTracker(model), Replay(msgs, strategy, model, entry_latency, data_latency).run() returning shadows, fills, position, cash and the mid path; Strategy.on_market(ctx, t, snapshot), on_fill; track_fifo(msgs, orders); C++20 firm::track_fifo in cpp/firm_lobreplay.hpp; Rust firm_lobreplay::track_fifo.

Rules. Shadow orders never change the replayed messages; our events at a message’s time act before it; every queue assumption lives in the model; the three implementations agree on the fixture.

Acceptance tests. code/firm/lobreplay/tests/: the book by hand; the fifo queue moving up on executions and cancellations ahead and filling on an execution behind, with and without a cancellation; the three models and both latencies on a hand stream; the fixture’s 29 fills from the reference; the C++20 and Rust tests on the same fixture.

Stretch. Pro-rata matching; iceberg orders; a reactive book in which other traders respond (Book 10).

Sources and further reading

  • C.-A. Lehalle and O. Mounjid, “Limit order strategic placement with adverse selection risk and the role of latency”, Market Microstructure and Liquidity 3(1), 2017.
  • R. Cont and A. de Larrard, “Price dynamics in a Markovian limit order market”, SIAM Journal on Financial Mathematics 4(1), 2013.
  • hftbacktest (open-source backtester), documentation, “Order fill”: exchange models and queue models.

18.9 Exercises

Exercise 18.1 ★

A shadow buy joins a bid queue of 300, 200 and 250 shares. Then 300 execute, 100 of the 200-share order cancel, an order of 400 joins, and 450 execute. How much of the shadow’s 100 fills under the fifo model?

Solution

Solution of Exercise 18.1.

The shadow joins behind 750. The 300 executed leave 450 ahead; the cancellation of 100 leaves 350; the order of 400 joins behind. Of the 450 executed, 350 consume the two orders still ahead and the last 100 execute against the order that joined behind: the shadow fills all 100.

Exercise 18.2 ★

Under the prob model a shadow has 600 ahead and 200 behind when 400 shares cancel at its price. What is its new position?

Solution

Solution of Exercise 18.2.

The share ahead is 600/800=0.75600/800 = 0.75: 300 of the cancelled 400 come from ahead and 100 from behind. New position: 300 ahead, 100 behind.

Exercise 18.3 ★

The strategy sees an event 3 milliseconds after it happens and its orders reach the exchange 5 milliseconds after it sends them. How old is the book its order meets?

Solution

Solution of Exercise 18.3.

The order was decided on a book 3 milliseconds old and reaches the exchange 5 milliseconds later: it meets a book 8 milliseconds newer than the one it was priced on.

Exercise 18.4 ★★

Why does the front-of-queue model’s mark-out stay positive at every latency while the fifo model’s turns negative?

Solution

Solution of Exercise 18.4.

The front model fills the shadow on every execution at its price, including the many small executions after which the price turns back: benign fills. The fifo model fills it only after the queue ahead is consumed, which happens mostly when the flow is heavy and the price is about to move through: adverse fills. Latency makes it worse, because the orders that still fill late are the ones faster traders did not want.

Exercise 18.5 ★★

The fifo quoter trades 522 lots an hour. With Kyle’s lambda at 0.10 ticks per 100 shares, what impact bound does the chapter’s rule give, in dollars an hour?

Solution

Solution of Exercise 18.5.

522×100×0.10/2=2 610522 \times 100 \times 0.10/2 = 2\,610 ticks of a cent: $26 an hour.

Exercise 18.6 ★★

Why does a shadow order placed at a price with no resting orders (inside the spread) never fill in a replay, and what does that say about replaying strategies that improve the price?

Solution

Solution of Exercise 18.6.

Shadows fill only by recorded executions at their price, and a price with no resting orders has none: an aggressor that would have hit the better-priced shadow is recorded executing elsewhere. Strategies that improve the price, or take liquidity in size, need a model of the aggressors’ response (a reactive simulator) or live evidence; the replay understates their fills.

Exercise 18.7 ★★★

Coding. Run the touch quoter under the fifo model with the market-data latency set to zero and only the order-entry latency varying. Does the edge depend more on seeing late or on acting late?

Solution

Solution of Exercise 18.7.

rs_lobreplay.see_or_act(1.0): with only a one-second order-entry latency the quoter fills 376 lots an hour with a mark-out of −0.075-0.075 ticks and loses $133; with only a one-second market-data latency it fills 416 lots with −0.132-0.132 ticks and loses $173. Seeing late is worse here: a stale view keeps orders at prices the market has already left, and the fills that come are the ones that have become bad.

Exercise 18.8 ★★★

Find the flaw. “Our market-making replay fills our quotes whenever a trade prints at our price, and it shows a Sharpe ratio of 9.”

Solution

Solution of Exercise 18.8.

Filling on every print at the quote’s price is the front-of-queue model: it assumes the quote was first in every queue and ignores the orders ahead of it, taking every benign fill a real order would miss. On the simulated tape that model turned a strategy with no edge into $1 106 an hour. Replay with queue positions and latency, look at the mark-outs, and bound the impact.

18.10 Problem: Where Did the Queue Go?

Problem 18.1

Weekend problem — a passive strategy, level 3

The chapter’s touch quoter on two simulated hours, and chapter 17’s bar quoter on four simulated days, replayed through firm.lobreplay.

Part I — The replay.

  1. How many messages does a simulated hour hold, and what does a market-by-order feed give the queue models that a market-by-price feed does not?
  2. What does a shadow order change in the replay, and what does it not?
  3. In what order do the strategy’s events and a market message at the same time act?
  4. How were the three implementations of the fifo model checked?

Part II — Queue position.

  1. What do the front and fifo models give the touch quoter with no latency: P&L, share of traded volume, lots?
  2. What does the prob model give, and why is it close to fifo?
  3. Why is the front model’s profit an artefact?
  4. What are the mark-outs of the front and fifo models?

Part III — Latency.

  1. How do the fifo model’s lots and mark-out change from zero to five seconds?
  2. At what latency does the quoter’s edge vanish?
  3. What did Lehalle and Mounjid find about latency?
  4. What does the front model do as latency grows, and why?

Part IV — The verdict.

  1. State the named result: the P&L of the passive strategy under each queue model and latency, and the latency at which its edge vanishes.
  2. What is the level-3 answer for chapter 17’s bar quoter, and which level-2 fill model was closer?
  3. What does the impact bound add, and when would it decide?
  4. What does a replay fundamentally miss, and where does the firm go to find it?
  5. What would you change in the strategy before replaying it again?
  6. What would you log live to calibrate the queue model?
  7. Why write the fill rule in three languages?
  8. In one sentence: what does a queue-position model assume?
Solution

Solution of Problem 18.1.

  1. 129 303. The identifier of every resting order, so the orders ahead of a shadow are known exactly; market-by-price gives only each level’s total size, and the shadow’s place must be modelled.
  2. It can be filled by recorded executions and changes the strategy’s position and cash; it does not change any recorded message: no impact.
  3. The strategy’s events (its orders’ arrivals and cancellations, its view of the market) act before a market message at the same time.
  4. On a shared fixture (8 513 messages, 75 shadow orders): the Python engine, the Python reference, the C++20 core and the Rust twin all give the same 29 fills.
  5. Front: $1 106 an hour, 61% of traded volume, 3 195 lots. Fifo: −$76-\$76, 10%, 522 lots.
  6. −$49-\$49 an hour and 499 lots: executions consume the quantity ahead first in both, and proportional cancellation is a fair guess when cancellations fall anywhere in the queue, as they do in the simulator.
  7. No real order is first in every queue: the model takes the benign executions a real order behind the queue never gets.
  8. Front: +0.39+0.39 ticks after 10 seconds; fifo: +0.008+0.008.
  9. Lots from 522 to 182 an hour; mark-out from +0.008+0.008 to −0.115-0.115 ticks.
  10. The mark-out turns negative at 50 milliseconds (−0.001-0.001); the P&L is within its noise of zero at every latency.
  11. That the value of predicting liquidity-consuming flows is eroded by latency when there is not enough time to cancel and reinsert the order.
  12. Its profit falls from $1 106 to $192 an hour and its share of volume from 61% to 11%, but its mark-out stays near +0.4+0.4: it still takes only benign fills, just fewer.
  13. Named result. Front of queue: $1 106, 778, 672, 498 and 192 an hour at 0, 50 ms, 200 ms, 1 s and 5 s. Fifo: −$76-\$76, −94-94, −89-89, −146-146 and −8-8. Prob: −$49-\$49, −42-42, −41-41, −26-26 and +33+33. With queue positions the quoter has no edge at any latency, and its mark-out turns negative at 50 milliseconds.
  14. 344 lots a day (47% of its orders), −$239-\$239 a day, a five-minute mark-out of −0.77-0.77 ticks; the penetration fill was closer than the touch fill.
  15. About $26 an hour with no latency ($9 at five seconds): small against this strategy’s noise, but decisive for one whose edge is tens of dollars.
  16. The market’s reaction to the strategy (impact, and other traders’ responses) and its real latency and fills: a reactive simulator (Book 10) and live trading at small size (chapters 19 and 21).
  17. Quote only when the queue is short or the imbalance favourable (chapter 8), and cancel faster when the flow turns toxic (chapter 9).
  18. Each order’s send, acknowledgement and fill times, the exchange’s queue position reports where available, and the market’s executions at the order’s price, to compare fills with the model’s.
  19. So that research and production fill orders by the same rule, checked on one fixture, and a discrepancy is a bug rather than a debate.
  20. Where the strategy’s order stands among the orders at its price, and how that changes as the book does.

18.11 Interview questions

Interview question 18.1 ★ researcher, trader

Why do market-making backtests usually overstate profits?

Solution

Solution of Interview question 18.1.

They assume fills a real order would not get: first place in the queue (fills on every touch or print), no latency, no impact. The fills the model adds are the benign ones (the price turns back), so the backtest understates adverse selection; on the chapter’s tape, front-of-queue fills made $1 106 an hour of a strategy that loses with queue positions.

Interview question 18.2 ★★ developer, researcher

How would you estimate your order’s position in the queue from a market-by-price feed?

Solution

Solution of Interview question 18.2.

Start at the level’s total size when the order arrives; move it up by every execution at the price; on each size decrease without an execution, attribute part of it to the orders ahead with a probability that depends on the position (proportional to the share ahead, or a power or log function of it, as hftbacktest’s probabilistic models do); calibrate against live fills or a market-by-order feed.

Interview question 18.3 ★★ developer

Design an order-book replay engine that supports latency and shadow orders. What are its invariants?

Solution

Solution of Interview question 18.3.

A single event queue merging market messages and the strategy’s delayed events; a book rebuilt by order identifier; shadow orders with queue positions under a pluggable model; latency queues for data and orders. Invariants: recorded messages are never altered; nothing acts before its timestamp; the strategy’s events at a time act before the market message at that time; runs are deterministic; the fill rule is the same as production’s.

Interview question 18.4 ★★ researcher, trader

How does latency affect a passive strategy’s fills and their quality?

Solution

Solution of Interview question 18.4.

Fewer fills (orders arrive later in the queue and are cancelled later), and worse ones: the fills that remain are those faster traders let through, so mark-outs fall. On the chapter’s quoter the mark-out goes from +0.008+0.008 ticks with no latency to −0.1-0.1 at one second; seeing late costs more than acting late.

Interview question 18.5 ★★ researcher

What does a replay backtest miss about market impact, and how would you bound it?

Solution

Solution of Interview question 18.5.

The strategy’s orders and trades would have moved prices and changed other traders’ behaviour; a replay keeps the recorded messages. Bound it with an impact coefficient (Kyle’s lambda times the strategy’s traded size), check the strategy’s share of the volume at its prices, and confirm at small size live.

Interview question 18.6 ★★★ developer

How would you make sure a research replay in Python and a production engine in C++ fill orders identically?

Solution

Solution of Interview question 18.6.

Write the fill rule once as a specification, implement it in each language, and check them on shared fixtures (messages, orders with their times, expected fills) bit for bit in continuous integration, as the chapter’s Python, C++20 and Rust implementations are; add live fills to the fixtures as they arrive.

Terms defined in this chapter

See all 2333 terms in the glossary