Microstructure and Execution · Execution
18Smart Order Routing
A router splits an order across five venues. Sent at once, the pieces arrive microseconds apart, and quotes on the later venues vanish in between: the first fill told fast traders what was coming. This chapter builds a router on the multi-venue simulator, measures what the order loses when its pieces arrive one after another and what it keeps when they are timed to arrive together, ranks venues by what they deliver rather than what they show, allocates a passive order across venues by Cont and Kukanov’s programme, and ends with the fees and the rules that protect an order from being read.
18.1 What a router decides
Definition 18.1 (Smart order router)
A smart order router splits a child order across venues and decides, for each piece, the venue, the price, the size, the order type and the moment it is sent, from its view of the venues’ quotes, fees, latencies and past fills.
Chapter 17 decided whether to rest or cross; the router decides where. Its inputs are the venues’ books (chapter 8’s fragmented market, protected quotes and the order protection rule of One Quant Book 1, chapter 9), their fees (chapter 7’s fee-adjusted price), its own latency to each venue (One Quant Book 7, chapter 18’s order-entry latency) and statistics of what past orders got. Foucault and Menkveld (2008) studied the entry of the London Stock Exchange against Euronext in Dutch stocks: the consolidated book was deeper after entry, and where more trades went through the entrant’s better quotes, it supplied less liquidity, which is why routers that actually reach every protected quote matter.
The experiment of this chapter uses five venues, A to E, in firm.exchsim. The router sits next to A and reaches the venues in 50, 120, 200, 280 and 360 microseconds. Each venue shows 500 shares at the ask of 100.01: 200 from slow providers and 300 from one fast provider, whose machine sits at A and reaches the other venues over faster links (45, 95, 145 and 195 microseconds). Behind them each venue has 1 000 shares at 100.02. The fast provider follows one rule: when one of its orders fills, it cancels its orders on every other venue. Take fees differ: 0.30 cents a share at A, B and D, 0.10 at C, and a rebate of 0.10 to the taker at E (an inverted venue, chapter 7).
18.2 Venue ranking and fill probabilities
Definition 18.2 (Venue ranking)
A venue ranking orders the venues by the expected cost of routing a share there, from the fill probability the router has measured on its own orders, the fee, and the cost of what does not fill.
A router that believes displayed quotes ranks venues by fee-adjusted price. One that keeps statistics learns what the display is worth. After 100 sprays (below) with a latency jitter of 10%, the router’s own orders filled 100% of the shares sent to A and 44%, 41%, 40% and 40% of those sent to B, C, D and E. With an unfilled share costing a tick later, the expected cost of a share sent to each venue ranks A first (0.30 cents), then E (0.56), C (0.63), B (0.69) and D (0.72): the displayed rebate of E is worth less than the certainty of A. Timed to arrive together, the same orders filled everywhere, and the ranking reverts to the fees: E (), C (0.10), then A, B and D (0.30). A venue ranking is a statement about the router’s own routing as much as about the venues.
def on_start(self, ctx):
self.names = {s.venue: s.name.split("@")[-1] for s in ctx.sessions}
d = send_delays([self.venues[v] for v, _, _ in self.legs], self.sync, self.margin)
for i, (v, _, _) in enumerate(self.legs):
ctx.set_timer(self.t0 + d[v] - ctx.now_ns, ("leg", i))
def on_timer(self, ctx, tag):
v, px, q = self.legs[tag[1]] if tag[0] == "leg" else tag[1:]
self.next_cl += 1 # ids are per session: unique across venues
cl = ctx.send(Order(side="B", qty=int(q), price=int(px), tif="I", venue=v,
cl_ord_id=self.next_cl))
self.cl_venue[cl] = v
if tag[0] == "leg":
self.pending.add(cl)
18.3 Aggressive routing and quote fade
Definition 18.3 (Spray order, quote fade)
A spray order sends marketable pieces to several venues at the same moment, so that they arrive in the order of the router’s latencies. Quote fade is the disappearance of displayed quotes between the moment a router sees them and the moment its orders arrive, typically because a liquidity provider cancels on seeing a fill elsewhere.
Definition 18.4 (Synchronised routing)
Synchronised routing delays each piece of a multi-venue order by the difference between its venue’s latency and the largest, so that all pieces arrive at their venues at about the same time and no fill can warn anyone before the last piece lands.
Figure 18.1 is the race. The spray reaches A at 50 microseconds and fills there; the fast provider learns of its fill 5 microseconds later and its cancels reach B at 100, C at 150, D at 200 and E at 250, each before the router’s piece. The router gets A’s 500 shares and the slow 200 on each other venue: 1 300 of 2 500, a fill ratio of 52%. The 1 200 shares left go to the next level in a clean-up sweep, and the whole order costs 0.73 cents a share above the displayed ask, fees included. Synchronised, every piece arrives at 360 microseconds: the fill ratio is 100% and the cost 0.18 cents, the average take fee.
mx_sor.race, from the latencies of the experiment. def on_report(self, ctx, rep):
if type(rep).__name__ != "Out_E" or self.faded:
return
self.faded = True
for o in ctx.working():
if o["cl"] != rep.cl_ord_id:
ctx.cancel(o["cl"], venue=o["venue"])
Synchronisation works only as well as the router knows its latencies. Figure 18.2 adds a random factor to each piece’s latency (lognormal, mean one), 100 trials per level. Up to 10% the synchronised sweep still fills everything; at 20% it fills 99.3% of the shares on average and the whole level on 94% of trials; at 40%, 90.3% and 37%. The spray gains from the noise, since it reshuffles the arrivals, but only to 68.6%. Royal Bank of Canada patented the method, in claims that describe timing parameters determined from latencies “so as to cause synchronized arrival” of the pieces at the venues (US patent 8 489 747, granted in 2013); its description names the problem as orders arriving at faster exchanges before slower ones, where others can detect them.
mx_sor.sweep_study.An order smaller than the displayed liquidity makes the router choose venues. For 1 000 shares (with 10% jitter), a router that picks the cheapest fees goes to E and C, the two farthest venues, and loses almost nothing even when it sprays (99.8%, 0.004 cents a share), because its first fill lands where the fast provider’s cancels are slow to reach. A router that picks the nearest venues goes to A and B and, spraying, fills 71.2% at 0.59 cents; synchronised, 100% at 0.30. Which piece arrives first matters more than which venue is cheapest.
18.4 Passive allocation across venues
A passive order faces the opposite problem: where to rest so as to be filled. Cont and Kukanov (2017) wrote it as a convex programme. With shares to buy, a market order and limit orders behind queues , and the shares that will leave the front of queue during the horizon (market orders and cancellations, random), the limit order at fills and the cost against the mid is
with the half-spread, the take fee, the rebate, the shares executed, and penalties , for executing too few or too many. Its expectation is convex when the penalties exceed the spread and fees; firm_sor.ck_allocate minimises the sample average over simulated outflows.
xi = np.atleast_2d(np.asarray(xi, float))
q, lim = np.asarray(queues, float), np.asarray(limits, float)
r = np.asarray(rebates, float)
g = np.zeros_like(r) if improve is None else np.asarray(improve, float)
fill = np.minimum(lim, np.maximum(xi - q, 0.0))
got = m + fill.sum(axis=1)
cost = ((h + take) * m - fill @ (h + r + g) + lam_u * np.maximum(x_total - got, 0)
+ lam_o * np.maximum(got - x_total, 0))
return cost, got
Take four lit venues and a dark one for a 1 000-share buy over a minute: queues of 1 500, 2 500, 300 and 800 shares at the lit bids and none in the dark pool; outflows with means of 2 000, 2 200, 900, 1 000 and 400 shares (lognormal, with a common factor); rebates of 0.20, 0.25, and 0.20 cents (the third venue is inverted); the dark venue fills at the mid, half a spread (0.5 cents) worse than a fill at the bid, without fee; a take fee of 0.30 cents, and penalties of 1.2 cents a share either way. The optimum is fitted on 400 scenarios and every rule is scored on 20 000 fresh ones (Table 18.1).
| rule | cost (cents a share) | filled | allocation (market; A, B, C, D, dark) |
|---|---|---|---|
| market order | 0.80 | 100.0% | 1 000; 0, 0, 0, 0, 0 |
| highest rebate | 0.77 | 22.0% | 0; 0, 1 000, 0, 0, 0 |
| shortest queue | 0.74 | 38.7% | 0; 0, 0, 0, 0, 1 000 |
| proportional to room | 0.20 | 61.3% | 0; 266, 171, 255, 140, 169 |
| Cont–Kukanov | 0.03 | 76.4% | 0; 405, 373, 406, 401, 193 |
| Cont–Kukanov, no dark venue | 0.06 | 69.7% | 0; 544, 478, 522, 478, 0 |
mx_sor.passive_study.The optimum posts 1 778 shares for 1 000 wanted: it overbooks, because each venue fills only part of its order and the penalty for too many is paid only on the scenarios where most venues fill. It sends nothing at the market. Chasing the highest rebate puts the whole order behind the longest queue and fills 22%; Battalio, Corwin and Jennings (2016) found in real data that limit orders routed to venues paying the largest rebates execute worse, and that brokers who route for rebates cannot have it all. The dark venue is worth 0.03 cents a share here: its fills are worse than a passive fill at the bid but come without a queue. Maglaras, Moallemi and Zheng (2021) showed that when many traders route this way, the queues of the venues move together, and the market behaves like one aggregate queue.
18.5 Fee-aware routing and anti-gaming
A fee-aware router compares prices net of fees (chapter 7): at equal displayed prices it takes the inverted venue first and rests on the venue with the best rebate. The two lessons above temper both: the cheapest venue to take from is worth nothing if its quote has faded by the time the order arrives, and the best rebate is worth nothing behind a queue that will not clear.
Definition 18.5 (Anti-gaming logic)
Anti-gaming logic is the part of a router that makes its orders hard to detect and exploit: synchronised arrivals, minimum fill quantities in dark venues, randomised child sizes and timing, avoiding venues whose fills are followed by adverse moves, and not routing to venues that let other participants learn about resting orders.
Chapter 9 measured how pinging orders find a large dark order and what the leak costs. A router answers with a dark-first leg that carries a minimum fill quantity, so that a 100-share ping cannot trade with it, and with child sizes drawn around their target instead of repeated (firm_sor.dark_first, child_sizes); it answers quote fade with synchronised arrivals; and it keeps, per venue, the markouts of its fills (One Quant Book 7, chapter 23) to stop routing to venues where they are consistently bad. Latency arbitrage, the business of the fast provider of this chapter, is One Quant Book 11’s subject.
18.6 Tutorial: a router on five venues
Goal. Sweep five simulated venues against a fading fast provider, with and without arrival-time alignment; rank the venues; allocate a passive order by Cont and Kukanov’s programme. End state: Figures 18.1 and 18.2, Table 18.1 and the numbers of sections 2 to 4.
- Venues.
mx_sor.VENUES,race();firm_sor.Venue,send_delays. - Sweep.
trial(sync, jitter, seed)withfirm_sor.RouterandFader;sweep_study(). - Ranking.
VenueStats.add,fill_ratio,rank;partial_study(). - Passive.
outflow(n, seed),ck_cost,ck_allocate;passive_study(); draw withfig_sor.py.
What to change next. Give the fast provider a second machine at C; let the router learn its latencies from acknowledgements; run the passive allocation on simulated venues instead of scenarios.
18.7 Build: smart order router
Purpose. The routing layer of chapter 28’s execution algorithm and of the desk of chapter 20: sweeps, passive allocation and venue statistics.
Interface. Venue(name, entry_ns, take, make, dark), send_delays(venues, sync, margin_ns), sweep(quotes, qty, fees), VenueStats (add, fill_ratio, rank), Router(legs, venues, t0_ns, sync, cleanup), Fader(price, qty, venues, start_ns), ck_cost, ck_allocate, child_sizes(qty, n, rng, spread), dark_first(qty, dark, min_qty).
Rules. Buys (sells mirror them); fees per share, positive when paid; client order identifiers unique across venues (the simulator numbers them per session); a clean-up sweep one tick higher after every leg has reported.
Acceptance tests. code/firm/sor/tests/: delays and fee ordering; venue statistics and ranking; on two simulated venues the fast provider’s quote fades before a spray and not before a synchronised sweep, and the clean-up completes the order; the Cont–Kukanov cost on a hand example and an optimum no worse than simple allocations; anti-gaming sizes.
Stretch. Learned latencies; a routing table by time of day; per-venue markouts in the ranking.
Sources and further reading
- T. Foucault and A. J. Menkveld, “Competition for order flow and smart order routing systems”, Journal of Finance 63(1), 2008.
- Royal Bank of Canada, “Synchronized processing of data by networked computing resources”, US patent 8 489 747 B2, 2013.
- R. Battalio, S. A. Corwin and R. Jennings, “Can brokers have it all? On the relation between make-take fees and limit order execution quality”, Journal of Finance 71(5), 2016.
- R. Cont and A. Kukanov, “Optimal order placement in limit order markets”, Quantitative Finance 17(1), 2017.
- C. Maglaras, C. C. Moallemi and H. Zheng, “Queueing dynamics and state space collapse in fragmented limit order book markets”, Operations Research 69(4), 2021.
18.8 Exercises
Exercise 18.1 ★
With the experiment’s latencies, how long must the router hold the piece for A in a synchronised sweep, and the piece for C?
Solution
Solution of Exercise 18.1.
The farthest venue is 360 microseconds away: A’s piece waits microseconds, C’s .
Exercise 18.2 ★
The fast provider’s machine moves to C. Its links reach A, B, D and E in 95, 50, 50 and 100 microseconds, and it hears of its fills at C in 5 microseconds. Which of its quotes does a spray from the router still find?
Solution
Solution of Exercise 18.2.
The spray still fills at A first (50 microseconds). The provider hears of it 95 microseconds later, at 145; its cancels reach C at 150, B at 195, D at 195 and E at 245. The spray reaches B at 120, before the cancel, and C, D and E at 200, 280 and 360, after it: it finds the fast quotes at A and B, 1 600 of the 2 500 shares (64%).
Exercise 18.3 ★
Compute the fee-adjusted cost of the spray: 1 300 shares at 100.01 (500 at A, 200 at each other venue) and a clean-up of 1 200 at 100.02 (1 000 at A and 200 at B), with the fees of the experiment.
Solution
Solution of Exercise 18.3.
At the ask, 1 300 shares pay only fees: cents. The clean-up pays a cent more on 1 200 shares (1 200 cents) and fees of cents. Above 100.01: cents a share.
Exercise 18.4 ★★
Why does the spray’s fill ratio rise with latency jitter?
Solution
Solution of Exercise 18.4.
The race is lost because A’s fill always comes first and the cancels are faster than the router’s later pieces. Random latencies sometimes make a far piece arrive before A’s fill has been reported, so it finds the fast quote; they help the spray exactly as they hurt the synchronised sweep.
Exercise 18.5 ★★
Why does the Cont–Kukanov optimum post more shares than it wants, and what stops it from posting many more?
Solution
Solution of Exercise 18.5.
Each venue fills only part of its order on most scenarios, so posting more than needed raises the expected shares executed toward . The over-fill penalty (1.2 cents a share, more than the 0.7 or so each passive fill earns) makes the scenarios in which most venues fill costly, which bounds the overbooking (1 778 shares posted).
Exercise 18.6 ★★
A router ranks venues by their displayed rebate. What does the passive study say it will get, and why?
Solution
Solution of Exercise 18.6.
It rests everything at the venue with the 0.25-cent rebate, which has the longest queue (2 500 shares): only 22.0% fills, and the penalty for the rest makes it cost 0.77 cents a share, almost as much as a market order.
Exercise 18.7 ★★★
Coding. Rerun the passive study with an over-fill penalty of 0.5 cents instead of 1.2. What does the optimum post, and what does it cost out of sample?
Solution
Solution of Exercise 18.7.
Posting pays: each passive fill earns more than the 0.5-cent over-fill penalty. The optimum posts about 4 030 shares (1 000 at A, B and D, 751 at C, 277 in the dark venue) and “costs” cents a share out of sample, a gain that exists only because the penalty undervalues the risk of buying too much.
Exercise 18.8 ★★★
Find the flaw. “Venue E showed 500 shares and we got 200: E’s quotes are fake, stop routing there.”
Solution
Solution of Exercise 18.8.
E’s 500 shares were there when the router looked; 300 of them belonged to a provider who cancelled after the router’s fill at A reached it. The router’s own spray caused the loss. Synchronised, E filled 100% and, with the taker rebate, became the cheapest venue.
18.9 Problem: The Order That Arrived After the Quotes Had Gone
Problem 18.1
Weekend problem — the order that arrived after the quotes had gone
A desk’s sweeps fill half of what the consolidated book shows. It asks why, and what a better router would change.
Part I — The router’s problem.
- Define a smart order router and list what it decides.
- Describe the five venues, their latencies and fees, and the fast provider.
- Define quote fade and a spray order.
- Why did Foucault and Menkveld’s evidence make routing to every protected quote matter?
Part II — The sweep.
- Draw the race of the spray at each venue.
- What fill ratio does the spray get, and what does the whole order cost?
- Define synchronised routing and compute the delays.
- State the named result: the fill ratio of the sweep with and without arrival-time alignment, and its fee-adjusted cost.
- How does latency jitter change both?
Part III — Ranking and choosing venues.
- Give the router’s measured fill ratios by venue after spraying.
- Rank the venues with and without synchronisation, and explain the difference.
- For 1 000 shares, compare routing by fee and by distance, sprayed and synchronised.
- Why did the fee-first spray lose almost nothing?
Part IV — Passive allocation and the answer.
- Write the Cont–Kukanov cost.
- Give the cost and fill of the market order, the highest rebate, the shortest queue, the proportional rule and the optimum.
- What is the dark venue worth here?
- What did Battalio, Corwin and Jennings find about rebates and execution quality?
- Name three anti-gaming rules and the attack each answers.
- What would you tell the desk?
- In one sentence: what does a router route around?
Solution
Solution of Problem 18.1.
1. It splits an order across venues and chooses venue, price, size, type and timing of each piece. 2. A to E at 50, 120, 200, 280 and 360 microseconds; take fees 0.30, 0.30, 0.10, 0.30 and cents; each shows 200 slow and 300 fast shares at 100.01 and 1 000 at 100.02; the fast provider sits at A with links of 45 to 195 microseconds and cancels everywhere after a fill. 3. See the definitions. 4. The consolidated book was deeper after the entrant arrived, and it supplied less liquidity where its better quotes were traded through. 5. Spray arrives at 50, 120, 200, 280, 360; cancels at 100, 150, 200 and 250 at B to E. 6. 52%, and 0.73 cents a share above the ask with the clean-up and fees. 7. Each piece waits the farthest latency minus its own: 310, 240, 160, 80 and 0 microseconds. 8. Named result: a spray fills 52% of the 2 500 displayed shares and costs 0.73 cents a share fee-adjusted; synchronised, it fills 100% at 0.18 cents, the average take fee. 9. Up to 10% jitter nothing changes; at 20% the synchronised sweep fills 99.3% (the whole level on 94% of trials) and at 40% 90.3%; the spray rises to 68.6%. 10. A 100%, B 44%, C 41%, D 40%, E 40%. 11. Spraying: A, E, C, B, D (0.30 to 0.72 cents); synchronised: E, C, then A, B, D ( to 0.30): a ranking measures the router’s own routing. 12. By fee (E, C): 99.8% sprayed, 100% synchronised, about zero cost; by distance (A, B): 71.2% sprayed at 0.59 cents, 100% synchronised at 0.30. 13. Its first fill came from far venues, from which the fast provider’s cancels could not overtake the other piece. 14. . 15. 0.80 (100%), 0.77 (22.0%), 0.74 (38.7%), 0.20 (61.3%) and 0.03 cents (76.4%). 16. 0.03 cents a share (0.06 without it). 17. Limit orders sent to the venues paying the largest rebates executed worse: routing for rebates does not maximise execution quality. 18. Synchronised arrivals (quote fade after a first fill), minimum fill quantities in the dark (pinging), randomised child sizes and timing (pattern detection). 19. Synchronise the sweeps, rank venues on measured fills, not displays, and allocate passive orders by expected fills, not rebates. 20. Around quotes that will not be there and queues that will not clear.
18.10 Interview questions
Interview question 18.1 ★ trader
Your sweep shows 5 000 shares across five venues and fills 2 600. What happened?
Solution
Solution of Interview question 18.1.
Quote fade: pieces arrived one after another, the first fills alerted fast providers who cancelled elsewhere. Check the arrival times per venue and the cancels just before them; synchronise.
What the interviewer is looking for: Arrival-time spread; the provider’s reaction.
Interview question 18.2 ★★ developer
How would you implement arrival-time synchronisation in a router, and where does it fail?
Solution
Solution of Interview question 18.2.
Measure one-way latency per venue (from acknowledgements, continuously), hold each piece by the difference to the slowest, and add a margin for jitter. It fails when latencies vary more than the reaction time of the fastest provider, or when a venue’s gateway queues.
What the interviewer is looking for: Latency estimation; jitter as the limit.
Interview question 18.3 ★★ researcher
How would you estimate a venue’s fill probability for passive orders, and why is it hard?
Solution
Solution of Interview question 18.3.
From the router’s own orders: fills against time resting, conditional on queue position and size; the difficulty is censoring (cancelled orders), selection (the router only posts where it expects fills) and the dependence on its own behaviour.
What the interviewer is looking for: Censoring and selection; conditioning on queue.
Interview question 18.4 ★★ researcher
Formulate the allocation of a passive order across venues as an optimisation problem.
Solution
Solution of Interview question 18.4.
Choose a market order and limits ; each limit fills what leaves the front of its queue beyond the queue ahead, up to ; minimise the expected fee-adjusted cost plus penalties for under- and over-execution (Cont and Kukanov), solved on scenarios or by stochastic approximation.
What the interviewer is looking for: Fill as ; penalties; convexity.
Interview question 18.5 ★★ bank
A client asks whether your router sends its limit orders to the venues that pay you the largest rebates. How do you answer, and how do you show it?
Solution
Solution of Interview question 18.5.
Answer with the routing logic (expected fills and cost, not rebates) and show it: fill rates, time to fill and markouts by venue for the client’s own orders, against the rebates, as Battalio, Corwin and Jennings’s analysis would.
What the interviewer is looking for: Conflict of interest; evidence by venue.
Interview question 18.6 ★★★ developer
Design the venue-statistics store of a router: what it records, how it ages the data, and how the router uses it.
Solution
Solution of Interview question 18.6.
Per venue, symbol group and order type: shares sent, filled, time to fill, fees, latencies and markouts; decayed with a half-life, with a prior for new venues; the router ranks venues by expected cost including non-fills and explores a little to keep estimates fresh.
What the interviewer is looking for: Fields; ageing; exploration.