Quantitative Finance · Book 11 · Market making

Market Making and High-Frequency Trading

Market Making and High-Frequency Trading · Market making

6Adverse Selection and Toxicity

A market maker’s fills from one class of trader lose almost a tick a share ten seconds later, and its fills from the others gain half a tick. Its quotes are the same for both until it decides they should not be. In this chapter’s simulated market 32% of a one-lot market maker’s fills come from informed traders, and those fills mark out at −0.95-0.95 ticks a share against +0.55+0.55 for the rest. Four features of the book, visible while the quote rests, forecast a fill’s mark-out with a correlation of 0.18 out of sample; withdrawing the side the forecast marks as toxic cuts the informed share of fills from 29% to 18%. How often to withdraw matters as much as whether to: at the natural threshold the market maker churns its orders to the back of their queues.

6.1 Mark-outs by counterparty, venue and order type

Adverse selection (One Quant Book 1, chapter 1) is measured by mark-outs (One Quant Book 2, chapter 15; One Quant Book 7, chapter 23): the move of a reference price after a fill, signed against the market maker. What makes the measure useful is grouping. A dealer answering requests for quote knows its counterparty and can mark out each client (One Quant Book 9, chapter 24, prices clients this way). On an anonymous exchange the market maker does not know who hit it, but it knows the venue, the order type that hit it (an immediate-or-cancel order routed across venues, a market order, an order that swept several levels), the size, and the time of day. Battalio, Corwin and Jennings found, in proprietary and public US equity data, that limit orders sent to venues paying large rebates were executed worse on several measures of quality, adverse selection among them: the venue is a group worth marking out by.

In a simulation the counterparty’s class is known. The simulator records, for each of the market maker’s fills, whether the aggressor was an informed trader, and the market maker never sees it; Figure 6.1 splits a one-lot quoter’s fills over six twenty-minute sessions. Informed fills turn against the quoter within two seconds and keep going: −0.95-0.95 ticks a share at ten seconds (standard error 0.14). Uninformed fills keep their half-spread: +0.55+0.55 (0.10). The informed traders are 32% of the fills and set the sign of the total, +0.07+0.07.

Mark-out of a one-lot market maker’s fills against the mid, by the class of the trader who hit it (the simulator’s truth), over six simulated sessions of twenty minutes; quantity-weighted. Data: fig_toxicity.py on hf_toxicity.run.
Figure 6.1. Mark-out of a one-lot market maker’s fills against the mid, by the class of the trader who hit it (the simulator’s truth), over six simulated sessions of twenty minutes; quantity-weighted. Data: fig_toxicity.py on hf_toxicity.run.

6.2 Detecting informed flow

The market maker cannot see the class of the next aggressor; it can see the circumstances in which informed traders arrive. Flow toxicity (One Quant Book 7, chapter 9) is the name for the share of the flow that adversely selects liquidity providers. Easley, López de Prado and O’Hara built VPIN, an estimate of it from order imbalance in volume time, and proposed it as a warning of toxicity-induced volatility. Andersen and Bondarenko found VPIN a poor volatility predictor, highest only after the flash crash, with its predictive content coming from a mechanical relation with trading intensity, and warned against adopting any stress metric before comparing it with simple benchmarks. The lesson for a market maker is to predict what it cares about, the mark-out of its own next fill, and to judge the predictor by that.

In a large-tick book the circumstances are few. Chapter 5 showed that a side whose queue is small against the other’s is about to be eaten; a burst of aggressive volume against a side announces the same thing; a mid that has not moved for a while is more likely to be stale; a spread wider than a tick says the book is being repriced. The chapter’s features are these four, computed from the order-by-order feed as the quote rests (Listing 6.1): the aggressive volume against the side over the last two seconds (lots), the other side’s share of the two best queues, the logarithm of one plus the seconds since the mid changed, and whether the spread exceeds a tick.

6.3 Toxicity scores

Definition 6.1 (Toxicity score)

A toxicity score is a market maker’s real-time estimate of the adverse selection it would suffer if one side of its quotes were filled now: here, the negative of a forecast of the fill’s mark-out at a fixed horizon, from features observable while the quote rests, fitted on the market maker’s own past fills.

Fitted by least squares on the 919 fills of six calibration sessions, the ten-second mark-out forecast is

mo^=1.43−0.029 flow−1.79 queue share of the other side−0.072ln⁡(1+quiet)+0.028 wide,\widehat{\mathrm{mo}}=1.43-0.029\,\text{flow}-1.79\,\text{queue share of the other side}-0.072\ln(1+\text{quiet})+0.028\,\text{wide},

in ticks a share. The other side’s share of the queues carries almost all of it, as chapter 5 predicted. The forecast correlates 0.19 with the realised mark-outs it was fitted on and 0.18 with those of six other sessions: small, and stable. Sorted into fifths by the forecast, the other sessions’ fills realise mark-outs from −0.86-0.86 to +0.83+0.83 ticks a share (Figure 6.2); a correlation of 0.18 on fills whose mark-outs vary by several ticks is enough to separate them.

Out of sample: the base quoter’s fills on six test sessions, sorted into fifths by the fitted forecast of their ten-second mark-out; mean forecast and realised mark-out per fifth. Data: hf_toxicity.calibration.
Figure 6.2. Out of sample: the base quoter’s fills on six test sessions, sorted into fifths by the fitted forecast of their ten-second mark-out; mean forecast and realised mark-out per fifth. Data: hf_toxicity.calibration.

6.4 Responses: widen, fade, skew, refuse

A score is worth what the market maker does with it. There are four responses, and the market structure decides which are available.

Definition 6.2 (Defensive widening)

Defensive widening moves a quote away from the fair price on the side, or both sides, where adverse selection is expected, keeping it in the book at a price that compensates for the expected mark-out; in a one-tick book it means stepping back a tick from the best price.

Fading withdraws the quote altogether; from the router’s side, One Quant Book 10, chapter 18 calls the quote that disappears as an order arrives a quote fade. Skewing moves both quotes in the direction of the expected move (chapter 7). Refusing is the dealer’s version: a dealer that knows its counterparty prices it out (client tiering, One Quant Book 9, chapter 24) or rejects it under last look (chapter 9). On an anonymous exchange a one-tick market maker has two of these: fade, or widen by stepping back (Figure 6.3).

The two responses a one-lot market maker has in a one-tick book when its bid is scored toxic (the book: bid queues in blue, ask queues in red, prices in ticks; our lot as a thick bar on top of a queue). Base: our lot rests at the best bid, 100, and at the best ask, 101. Fade: the bid is withdrawn. Widen: the bid steps back to 99, where it is almost never filled.
Figure 6.3. The two responses a one-lot market maker has in a one-tick book when its bid is scored toxic (the book: bid queues in blue, ask queues in red, prices in ticks; our lot as a thick bar on top of a queue). Base: our lot rests at the best bid, 100, and at the best ask, 101. Fade: the bid is withdrawn. Widen: the bid steps back to 99, where it is almost never filled.

The quoter of Listing 6.2 applies either response whenever a side’s score exceeds a threshold. On the six test sessions, at a threshold of zero (a side is withdrawn whenever its fill is forecast to lose):

responsesharesinformed shareten-second mark-outmessagesP&L (sd)
a sessionof fills(ticks a share, se)a session($ a session)
none19 16729.4%−0.005-0.005 (0.096)696−75.00-75.00 (178.23)
fade6 55018.3%0.136 (0.199)1 01237.75 (134.30)
widen6 68320.0%0.112 (0.222)1 513−2.08-2.08 (105.59)

Both responses remove a third of the informed share of fills, which the simulator measures exactly; their mark-outs improve by about 0.1 to 0.15 ticks a share, but the standard errors are of the same size, and six sessions cannot rank fading against widening. Both cost two thirds of the volume and send many more messages. Varying the threshold shows why:

fade above a score ofsharesinformed sharemark-out (se)messages
06 55018.3%0.136 (0.199)1 012
0.210 73317.9%0.214 (0.146)952
0.415 01724.1%0.186 (0.123)841
0.618 01730.1%−0.060-0.060 (0.103)740

At a threshold of zero the score of a side crosses the threshold back and forth, and every crossing cancels the order and sends it to the back of its queue when it returns; chapter 5 showed that the back of a queue is where the adverse fills are. A threshold of 0.2 keeps the informed share as low with 64% more volume. From 0.4 the market maker fades too rarely and the informed share returns.

6.5 Toxicity is not permanent

The same quoter’s fills in the news window of each calibration session, and before and after it:

before the newsduring the newsafter the news
fills a minute6.9717.896.69
informed share of fills34.7%23.6%32.3%
ten-second mark-out (ticks a share)0.0250.0590.132

In this market the news window brings three times the noise flow as well as a faster efficient price, so the fills during it are less toxic, not more: a rule that widens on every news window would give up the best flow of the session. The point generalises: toxicity is a property of the flow at a time, not of an instrument, a venue or a client for ever. A score is refitted on recent fills, its decay is measured as for any predictor (One Quant Book 7, chapter 13), and a counterparty tiered as toxic is re-tested.

6.6 Tutorial: scoring the next fill

Goal. Measure who is on the other side of a market maker’s fills, build a toxicity score from observable features, and test fading and widening on it. End state: the three tables and Figures 6.1 and 6.2.

  1. Features from the feed. firm.toxicity.FlowTracker updates on every message (the harness now passes the message behind each update, ctx.last).

        def update(self, t: float, last: dict | None, ext: dict) -> None:
            if last is not None and last["kind"] == b"E":
                v = last["agg"] * last["qty"] / self.lot
                self.trades.append((t, v))
                self.net += v
            while self.trades and self.trades[0][0] < t - self.window:
                self.net -= self.trades.popleft()[1]
            mid = 0.5 * (ext["bid"] + ext["ask"])
            if mid != self.mid:
                self.mid, self.t_mid = mid, t
            self.qb, self.qa = max(int(ext["bid_qty"]), 0), max(int(ext["ask_qty"]), 0)
            self.spread = int(ext["ask"]) - int(ext["bid"])
            self.now = t
    
        def features(self, side: int) -> np.ndarray:
            tot = self.qb + self.qa
            other = (self.qa if side == 1 else self.qb) / tot if tot else 0.5
            return np.array([-side * self.net, other, math.log1p(self.now - self.t_mid), float(self.spread > 1)])
    Listing 6.1. Signed flow against each side, the other side’s queue share, quiet time and spread. code/firm/toxicity/firm_toxicity.py
  2. Calibrate. hf_toxicity.fitted() regresses the ten-second mark-outs of the base quoter’s fills on the features at each fill, on six calibration sessions.
  3. Respond. ToxicQuoter fades or widens a side whose score exceeds the threshold.

        def on_market(self, ctx, t, top):
            x = ctx.external(top)
            self.flow.update(t, ctx.last, x)
            px = {1: int(x["bid"]), -1: int(x["ask"])}
            qty = {1: self.size if ctx.position + self.size <= self.limit else 0,
                   -1: self.size if ctx.position - self.size >= -self.limit else 0}
            for s in (1, -1):
                f = self.flow.features(s)
                self.cur[s] = f
                if self.model is not None and self.mode != "base":
                    sc = self.model.score(f)
                    if self.restore is None:
                        self.faded[s] = sc > self.threshold
                    elif sc > self.threshold:
                        self.faded[s] = True
                    elif sc < self.restore:
                        self.faded[s] = False
                    if self.faded[s]:
                        if self.mode == "fade":
                            qty[s] = 0
                        else:
                            px[s] -= s
            ctx.quote(px[1], qty[1], px[-1], qty[-1])
    Listing 6.2. Score each side on every update; fade or step back the toxic side. code/firm/toxicity/firm_toxicity.py
  4. Evaluate with responses() and fade_sweep() on six other sessions; by_class() and by_period() use the simulator’s truth.

What to change next. Add hysteresis (fade above 0.2, return below 0.1) and measure the messages saved; score the whole book’s toxicity with VPIN and compare it with the fill-level score.

6.7 Build: the toxicity toolkit

Purpose. Grouped mark-outs, streaming toxicity features, a fitted score and responses that any Quoter can wrap.

Interface. grouped(markouts, qty, groups); FlowTracker(window).update(t, last, ext), .features(side); ToxicityModel.fit(X, y), .predict, .score; ToxicQuoter(model, mode, threshold, size, limit, record). In firm.mmharness: ctx.last (the message behind each update) and the informed entry of Result.extra (evaluation only).

Rules. Features use only what the feed has shown at the time. A score is fitted on sessions different from those it is judged on. The simulator’s counterparty classes never reach the Quoter.

Acceptance tests. code/firm/toxicity/tests/: features by hand on a three-message feed (window, quiet time, spread); the fit recovers a planted linear mark-out; grouped means by hand. code/firm/mmharness/tests/: the Quoter sees add, cancel and execution messages; the counterparty log has one entry per fill and both classes.

Stretch. Hysteresis; a logistic score of the informed probability; per-venue scores when firm.exchsim runs several venues.

Sources and further reading

  • D. Easley, M. López de Prado and M. O’Hara, “Flow toxicity and liquidity in a high-frequency world”, Review of Financial Studies 25(5), 2012.
  • T. G. Andersen and O. Bondarenko, “VPIN and the flash crash”, Journal of Financial Markets 17, 2014.
  • R. Battalio, S. A. Corwin and R. Jennings, “Can brokers have it all? On the relation between make-take fees and limit order execution quality”, Journal of Finance 71(5), 2016.

6.8 Exercises

Exercise 6.1 ★

32% of a quoter’s fills are informed, with a mark-out of −0.95-0.95 ticks; the rest mark out at +0.55+0.55. What is the mark-out of all its fills?

Solution

Solution of Exercise 6.1.

0.32×(−0.95)+0.68×0.55=−0.304+0.374=+0.070.32\times(-0.95)+0.68\times0.55=-0.304+0.374=+0.07 ticks a share.

Exercise 6.2 ★

With the fitted forecast, compute the bid’s score when 2 lots were sold against it in the last two seconds, the ask holds 80% of the two best queues, the mid last moved half a second ago and the spread is one tick. Does the quoter fade at a threshold of 0? Of 0.2?

Solution

Solution of Exercise 6.2.

Forecast 1.432−0.029×2−1.795×0.8−0.072ln⁡1.5+0=−0.091.432-0.029\times2-1.795\times0.8-0.072\ln1.5+0=-0.09: the score is 0.09. The quoter fades at a threshold of 0 and keeps the bid at 0.2.

Exercise 6.3 ★

A rule earns 0.2 ticks a share on 10 000 shares a session with a one-cent tick. What is its ten-second edge a session?

Solution

Solution of Exercise 6.3.

0.2×10 000×$0.01=$200.2\times10\,000\times\$0.01=\$20.

Exercise 6.4 ★★

Why do fading and widening both send more messages than the base quoter?

Solution

Solution of Exercise 6.4.

Every change of a side’s state cancels an order and every return sends a new one; the base quoter only reprices when the best price moves.

Exercise 6.5 ★★

Why does the other side’s share of the queues carry almost all of the forecast in this market?

Solution

Solution of Exercise 6.5.

In a one-tick book the next price move happens when a best queue is emptied; a side holding a small share of the two queues is the one about to be emptied, which is when informed traders are taking it. The other features are weaker versions of the same information.

Exercise 6.6 ★★

From Figure 6.2, what is the realised mark-out of the fifth of fills with the best forecast, and is the difference with the worst fifth larger than its standard errors?

Solution

Solution of Exercise 6.6.

+0.83+0.83 ticks a share (standard error 0.25) against −0.86-0.86 (0.12) for the worst fifth: a gap of 1.69 ticks, several times the standard errors.

Exercise 6.7 ★★★

Coding. Run the fading quoter with hysteresis, hysteresis(0.2, 0.0): a side is faded when its score exceeds 0.2 and restored only when it falls below 0. Compare its messages, volume and informed share with fading above 0.2 without hysteresis.

Solution

Solution of Exercise 6.7.

With hysteresis: 716 messages a session against 952, 9 767 shares against 10 733, an informed share of 21.2% against 17.9%, a ten-second mark-out of 0.102. It saves a quarter of the messages but does not lower toxicity: a side kept out longer also misses the good moments and returns at the back of its queue.

Exercise 6.8 ★★★

Find the flaw. “VPIN is at its highest of the year: widen every quote on every venue until it falls.”

Solution

Solution of Exercise 6.8.

VPIN’s predictive content comes largely from trading intensity (Andersen and Bondarenko), and a market-wide stress metric says nothing about which of one’s own quotes will be picked off. Widening everywhere gives up the good flow with the bad; score one’s own fills, side by side, and act on the sides the score flags.

6.9 Problem: Who Is on the Other Side

Problem 6.1

Weekend problem — who is on the other side

A one-lot market maker in a one-tick stock wants to stop trading with the people who know more than it does, without stopping trading.

Part I — Measuring.

  1. What groups can a market maker mark out by on an anonymous exchange, and on a request-for-quote platform?
  2. What did Battalio, Corwin and Jennings find about venues?
  3. Give the mark-outs of informed and uninformed fills and the informed share.
  4. Why is the total mark-out so close to zero?

Part II — Predicting.

  1. What is VPIN, and what did Andersen and Bondarenko find?
  2. Define a toxicity score and list the chapter’s features.
  3. Give the fitted forecast and its in-sample and out-of-sample correlations.
  4. What does the calibration by fifths show?

Part III — Responding.

  1. Define defensive widening and distinguish it from fading, skewing and refusing.
  2. Which responses does a one-tick market maker have on an anonymous exchange?
  3. Give the three responses’ shares, informed shares and mark-outs.
  4. Why can six sessions not rank fading against widening?

Part IV — The verdict.

  1. State the named result: the change in the informed share of fills and in the ten-second mark-out when the quoter fades on its score at a threshold of 0.2, and the volume given up.
  2. Why is a threshold of zero worse than 0.2?
  3. What happens above 0.4?
  4. What did the news window show about toxicity over time?
  5. How would you re-test a counterparty tiered as toxic?
  6. How would the responses change on a small-tick asset?
  7. What would a dealer do with the same score?
  8. In one sentence: what is a toxicity score for?
Solution

Solution of Problem 6.1.

  1. Anonymous exchange: venue, order type, size, time of day, book state; request for quote: the counterparty itself.
  2. Limit orders routed to venues paying large rebates were executed worse on several quality measures, adverse selection among them.
  3. Informed −0.95-0.95 (se 0.14), uninformed +0.55+0.55 (0.10) ticks a share at ten seconds; 32% of fills informed.
  4. Informed fills are a third of the fills and lose almost twice what the others earn: +0.07+0.07 in total.
  5. An estimate of flow toxicity from order imbalance in volume time; a poor volatility predictor, highest only after the flash crash, its content mechanical in trading intensity.
  6. See Definition 6.1; flow against the side, the other side’s queue share, quiet time, a wide spread.
  7. 1.43−0.029 flow−1.79 share−0.072ln⁡(1+quiet)+0.028 wide1.43-0.029\,\text{flow}-1.79\,\text{share}-0.072\ln(1+\text{quiet})+0.028\,\text{wide}; 0.19 in sample, 0.18 out of sample.
  8. Realised mark-outs rise across fifths, from −0.86-0.86 to +0.83+0.83: the score orders fills correctly out of sample.
  9. See Definition 6.2; fading withdraws, skewing moves both quotes, refusing prices out or rejects a known counterparty.
  10. Fading and widening by a tick.
  11. None: 19 167 shares, 29.4%, −0.005-0.005; fade: 6 550, 18.3%, 0.136; widen: 6 683, 20.0%, 0.112.
  12. The mark-outs’ standard errors (0.10 to 0.22) are as large as the differences.
  13. Informed share from 29.4% to 17.9%, ten-second mark-out from −0.005-0.005 to 0.214 ticks a share (standard error 0.146), volume from 19 167 to 10 733 shares a session: 44% given up.
  14. Near zero the score crosses the threshold back and forth; each crossing sends the order to the back of its queue, where fills are worst.
  15. The quoter fades too rarely: at 0.4 the informed share is 24.1%, at 0.6 30.1%.
  16. Fewer informed fills during the news window (23.6% against 34.7% before): noise flow rises with the news; toxicity varies over time.
  17. Quote it again at a small size or a normal price for a period, mark out its fills, and restore it if they are no worse than the average.
  18. The choice of depth returns: widening becomes a continuous response (chapter 4’s half the expected move), and fading is rarely needed.
  19. Price the counterparty (a wider quote or a lower win rate), or reject under last look.
  20. To decide, side by side and moment by moment, where to show liquidity.

6.10 Interview questions

Interview question 6.1 ★ trader

Your fills on one venue mark out much worse than on another. List three explanations and how you would tell them apart.

Solution

Solution of Interview question 6.1.

Different counterparties (informed flow concentrates on the venue); different fees (rebate venues attract toxic takers, as Battalio, Corwin and Jennings found); different latency (you are slower on that venue and picked off). Split mark-outs by order type and time, compare fill times with your latency, and check whether the gap moves with the fee schedule.

What the interviewer is looking for: several hypotheses and a way to separate them.

Interview question 6.2 ★★ researcher

You only observe the mark-outs of fills you got. How does that bias a toxicity model, and what can you do about it?

Solution

Solution of Interview question 6.2.

Selection: fills happen more when quotes are stale or when the score was low, and quotes you withdrew have no mark-out. Randomise part of the policy, weight by the probability of being filled, or model the fill and the mark-out jointly.

What the interviewer is looking for: recognising selection and a remedy.

Interview question 6.3 ★★ researcher

A toxicity forecast correlates 0.18 with realised mark-outs. Is that useful? What else do you need to know?

Solution

Solution of Interview question 6.3.

It can be: mark-outs are very noisy, so a small correlation still sorts fills by a tick or more. What matters is the spread of realised mark-outs across score buckets out of sample, the volume given up, and the cost of acting (messages, lost queue place).

What the interviewer is looking for: thinking in buckets and costs, not in R2R^2.

Interview question 6.4 ★★ developer

A score crosses its threshold several times a second. What does that do to the quoter, and how do you fix it in code?

Solution

Solution of Interview question 6.4.

The order is cancelled and re-entered at every crossing, losing its queue place and flooding the gateway. Add hysteresis or a minimum hold, and count messages per side in the tests.

What the interviewer is looking for: priority loss and hysteresis.

Interview question 6.5 ★★ risk

A market maker with quoting obligations wants to fade whenever its score is high. What constrains it?

Solution

Solution of Interview question 6.5.

Its quoting obligation: a minimum time at or near the best prices and maximum widths (One Quant Book 1, chapter 24). It can widen within the allowed width or reduce size, but not fade for long.

What the interviewer is looking for: obligations as a constraint on the response.

Interview question 6.6 ★★★ researcher

A fraction π\pi of fills are informed and lose LL; the rest gain GG. A score flags informed fills with probability aa and uninformed ones with probability bb. Fading every flagged fill, when does the expected mark-out per fill improve, and what happens to volume?

Solution

Solution of Interview question 6.6.

Remaining fills: informed π(1−a)\pi(1-a), uninformed (1−π)(1−b)(1-\pi)(1-b); mean mark-out [(1−π)(1−b)G−π(1−a)L]/[π(1−a)+(1−π)(1−b)]\bigl[(1-\pi)(1-b)G-\pi(1-a)L\bigr]/\bigl[\pi(1-a)+(1-\pi)(1-b)\bigr], which improves on (1−π)G−πL(1-\pi)G-\pi L exactly when a>ba>b. Volume falls to π(1−a)+(1−π)(1−b)\pi(1-a)+(1-\pi)(1-b) of its level.

What the interviewer is looking for: the mixture algebra and the condition a>ba>b.

Terms defined in this chapter

See all 2333 terms in the glossary