---
title: "Child-Order Placement"
book: "Microstructure and Execution"
subject: quant
language: en
chapter: 17
exercises: 8
source: https://one-course.com/books/quant/10/en/chapter/17-child-order-placement
---

# Chapter 17 — Child-Order Placement

An algorithm must buy 800 shares in the next minute. It can join the queue at the bid and hope, or pay the spread now. Waiting saves half a spread if it fills and costs more than that if the price walks away while it waits. This chapter states the choice as a problem, measures the obvious policies in the simulated market, solves a Markov model of the best quotes by dynamic programming, lets a tabular learner find the same answer, and looks for the queue imbalance above which crossing at once is cheaper.

## 17.1 The placement problem

**Definition 17.1 (Order placement problem).**

The *order placement problem* is the choice, for a [child order](https://one-course.com/books/quant/10/en/chapter/14-the-almgrenchriss-framework#def-mx-the-almgren-chriss-framework-parent) that must be done by a deadline, of where and when to post it (at the best quote, behind it, inside the spread, or across it as a marketable order) and of how to revise it, so as to minimise the expected cost against the mid at the decision, penalised by its dispersion.

**Definition 17.2 (Non-execution risk, clean-up trade).**

*Non-execution risk* is the chance that a passive order is not filled by its deadline, and the cost of what must then be done. The *clean-up trade* is that remedy: at the deadline the unfilled remainder is cancelled and sent as a marketable order, at whatever the price has become.

A [child order](https://one-course.com/books/quant/10/en/chapter/14-the-almgrenchriss-framework#def-mx-the-almgren-chriss-framework-parent) (chapter 14) inherits a quantity and a deadline from the schedule of chapter 16. Its cost is one of three things. Crossing at once pays half the spread plus whatever the order walks through beyond the best quote. Joining the bid and filling earns half the spread; its fill probability is the queue’s business (chapter 6’s [queue value](https://one-course.com/books/quant/10/en/chapter/6-tick-size-queues-and-priority#def-mx-tick-size-queues-and-priority-queuevalue), One Quant Book 7, chapter 18’s passive fill probability). Joining and not filling adds the [clean-up trade](#def-mx-child-order-placement-risk) after the price has moved away, the adverse selection (One Quant Book 1, chapter 1) of the orders that do not fill. Harris and Hasbrouck (1996) measured the trade-off on NYSE SuperDOT orders and found that [limit orders](https://one-course.com/books/quant/10/en/chapter/1-the-limit-order-book#def-mx-the-limit-order-book-orders) placed at or better than the quote beat [market orders](https://one-course.com/books/quant/10/en/chapter/1-the-limit-order-book#def-mx-the-limit-order-book-orders) even after a penalty for the orders that did not execute.

The simulated market of this chapter is `firm.agentmkt` with calmer liquidity providers than chapter 16’s: on a 30-minute session the spread is one tick 97.7% of the time, the best quotes hold 14.6 lots (of 100 shares) on average, the mid moves once every 9.4 seconds and 137 shares trade a second, so the 800-share slice is about a tenth of a minute’s volume.

## 17.2 Passive against aggressive

Five policies work the same nine slices a session (buys and sells in turn, two minutes apart), on 24 sessions with the same order flow for every policy:

- *cross* : a marketable order for the 800 shares at once;
- *join* : a [limit order](https://one-course.com/books/quant/10/en/chapter/1-the-limit-order-book#def-mx-the-limit-order-book-orders) at the bid, left alone, cleaned up at the deadline;
- *behind* : the same one tick below the bid;
- *reprice* : join, and follow the bid when it moves away;
- *imbalance* : cross at once if the queue imbalance on the order’s side, $(B-A)/(B+A)$ with $B$ the size at the bid and $A$ at the ask for a buy (One Quant Book 7, chapter 8), exceeds 0.5, a common rule of thumb; otherwise reprice.

**Definition 17.3 (Repricing rule).**

A *repricing rule* says when a resting [child order](https://one-course.com/books/quant/10/en/chapter/14-the-almgrenchriss-framework#def-mx-the-almgren-chriss-framework-parent) is cancelled and replaced at a new price: here, whenever the best quote on its side moves away from it, at the cost of its place in the queue.

| policy | cost (s.e.) | s.d. | against crossing (s.e.) | filled passively | cleaned up |
| --- | --- | --- | --- | --- | --- |
| cross | 0.70 (0.04) | 0.53 | — | 0% | 0% |
| join | $-0.24$ (0.06) | 0.89 | $-0.94$ (0.07) | 91.0% | 9.0% |
| behind | 0.76 (0.11) | 1.59 | 0.06 (0.11) | 16.7% | 83.3% |
| reprice | $-0.22$ (0.05) | 0.72 | $-0.92$ (0.06) | 97.3% | 2.7% |
| imbalance | $-0.19$ (0.05) | 0.78 | $-0.89$ (0.06) | 90.8% | 2.7% |

***Table 17.1.** Working an 800-share [child order](https://one-course.com/books/quant/10/en/chapter/14-the-almgrenchriss-framework#def-mx-the-almgren-chriss-framework-parent) within 60 seconds in the simulated market: 216 slices, cost per share in ticks against the mid at the decision (positive: paid more), its standard deviation, the paired difference to crossing, and the shares filled passively or by the [clean-up trade](#def-mx-child-order-placement-risk). Data: `mx_place.sim_study`.*

Joining the bid costs 0.94 ticks a share less than crossing, almost the whole spread: the queue fills 91% of the order in the minute. Crossing is expensive not for its half-tick but for the 3.2 lots of the second level it walks into (0.70 ticks on average). Stepping one tick behind almost never fills (16.7%) and leaves 83% to the [clean-up trade](#def-mx-child-order-placement-risk) after the price has moved: it is the worst policy and the most dispersed. Repricing trades a few hundredths of a tick of mean for a smaller spread of outcomes (0.72 against 0.89): the order it saves from the clean-up is the one the price ran away from. The imbalance rule crossed at the start of 6.5% of the slices and paid for it: 0.03 ticks worse than repricing on average.

## 17.3 Queue management and repricing

Two facts drive a resting order: where it is in the queue and whether the queue will be there when its turn comes. Its place is known from the feed (`firm.exchsim`’s `working()` counts the shares ahead in the agent’s own view of the book); the queue’s future is what the imbalance predicts: a thin ask facing a deep bid tends to be consumed first, and the price ticks up away from a resting buy. Lehalle and Mounjid (2017) built limit-order strategies that monitor this imbalance to reduce adverse selection.

**Definition 17.4 (Queue-aware placement).**

*Queue-aware placement* conditions each decision (wait, reprice, cross) on the order’s position in its queue and on the sizes of the queues on both sides, not only on the prices and the time left.

A [repricing rule](#def-mx-child-order-placement-reprice) has two failure modes. Too eager, it gives up queue position for nothing when the quote flickers; too lazy, it leaves the order behind the market, where *behind* ended. Real engines add a minimum resting time, a tolerance of a tick before chasing, and a limit on how far to chase before crossing; on this market the simple rule already captures most of the value.

## 17.4 Placement as a Markov decision problem

Model the best quotes as queues (the birth–death processes of One Quant Book 4, chapter 8, and Huang, Lehalle and Rosenbaum’s queue-reactive model, which makes the rates depend on the queues’ sizes). A buy of $r$ lots rests behind $n$ lots at the bid; the ask holds $a$ lots. In a step $dt$, a market sell arrives with probability $\mu\,dt$ and takes one lot from the front (one of ours when $n=0$, earning half a tick against the mid), a lot ahead of us cancels with probability $\theta n\,dt$, the ask loses a lot with probability $(\mu+\theta a)\,dt$ and gains one with $\lambda\,dt$. When the ask empties, the price moves up a tick and the remaining lots start again from fresh queues, a tick worse. At each step the order may cross: $r$ lots against $a$ at the ask cost half a tick each for the first $a$ and 1.5 for the next level. Dynamic programming (One Quant Book 4, chapter 9) gives the value $V_k(r,n,a)$ with $k$ steps left:

$$
V_k=\min\bigl(C(r,a),\ \E[V_{k-1}\mid r,n,a]\bigr),\qquad V_0=C(r,a).
$$

```python
        new, flag = [np.zeros_like(v[0])], np.zeros((r_max + 1, n_max + 1, a_max + 1), bool)
        for r in range(1, r_max + 1):
            x = v[r]
            ahead = np.vstack([x[:1] * 0, x[:-1]])               # n - 1
            fill = -0.5 + v[r - 1][0][None, :]                   # our front lot fills
            sell = np.where(n > 0, ahead, fill)
            down = np.hstack([x[:, :1] * 0, x[:, :-1]])          # a - 1
            down[:, 1] = r * 1.0 + w[r]                          # the ask empties: a tick worse
            up = np.hstack([x[:, 1:], x[:, -1:]])
            cont = (p_sell * sell + p_canc * ahead + p_down * down + p_up * up
                    + p_none * x)
            cross = cc[r] <= cont
            new.append(np.where(cross, cc[r], cont))
```

***Listing 17.1.** The backward induction of the placement problem: for each number of lots left, the value of waiting one step against crossing now, on the grid of lots ahead and lots at the ask. code/firm/placement/firm_placement.py*

Fitted on the feed of the calibration session (`firm_placement.calibrate`), $\lambda=0.99$ lots a second join each best quote, each lot at the best cancels at $\theta=0.021$ a second, [market orders](https://one-course.com/books/quant/10/en/chapter/1-the-limit-order-book#def-mx-the-limit-order-book-orders) take $\mu=0.69$ lots a second from each side, and the second level holds 3.2 lots. For the 800-share slice the plan never crosses at the start: on the model, joining is worth $-0.40$ ticks a share against 0.66 for crossing. Crossing becomes right only for small, urgent orders. For a single lot the plan crosses when the imbalance exceeds a threshold that falls as the deadline nears ([Figure 17.1](#fig-mx-child-order-placement-threshold)): 0.86 with fifteen seconds left, 0.79 with five, 0.66 with three. For eight lots it never crosses at the start, with fifteen seconds or sixty.

![The queue imbalance above which the plan crosses a single lot at once, against the time left: the single-threshold rule that best matches the dynamic-programming decision over the model’s fresh queues. Data: mx_place.threshold_curve.](https://one-course.com/images/onecourse/chapters/quant-10/mx-child-order-placement/fig-e92508c92ddc.svg)

***Figure 17.1.** The queue imbalance above which the plan crosses a single lot at once, against the time left: the single-threshold rule that best matches the dynamic-programming decision over the model’s fresh queues. Data: `mx_place.threshold_curve`.*

```python

class DPPolicy:
    def __init__(self, plan: Plan):
        self.plan = plan

    def act(self, st):
        lots = math.ceil(st["left"] / LOT)
        at_best = st["price"] == st["near"]
        ahead = st["ahead"] if at_best and st["ahead"] is not None else st["near_qty"]
        n = ahead // LOT                # lots ahead: ours, or a place at the back
        if self.plan.decide(lots, n, st["far_qty"] // LOT, st["seconds_left"]):
            return "cross"
```

***Listing 17.2.** The plan as a policy of the slice executor: the lots left, the lots ahead (or a fresh place at the back after the price moved) and the lots at the far quote, looked up in the dynamic-programming table. code/firm/placement/firm_placement.py*

Run in the simulated market, the plan’s decisions coincide with repricing on every slice: it never found crossing worthwhile before the deadline. The model is too optimistic about waiting, though: it expects $-0.40$ ticks a share where the market gives $-0.22$, because its price moves only when a queue empties while the simulated fundamentalists push the price toward a value the model does not see.

Is the model’s threshold real? One lot with five seconds left, 1 824 slices, crossing and repricing paired on the same flow: [Figure 17.2](#fig-mx-child-order-placement-urgent) shows the difference by imbalance. Waiting is 0.26 ticks cheaper when the imbalance is below $-0.5$; the saving falls to a few hundredths above zero and is indistinguishable from nothing above 0.5 ($+0.03\pm0.07$ between 0.5 and 0.75, $-0.11\pm0.08$ above). The simulated market agrees that the advantage of waiting vanishes where the model says it does, but at no imbalance is crossing significantly cheaper.

![One lot with five seconds left: the cost of joining the bid (with repricing and the clean-up trade) minus the cost of crossing at once, by queue imbalance at the start, with standard errors; 1 824 paired slices. Data: mx_place.urgent_study.](https://one-course.com/images/onecourse/chapters/quant-10/mx-child-order-placement/fig-feda080c8098.svg)

***Figure 17.2.** One lot with five seconds left: the cost of joining the bid (with repricing and the [clean-up trade](#def-mx-child-order-placement-risk)) minus the cost of crossing at once, by queue imbalance at the start, with standard errors; 1 824 paired slices. Data: `mx_place.urgent_study`.*

Cont and Kukanov (2017) posed the placement problem across several venues, with fees and rebates, as a convex optimisation of how many shares to post where and how many to send as [market orders](https://one-course.com/books/quant/10/en/chapter/1-the-limit-order-book#def-mx-the-limit-order-book-orders); chapter 18 takes up the routing half.

## 17.5 Learning a placement policy

The plan needed the model’s rates. A learner can find a policy from experience alone: tabular Q-learning (One Quant Book 12, chapter 17, develops reinforcement learning) estimates the cost-to-go $Q(s,a)$ of each state and action from simulated episodes and acts greedily. Nevmyvaka, Feng and Kearns (2006) applied reinforcement learning to trade execution on NASDAQ order-book data, with a state of time left, inventory and market variables. Here `firm_placement.qlearn` runs 200 000 episodes of the Markov book with the state (lots left, lots ahead, lots at the ask, five-second bucket of time) and the actions wait or cross. Its greedy policy costs $-0.39$ ticks a share on 40 000 fresh episodes against $-0.40$ for the plan: close, in a state space of which it visited only 4.0%. The learner reached the plan’s answer (join and wait) without being told the rates; it could not do better, and in a larger state space it would need far more episodes, or a function approximator, which is Book 12’s subject.

## 17.6 Tutorial: working a child order

**Goal.** Compare placement policies for a one-minute slice in the simulated market, solve the placement problem by dynamic programming on a fitted Markov book, and learn it with Q-learning. **End state:** [Table 17.1](#tab-mx-child-order-placement-policies), Figures [17.1](#fig-mx-child-order-placement-threshold) and [17.2](#fig-mx-child-order-placement-urgent) and the numbers of sections 4 and 5.

1. **Market.** `mx_place.MARKET` , `market_stats()` ; `firm_placement.calibrate(res)` .
2. **Policies.** `SliceExecutor(slices, policy)` with `Cross` , `Post(offset, reprice)` , `Imbalance` , `DPPolicy` ; `sim_study()` .
3. **Plan.** `solve(model, r_max, horizon_s)` , `imbalance_threshold(plan, r, k)` ; `dp_thresholds()` , `threshold_curve()` .
4. **Check.** `urgent_study()` : one lot, five seconds, paired by imbalance.
5. **Learn.** `qlearn(model, …)` , `q_policy(q)` , `simulate(model, policy, …)` ; `model_study()` ; draw with `fig_place.py` .

**What to change next.** Give the fundamentalists more weight (stronger short-term drift) and find the imbalance at which crossing becomes cheaper in the simulated market; add a mid-price drift to the Markov book and see the plan’s threshold fall; replace the Q-table by a small network.

## 17.7 Build: placement

**Purpose.** The child-order layer under chapter 16’s schedules and chapter 28’s [execution algorithm](https://one-course.com/books/quant/10/en/chapter/16-benchmark-algorithms#def-mx-benchmark-algorithms-algo): policies, a fitted book model, its optimal plan and a learned baseline.

**Interface.** `BookModel(lam, theta, mu, fresh, depth2)`, `calibrate(res)`, `cross_cost(r, a, depth2)`, `solve(model, r_max, horizon_s, dt, n_max, a_max)` (`Plan.decide`, `.value`, `.fresh_value`), `imbalance_threshold(plan, r, k)`, `simulate(model, policy, r, horizon_s, …)`, `qlearn(…)`, `q_policy(q)`, `SliceExecutor(slices, policy, check_s)`, `Cross`, `Post`, `Imbalance`, `DPPolicy`.

**Rules.** Lots of 100 shares; costs in ticks against the mid at the decision, per share or summed over lots; the clean-up at the deadline is part of every passive policy; sells are the mirror image.

**Acceptance tests.** `code/firm/placement/tests/`: crossing walks the book; the plan’s value equals its own Monte Carlo cost and is no worse than always or never crossing; the crossing region grows with the lots ahead and shrinks with the ask; Q-learning approaches the plan; the executor completes cross and repricing slices.

**Stretch.** Queue-reactive rates; a drift signal in the state; several venues (chapter 18).

Sources and further reading

- L. Harris and J. Hasbrouck, “Market vs. limit orders: the SuperDOT evidence on order submission strategy”, *Journal of Financial and Quantitative Analysis* 31(2), 1996.
- Y. Nevmyvaka, Y. Feng and M. Kearns, “Reinforcement learning for optimized trade execution”, *Proceedings of the 23rd International Conference on Machine Learning* , 2006.
- W. Huang, C.-A. Lehalle and M. Rosenbaum, “Simulating and analyzing order book data: the queue-reactive model”, *Journal of the American Statistical Association* 110(509), 2015.
- R. Cont and A. Kukanov, “Optimal order placement in limit order markets”, *Quantitative Finance* 17(1), 2017.
- C.-A. Lehalle and O. Mounjid, “Limit order strategic placement with adverse selection risk and the role of latency”, *Market Microstructure and Liquidity* 3(1), 2017.

## 17.8 Exercises

**Exercise 17.1 ★.**

A buy of 8 lots crosses an ask of 5 lots with 3.2 lots at the next level. What does it cost per share against the mid, with a one-tick spread?

**Solution of Exercise 17.1.**

Five lots at half a tick and three at 1.5 ticks: $2.5+4.5=7.0$ ticks for 8 lots, 0.875 ticks a share.

**Exercise 17.2 ★.**

With the fitted rates, how long on average does a lot at the back of a 15-lot bid queue wait for the queue ahead to clear, if the queue ahead only shrinks (market sells, and cancellations at $\theta$ per lot ahead) and the price does not move?

**Solution of Exercise 17.2.**

With $n$ lots ahead the queue loses one at rate $\mu+\theta n$, so the expected time is $\sum_{n=1}^{15}1/(0.69+0.021n)=17.8$ seconds.

**Exercise 17.3 ★.**

Why does stepping one tick behind the bid cost more than crossing, when it pays no spread when it fills?

**Solution of Exercise 17.3.**

It fills only when the whole bid queue has been consumed or the price has fallen to it (16.7% of the shares); otherwise the price has usually moved away by the deadline and the [clean-up trade](#def-mx-child-order-placement-risk) (83% of the shares) pays the spread from a worse mid.

**Exercise 17.4 ★★.**

Why does the plan’s threshold fall as the deadline nears?

**Solution of Exercise 17.4.**

With less time left, a lot that waits is less likely to fill before the deadline, and the [clean-up trade](#def-mx-child-order-placement-risk) then pays the spread after any adverse move. The option to wait is worth less, so a smaller predicted risk of an up-move (a lower imbalance) is enough to cross.

**Exercise 17.5 ★★.**

The model expects $-0.40$ ticks a share for joining and the simulated market gives $-0.22$. Name two things the model leaves out that explain the gap.

**Solution of Exercise 17.5.**

Price moves not caused by a queue emptying (the fundamentalists trade toward a hidden value and move the price whatever the queues), and the correlation between order flow and those moves: passive buys fill when the price is falling toward them. [Market orders](https://one-course.com/books/quant/10/en/chapter/1-the-limit-order-book#def-mx-the-limit-order-book-orders) larger than a lot, and rates that depend on the queues, also matter.

**Exercise 17.6 ★★.**

Why does repricing lower the standard deviation of the cost more than its mean?

**Solution of Exercise 17.6.**

Repricing matters only on the slices where the price ran away: it turns their costly [clean-up trades](#def-mx-child-order-placement-risk) into later passive fills. Those slices are the tail of the cost distribution, so removing them shrinks the dispersion (0.72 against 0.89) while the mean moves by little, since most slices fill either way.

**Exercise 17.7 ★★★.**

*Coding.* Solve the plan for one lot with five seconds left. In what share of the fresh-queue states (weighted by their law) does it cross at once, with the fitted rates and with $\lambda$ halved (fewer sellers joining the ask, which then empties sooner)?

**Solution of Exercise 17.7.**

In 6.0% of fresh-queue states with the fitted rates, and in 18.6% with $\lambda$ halved: a thinner, faster-emptying ask makes an up-move likelier, and crossing at once pays more often.

**Exercise 17.8 ★★★.**

*Find the flaw.* “Our passive fills cost minus half a tick against the mid at the fill, so passive execution earns half a tick on every share.”

**Solution of Exercise 17.8.**

The fill is measured against the mid at the fill, not at the decision: the passive order filled because the price came to it, often just before falling further (adverse selection), and the orders that did not fill were cleaned up after the price moved away. Against the mid at the decision, joining earned $-0.24$ ticks a share in the simulated market, not $-0.5$.

## 17.9 Problem: Join, Step Ahead or Cross?

**Problem 17.1.**

Weekend problem — join, step ahead or cross?

An algorithm’s [child orders](https://one-course.com/books/quant/10/en/chapter/14-the-almgrenchriss-framework#def-mx-the-almgren-chriss-framework-parent) cross the spread every time; the desk asks whether they should rest instead, and when crossing is right.

**Part I — The problem.**

1. State the [order placement problem](#def-mx-child-order-placement-problem) and its cost.
2. Define [non-execution risk](#def-mx-child-order-placement-risk) and the [clean-up trade](#def-mx-child-order-placement-risk) .
3. Why does stepping ahead of the bid hardly exist in the simulated market?
4. Describe the simulated market’s spread, queues and pace.

**Part II — Policies in the simulated market.**

5. Give the cost per share and its standard deviation for crossing, joining, stepping behind and repricing.
6. Why is crossing 800 shares more expensive than half a tick?
7. What share of the order does joining fill passively, and what share goes to the [clean-up trade](#def-mx-child-order-placement-risk) ?
8. What did the imbalance rule of thumb cost, and how often did it cross?

**Part III — The plan.**

9. Write the Markov book and its transitions.
10. Write the backward induction.
11. Give the fitted rates.
12. What does the plan do with the 800-share slice, and what value does it expect?
13. Give the plan’s imbalance threshold for one lot with fifteen, five and three seconds left.
14. *State the named result* : the expected cost per share of each placement policy, and the queue-imbalance threshold above which crossing at once is cheaper.
15. What does the simulated market say about that threshold for one lot with five seconds left?

**Part IV — Learning and the answer.**

16. How does tabular Q-learning estimate the policy, and on what state?
17. How close does it come to the plan, and how much of its table did it visit?
18. Why would it struggle on a real book?
19. What would you tell the desk?
20. In one sentence: when is crossing at once right?

**Solution of Problem 17.1.**

**1.** Choose where and when to post a [child order](https://one-course.com/books/quant/10/en/chapter/14-the-almgrenchriss-framework#def-mx-the-almgren-chriss-framework-parent) due by a deadline, and how to revise it, minimising the cost against the mid at the decision. **2.** The chance of not being filled by the deadline, and the marketable order that then does the remainder. **3.** The spread is one tick 97.7% of the time: there is no price between the bid and the ask. **4.** A one-tick spread 97.7% of the time, 14.6 lots at the best on average, a mid move every 9.4 seconds, 137 shares a second. **5.** Cross 0.70 (0.53), join $-0.24$ (0.89), behind 0.76 (1.59), reprice $-0.22$ (0.72) ticks a share. **6.** It walks past the ask into the second level (3.2 lots on average). **7.** 91.0% passively, 9.0% by the [clean-up trade](#def-mx-child-order-placement-risk). **8.** It crossed on 6.5% of the slices and cost 0.03 ticks a share more than repricing. **9.** A market sell at rate $\mu$ takes the bid’s front lot, lots ahead cancel at $\theta n$, the ask shrinks at $\mu+\theta a$ and grows at $\lambda$; an empty ask moves the price up a tick and restarts from fresh queues. **10.** $V_k=\min(C(r,a),\E[V_{k-1}\mid r,n,a])$, $V_0=C(r,a)$. **11.** $\lambda=0.99$, $\theta=0.021$, $\mu=0.69$, a second level of 3.2 lots. **12.** It never crosses at the start and expects $-0.40$ ticks a share (crossing: 0.66). **13.** 0.86, 0.79 and 0.66. **14.** *Named result*: for an 800-share one-minute slice, crossing costs 0.70 ticks a share, joining $-0.24$, joining with repricing (the plan’s choice) $-0.22$, one tick behind 0.76, the imbalance rule $-0.19$; the plan never crosses such a slice at the start, and for a single lot it crosses above an imbalance of 0.86 with fifteen seconds left, 0.79 with five and 0.66 with three. **15.** Waiting’s advantage falls from 0.26 ticks below $-0.5$ to nothing measurable above 0.5; crossing is never significantly cheaper. **16.** By updating $Q(s,a)$ toward the observed cost plus the best next value on simulated episodes; the state is lots left, lots ahead, lots at the ask and a five-second time bucket. **17.** $-0.39$ ticks a share against $-0.40$, having visited 4.0% of its table. **18.** The real state is larger (more queues, flow, signals), rewards are noisy and episodes costly, so a table cannot be filled. **19.** Rest at the best quote and reprice when it moves; cross only small remainders near the deadline when the queues say the price is about to leave. **20.** When little time is left, the order is small, and the queue imbalance says the price is about to move away.

## 17.10 Interview questions

**Interview question 17.1 ★ trader.**

You must buy 1 000 shares in the next minute in a stock with a one-tick spread and deep queues. [Market order](https://one-course.com/books/quant/10/en/chapter/1-the-limit-order-book#def-mx-the-limit-order-book-orders) or [limit order](https://one-course.com/books/quant/10/en/chapter/1-the-limit-order-book#def-mx-the-limit-order-book-orders)?

**Solution of Interview question 17.1.**

A [limit order](https://one-course.com/books/quant/10/en/chapter/1-the-limit-order-book#def-mx-the-limit-order-book-orders) at the bid, repriced if the bid moves away and crossed near the deadline: with deep queues and a minute, the fill probability is high and the half-spread saved is most of the cost.

*What the interviewer is looking for: Fill probability versus the spread; a clean-up plan.*

**Interview question 17.2 ★★ researcher.**

How does queue imbalance enter the decision to cross, and why?

**Solution of Interview question 17.2.**

A deep bid against a thin ask predicts that the ask empties first and the price rises: a resting buy is then likely to be left behind. It lowers the value of waiting, most when little time is left; crossing is right above a threshold that falls toward the deadline.

*What the interviewer is looking for: Prediction of the next move; interaction with time left.*

**Interview question 17.3 ★★ researcher.**

Set up child-order placement as a dynamic programme. What is the state?

**Solution of Interview question 17.3.**

State: quantity left, position in the queue, the queues’ sizes (and any signal), time left; actions: wait, reprice, cross; cost: fills against the decision mid plus the terminal clean-up; backward induction on a fitted Markov model of the queues.

*What the interviewer is looking for: A complete state; the terminal condition.*

**Interview question 17.4 ★★ developer.**

Your repricing logic cancels and replaces an order every time the bid flickers. What goes wrong, and what would you add?

**Solution of Interview question 17.4.**

It loses queue priority each time and multiplies messages (throttles, fees, message-to-trade limits). Add a minimum resting time, a tolerance before chasing, and state that remembers the queue position it would give up.

*What the interviewer is looking for: Priority loss; message costs; hysteresis.*

**Interview question 17.5 ★★ mle.**

Would you use reinforcement learning for order placement? What would you compare it with, and how?

**Solution of Interview question 17.5.**

Only against a benchmark: the dynamic-programming plan on a fitted model and simple rules, evaluated on the same simulated flow (common random numbers) and then on live A/B slices. The learner must beat them out of sample, not on the episodes it trained on.

*What the interviewer is looking for: Baselines; paired evaluation; out-of-sample.*

**Interview question 17.6 ★★★ trader.**

Your passive fill rate is 90% but your execution cost against arrival is worse than a colleague who crosses more. How can that be?

**Solution of Interview question 17.6.**

The passive fills are adversely selected (filled when the price was about to go the wrong way), and the 10% that do not fill are cleaned up after the price ran away; a colleague crossing when the imbalance warns of a move avoids both.

*What the interviewer is looking for: Adverse selection of fills; the cost of the unfilled.*
