---
title: "Order Gateway and Order Management"
book: "Low-Latency Software"
subject: quant
language: en
chapter: 21
exercises: 8
source: https://one-course.com/books/quant/13/en/chapter/21-order-gateway-and-order-management
---

# Chapter 21 — Order Gateway and Order Management

A strategy cancels a resting buy order of 100 lots and, in the same microsecond, sends its replacement one tick lower. On the way to the exchange the cancel crosses a fill: a seller has just taken the old order. The exchange rejects the cancel as too late, accepts the new order, and a second seller takes it too. For the few microseconds between the strategy’s decision and the fill report, the firm held twice the position it believed it held, and a gateway that counted only the orders it knew to be live would have let the strategy send a third. Nothing in this story is a bug in the ordinary sense: every message was valid, and every party did what the protocol says. It is what happens when two machines a network apart each change an order’s state without waiting for the other. This chapter builds the component that keeps the firm’s view of its orders honest while this goes on: the [order gateway](#def-ll-order-gateway-and-order-management-gateway), with a state machine that has names for “I asked and have not heard back”, an exposure count that assumes the worst about every order still in flight, a throttle that keeps the venue’s message budget, and a reconciliation against the exchange’s own copy of the day.

## 21.1 The order state machine

**Definition 21.1 (Order gateway, order state machine).**

An *order gateway* is the component between a firm’s strategies and a venue’s order-entry session: it turns the strategies’ requests into the venue’s messages, the venue’s reports into order states, and refuses what the firm’s limits or the venue’s rules forbid. An *order state machine* is the set of states an order can be in, as the firm sees it, with the events that move it from one to another; the firm’s view includes the *pending* states, entered when the firm has sent a request and not yet heard its answer.

The venue’s order has few states: live, partially filled, and then filled, cancelled or expired (Book 10’s simulator, its `PROTOCOL.md`). The firm’s has more, because the firm is always a little behind. After it sends a new order, the order is *pending new* until the acknowledgement arrives; after it asks for a cancel or a replace, it is *pending cancel* or *pending replace* until the answer arrives, and in the meantime the order can still fill. FIX names the same states in its order status field (*PendingNew*, *PendingCancel*, *PendingReplace*, next to *New*, *PartiallyFilled*, *Filled*, *Canceled* and *Rejected*; chapter 15): the protocol itself admits that the two sides disagree for a while. [Figure 21.1](#fig-ll-order-gateway-and-order-management-states) draws the firm’s machine.

![The firm’s view of an order. Dashed boxes are the pending states, entered when a request is sent and left when its answer arrives; in red, the fills that reach an order after the firm has asked to cancel it. Not drawn: every open state can also be cancelled by the venue on its own (a lost connection, a halt, an expiry), and pending new can fill or be cancelled before its acknowledgement. Only the terminal states, doubled, are final; the table of is the whole machine.](https://one-course.com/images/onecourse/chapters/quant-13/ll-order-gateway-and-order-management/fig-3c5c8a663982.svg)

***Figure 21.1.** The firm’s view of an order. Dashed boxes are the pending states, entered when a request is sent and left when its answer arrives; in red, the fills that reach an order after the firm has asked to cancel it. Not drawn: every open state can also be cancelled by the venue on its own (a lost connection, a halt, an expiry), and pending new can fill or be cancelled before its acknowledgement. Only the terminal states, doubled, are final; the table of [Listing 21.1](#lst-ll-order-gateway-and-order-management-table) is the whole machine.*

The build writes the machine as a table, one row per state and one column per event, with a marker for the transitions that are not allowed ([Listing 21.1](#lst-ll-order-gateway-and-order-management-table)). A table has three advantages over a nest of conditions. It is complete by construction: every pair of state and event has an entry, so the question “what if a fill arrives while a replace is pending?” has an answer that someone wrote down. It is data: the Python reference, the C++ gateway and the Rust gateway each hold a copy, and the tests check that the three copies are the same table, entry for entry, against `data/table.txt`. And an impossible event is detected, not absorbed: a report that the table does not allow is returned as an error, which is how the first version of the build found that it had no arrow from *live* to *cancelled*. The venue cancels orders on its own when a session is lost, when trading halts or when an order expires, and the conformance test against Book 10’s live server, which cuts a connection on purpose, hit the missing entry on its first run.

```cpp
inline constexpr std::uint8_t kTable[NStates][NEvents] = {
    //            Ack  Reject    Fill  FillAll  CancelReq  CancelAck  TooLate  ReplaceReq  ReplaceAck
    /*PendNew*/ {Live, Rejected, Partial, Filled, PendingCancel, Cancelled, X, X, X},
    /*Live*/    {X, X, Partial, Filled, PendingCancel, Cancelled, X, PendingReplace, X},
    /*Partial*/ {X, X, Partial, Filled, PendingCancel, Cancelled, X, PendingReplace, X},
    /*PendCxl*/ {PendingCancel, X, PendingCancel, Filled, X, Cancelled, Filled, X, X},
    /*PendRpl*/ {X, X, PendingReplace, Filled, X, Cancelled, Filled, X, Live},
    /*Filled*/  {X, X, X, X, X, X, Filled, X, X},
    /*Cxl*/     {X, X, X, X, X, X, Cancelled, X, X},
    /*Rej*/     {X, X, X, X, X, X, X, X, X},
};
```

***Listing 21.1.** The order state machine as a table: the next state for every state and event, and a marker for what is not allowed. code/firm/ordergw/cpp/firm_ordergw.hpp*

A replace deserves a word. The simulator, like most venues, gives the replaced order a new client identifier, and the gateway keeps the order’s history under it; while the replace is pending, the order can fill under its old identifier, so the gateway keeps both until the answer arrives. Whether the order keeps its place in the queue is the venue’s rule (the report says whether priority was kept); the gateway only records it.

## 21.2 Acknowledgements, in-flight orders and races

**Definition 21.2 (In-flight order, cancel–fill race).**

An *in-flight order* is an order in a pending state: a request about it has been sent and its answer has not arrived, so the order may have changed at the venue in ways the firm does not know yet. A *cancel–fill race* occurs when a fill and a request to cancel or replace the same order cross: the fill happens at the venue before the request arrives there, and its report reaches the firm after the request was sent.

[Figure 21.2](#fig-ll-order-gateway-and-order-management-race) draws one. The strategy sends its cancel at time $c$; it arrives at the venue a one-way latency $L$ later; each fill report takes $d$ to come back. Any fill in the interval from $c - d$ to $c + L$ races the cancel: the fills before $c - d$ were known when the strategy decided, and the ones after $c + L$ cannot happen, because the order is gone. If fills arrive at a Poisson rate $\lambda$ on a resting order, the probability that a cancel meets at least one is

$$
P(\text{race}) = 1 - e^{-\lambda (L + d)} \approx \lambda (L + d),
$$

as long as the order has rested longer than $L + d$ before the cancel (an order cancelled sooner has had less time exposed, and the weekend problem computes the exact probability). The race is not a failure of the venue or of the gateway. It is the latency of the path, seen from the order’s point of view, and the only ways to make it rarer are to shorten the path or to cancel less.

![A cancel–fill race. The cancel leaves at c and arrives L later; a fill at the venue between c - d and c + L is reported after the cancel was sent. The gateway, in the state pending cancel, applies the fill and then the “too late” answer, and ends in filled.](https://one-course.com/images/onecourse/chapters/quant-13/ll-order-gateway-and-order-management/fig-51c3ca9aa12d.svg)

***Figure 21.2.** A [cancel–fill race](#def-ll-order-gateway-and-order-management-inflight). The cancel leaves at $c$ and arrives $L$ later; a fill at the venue between $c - d$ and $c + L$ is reported after the cancel was sent. The gateway, in the state *pending cancel*, applies the fill and then the “too late” answer, and ends in *filled*.*

The build measures it. `firm_ordergw_sim.py` runs the Python gateway against Book 10’s matching engine: a strategy keeps one buy order working, holds it for an exponentially distributed time of mean $1\,\mathrm{m}\mathrm{s}$ and cancels it; a counterparty sells into it at random, 200 times a second; every report takes $20\,\text{µ}\mathrm{s}$ to come back. [Figure 21.3](#fig-ll-order-gateway-and-order-management-races) gives the fraction of cancels that met a fill, for one-way latencies from $10\,\text{µ}\mathrm{s}$ to $3\,\mathrm{m}\mathrm{s}$. At $100\,\text{µ}\mathrm{s}$, 2.1% of the cancels raced, against 2.2% from the exact formula; at $10\,\text{µ}\mathrm{s}$ it is 0.65%. At $1\,\mathrm{m}\mathrm{s}$ the simple formula would say 18% and the engine says 12%: the orders are held for about as long as the path takes, so the window is often shorter than $L + d$.

![Cancel–fill races against the latency of the cancel, with fills at = 200 per second and reports delayed d = 20\, µ s. The points are the Python gateway against Book 10’s matching engine (two seeded runs of 2 000 orders each per point); the solid curve is the exact probability for orders held 1\, m s on average, the dashed one the formula for orders that have rested at least L + d when cancelled. Data: fig_races.py (deterministic).](https://one-course.com/images/onecourse/chapters/quant-13/ll-order-gateway-and-order-management/fig-dc5e432117f2.svg)

***Figure 21.3.** [Cancel–fill races](#def-ll-order-gateway-and-order-management-inflight) against the latency of the cancel, with fills at $\lambda = 200$ per second and reports delayed $d = 20\,\text{µ}\mathrm{s}$. The points are the Python gateway against Book 10’s matching engine (two seeded runs of 2 000 orders each per point); the solid curve is the exact probability for orders held $1\,\mathrm{m}\mathrm{s}$ on average, the dashed one the formula for orders that have rested at least $L + d$ when cancelled. Data: `fig_races.py` (deterministic).*

What the gateway does with a race decides whether it is harmless. The hook’s story is the classic failure: the gateway counted an order as gone when it *sent* the cancel, so the replacement fitted under the position limit, and both filled. The rule that prevents it is *worst-case accounting*: until the venue has answered, an order counts for everything it could still do. An order pending cancel counts its whole remaining quantity; an order pending replace counts the larger of its old and its new quantity, because either may fill; an order pending new counts in full. The long exposure is then the position plus the worst remaining quantity of every open buy order, and a new order is sent only if it fits under the limit on top of it. With a limit of 100 lots, the hook’s gateway refuses the replacement until the cancel’s answer arrives, and the firm’s position never exceeds 100. The price is paid in opportunity: for one round trip, the strategy cannot use the quantity it is trying to withdraw.

The gateway keeps this exposure as running sums, one per side, updated whenever an order changes: take the order’s worst remaining quantity out, change the order, put it back ([Listing 21.2](#lst-ll-order-gateway-and-order-management-change)). The check before a new order is then one addition and one comparison ([Listing 21.3](#lst-ll-order-gateway-and-order-management-new)), whatever the number of open orders. A gateway that recomputes the exposure by walking its orders pays for every one of them on every new order: [Figure 21.4](#fig-ll-order-gateway-and-order-management-exposure) measures about $0.4\,\text{µ}\mathrm{s}$ for a thousand open orders kept in an array and about $1.6\,\text{µ}\mathrm{s}$ in a hash map, where the running sums cost less than the timer can resolve. A market maker quoting a few hundred instruments has thousands of open orders all day.

```cpp
    // Keep the exposure sums current: take the order's worst remaining quantity out,
    // change the order, put its new worst remaining quantity back.
    void add(const Order& o, int sign) { (o.side == 'B' ? exp_b_ : exp_s_) += sign * o.worst_leaves(); }
    template <class F> void change(Order& o, F&& f) {
        add(o, -1);
        f(o);
        add(o, +1);
    }
```

***Listing 21.2.** Worst-case exposure kept as running sums: every change to an order goes through `change`. code/firm/ordergw/cpp/firm_ordergw.hpp*

```cpp
    std::size_t new_order(std::int64_t t, std::uint64_t cl, char side, std::int64_t qty,
                          std::int64_t price, std::uint8_t* out) {
        if ((side == 'B' && exp_b_ + position_ + qty > lim_.max_long) ||
            (side == 'S' && exp_s_ - position_ + qty > lim_.max_short))
            return refuse(Why::Exposure);
        Order& o = orders_[cl & mask_];
        if (o.cl && o.state < Filled) return refuse(Why::State);  // slot holds an open order
        if (!throttle(t)) return 0;
        o = Order{cl, side, PendingNew, qty, price, 0, 0, 0, 0};
        add(o, +1);
        ++live_;
        wire::InOWriter(out).cl_ord_id(cl).locate(1).side(side)
            .qty(static_cast<std::uint32_t>(qty)).price(static_cast<std::uint32_t>(price))
            .tif('D').display('Y').post_only('N').stp_mode('N');
        return wire::InO::kLength;
    }
```

***Listing 21.3.** A new order: the exposure check against the worst case, the throttle, then the venue’s message written in place with the codec of chapter 16. code/firm/ordergw/cpp/firm_ordergw.hpp*

![The cost of the worst-case exposure check before each new order, recomputed by scanning the open orders, against their number (median of timed calls). Kept as running sums, as in firm.ordergw, it costs one addition, below the timer’s resolution at every size. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2. Data: bench_gateway.py.](https://one-course.com/images/onecourse/chapters/quant-13/ll-order-gateway-and-order-management/fig-e8bc234078c8.svg)

***Figure 21.4.** The cost of the worst-case exposure check before each new order, recomputed by scanning the open orders, against their number (median of timed calls). Kept as running sums, as in `firm.ordergw`, it costs one addition, below the timer’s resolution at every size. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2. Data: `bench_gateway.py`.*

The same benchmark times the gateway’s whole work for each step of an order’s life: writing a new order, applying its acknowledgement, a partial fill, writing the cancel and applying its answer, each decoded from the venue’s binary report. Every step costs 4 to $5\,\mathrm{n}\mathrm{s}$ at the median and 5 to $6\,\mathrm{n}\mathrm{s}$ at the 99th percentile: the gateway’s logic is not where the microseconds of an order’s path go. They go to the network and the kernel (chapter 13 and One Quant Book 14), and to the risk checks of the next chapter.

## 21.3 Throttles and message budgets

**Definition 21.3 (Token bucket).**

A *token bucket* is a rate limiter that holds up to $b$ tokens and gains $r$ tokens a second; a message may be sent only if a token is available, and sending it spends one. It allows bursts of up to $b$ messages and, over any interval of length $T$, at most $b + rT$ messages.

Venues limit how many messages a session may send: a rate per second, sometimes a burst, sometimes a weight per message type or a ratio of orders to trades (the message throttle of One Quant Book 11, chapter 10, and Book 10’s simulator, which rejects the excess with reason `T`). A gateway that relies on the venue to enforce the limit learns about it from rejections, after the fact, and a venue may disconnect a session that keeps exceeding it. The gateway therefore keeps its own budget, set a little below the venue’s, and refuses locally what would be rejected remotely. The build’s bucket counts in integer nano-tokens, so that the Python, C++ and Rust gateways agree to the last message on the same journal ([Listing 21.4](#lst-ll-order-gateway-and-order-management-bucket)).

```cpp
    bool allow(std::int64_t t) {
        if (started_) tokens_ = std::min(cap_, tokens_ + rate_ * (t - last_));
        started_ = true;
        last_ = t;
        if (tokens_ < kOne) return false;
        tokens_ -= kOne;
        return true;
    }
```

***Listing 21.4.** The token bucket in integer nano-tokens: refill for the time elapsed, then spend one token or refuse. code/firm/ordergw/cpp/firm_ordergw.hpp*

What to do with a refused message is the strategy’s decision, not the gateway’s. A cancel refused by the throttle is the worst case: the order stays in the market while the strategy wanted it out. Gateways therefore keep part of the budget for cancels, or let cancels through ahead of new orders when the budget is short; the build’s gateway reports the reason for each refusal (throttle, exposure, state) so that the strategy can tell a full budget from a limit. The build’s replace fixture runs a strategy that replaces its order every millisecond on average against a budget of 300 messages a second with a burst of two: the gateway refuses 111 of its requests for the budget, and the strategy retries.

## 21.4 Sessions: login, sequence numbers, cancel on disconnect

The venue’s order-entry session (Book 10, chapter 26) carries the gateway’s messages. Book 10’s simulator speaks SoupBinTCP: the client logs in with a user name and a password; its messages go unsequenced; the server’s reports are sequenced, numbered from one per session, but the numbers are not on the wire: each side counts. A client that loses its connection logs in again asking for the next report it has not seen, and the server replays every report from that number. The client’s half of the protocol is short ([Listing 21.5](#lst-ll-order-gateway-and-order-management-session)): count the sequenced messages, ask for the next one at login, check that the server starts there, send a [heartbeat](https://one-course.com/books/quant/13/en/chapter/15-protocols-i-fix#def-ll-protocols-i-fix-session) after a second of silence, and consider the server dead after fifteen.

```python
    def on_frame(self, t, typ, payload):
        self.last_recv = t
        if typ == "A":
            first = int(payload[10:30])
            if first != self.next_seq:     # the server starts elsewhere: messages would be lost or applied twice
                raise ValueError(f"login accepted at {first}, expected {self.next_seq}")
            self.logged_in = True
        elif typ == "J":
            raise ConnectionError(f"login rejected: {payload.decode()}")
        elif typ == "S":
            self.next_seq += 1
            return payload
        elif typ == "Z":
            self.logged_in = False
        return None
```

***Listing 21.5.** The client side of the order-entry session: a login accepted at the wrong sequence number is an error, not a detail. code/firm/ordergw/firm_ordergw.py*

A lost connection is where worst-case accounting earns its keep. While the gateway is disconnected, it does not know what its orders are doing: they may be filling, or the venue may have cancelled them. Most venues offer, and many require, a cancel on disconnect: the venue cancels the session’s resting orders when it detects that the connection is lost (the simulator does it by default, with cancel reason `D`). The gateway still counts them in full until the reports say otherwise, because cancel on disconnect is a feature of the venue, not a guarantee to the firm: it acts only on the orders it covers, when it detects the loss, and a fill can precede it. The conformance test does exactly this against Book 10’s live server on the loopback interface: three orders go out, one fills, the connection is cut, the exposure stays at 300 lots while the gateway is away, and after the new login the replayed reports cancel the other two and bring it back to 100.

**As of September 2026 — Cancel on disconnect and drop copy at one exchange.**

CME Group’s client documentation (consulted September 2026) describes Cancel On Disconnect for its iLink order-entry sessions: when a lost connection is detected, it cancels all resting futures and options orders of the disconnected session, except good-till-cancelled and good-till-date orders, and new sessions are created with it enabled by default. Its Drop Copy service sends real-time copies of the execution reports and acknowledgements sent over iLink sessions on a separate, dedicated path.

## 21.5 Drop copies and reconciliation

**Definition 21.4 (Drop copy).**

A *drop copy* is a separate session on which the venue sends copies of the reports of a firm’s trading sessions (at least its executions and cancellations) as it sends the originals; it is read-only, and serves as an independent record of the firm’s orders and fills.

The gateway’s view of the firm’s orders is built from the reports it received, applied in order; any lost, duplicated or misapplied report makes it wrong, and a wrong view is found only by comparing it with another. The [drop copy](#def-ll-order-gateway-and-order-management-dropcopy) is that other view, from the venue, on a different connection, often read by a different process (Book 10’s simulator offers one session of this kind per firm). Reconciliation compares the two: for every order, the filled quantity and whether it was cancelled ([Listing 21.6](#lst-ll-order-gateway-and-order-management-reconcile)). Each difference is a *break*, and a break is investigated, not fixed by overwriting one side: an order the [drop copy](#def-ll-order-gateway-and-order-management-dropcopy) shows filled and the gateway shows live is a lost report and an unknown position; the opposite is a report applied twice.

```python
def reconcile(gw, drop_copy):
    """Compare the gateway's filled quantities and cancellations with the venue's drop copy of E and C reports."""
    filled, cancelled = {}, set()
    for kind, cl, qty in drop_copy:
        if kind == "E":
            filled[cl] = filled.get(cl, 0) + qty
        elif kind == "C":
            cancelled.add(cl)
    breaks = []
    for cl, o in gw.orders.items():
        if o.filled != filled.get(cl, 0):
            breaks.append(f"order {cl}: gateway filled {o.filled}, drop copy {filled.get(cl, 0)}")
        if (o.state == "cancelled") != (cl in cancelled):
            breaks.append(f"order {cl}: gateway {o.state}, drop copy {'cancelled' if cl in cancelled else 'not'}")
    return breaks
```

***Listing 21.6.** Reconciliation against the drop copy: filled quantities and cancellations, order by order; every difference is a break. code/firm/ordergw/firm_ordergw.py*

In the build, every simulated run and the live test end with zero breaks, and the property tests make sure that this is not luck: over random interleavings of fills, cancels, replaces and latencies from $5\,\text{µ}\mathrm{s}$ to $2\,\mathrm{m}\mathrm{s}$, no order ever has more filled than ordered, no remaining quantity is negative, the position equals the sum of the fills, and the worst-case exposure is never below the position; and with a limit of 100 lots, the position against the matching engine never exceeds 100. The firm reconciles at several speeds: continuously against the [drop copy](#def-ll-order-gateway-and-order-management-dropcopy), every morning against the broker’s records (One Quant Book 1, chapter 7), and the order management system above the gateway (Book 10, chapter 20) against the portfolio’s books.

## 21.6 Tutorial: races, throttles and a reconciliation

**Goal.** Run the gateway against the matching engine, count its races, replay its journal in three languages, and reconcile it against the [drop copy](#def-ll-order-gateway-and-order-management-dropcopy) of a live session. **End state:** [Figure 21.3](#fig-ll-order-gateway-and-order-management-races), the replays equal to `data/expected.txt`, and a reconciliation with zero breaks.

1. **Races.** `python fig_races.py` runs `firm_ordergw_sim.run` at six latencies and writes the fractions of [Figure 21.3](#fig-ll-order-gateway-and-order-management-races) , next to the exact formula of `ll_gateway.py` .
2. **Journals.** `make_ordergw_fixtures.py` records two runs (cancel and new; cancel–replace under a tight budget) as journals of requests and reports; the C++ and Rust gateways replay them, re-making every request with the same result, and print the summaries the Python reference printed. The C++ test also feeds the reports as encoded binary messages, and counts no allocation.
3. **Session.** `tests/test_live_session.py` starts Book 10’s live server, logs in a trading session, a [drop copy](#def-ll-order-gateway-and-order-management-dropcopy) and a counterparty, cuts the trading session’s connection, logs in again, and reconciles.
4. **Costs.** `python bench_gateway.py` times each step of an order’s life and the exposure check against the number of open orders.

**What to change next.** Lower the replace fixture’s budget to 100 messages a second and watch the refusals; make the counterparty’s fills arrive in bursts instead of at a constant rate and compare the race count with the formula.

## 21.7 Build: the order gateway

**Purpose.** The firm’s single path to the venue’s order entry: every order from the strategy engine (chapter 20) passes through it, after the risk gate (chapter 22); its journal is logged (chapter 23), and its state survives a failover (chapter 24).

**Interface.** Python reference `firm_ordergw`: `TABLE`, `transition`, `TokenBucket`, `Gateway(rate, burst, max_long, max_short)` with `new`, `cancel`, `replace`, `on_report`, `worst_long`, `worst_short`, `races`; `Session`, `frames`, `reconcile`, `replay`, `summary`; `firm_ordergw_sim.run` against Book 10’s engine. C++20 `firm::ogw`: `Gateway(Limits)` with `new_order`, `cancel`, `replace` writing the venue’s messages in place, `on_report` and `on_wire`, `summary`. Rust `firm_ordergw`: `Gateway`, `replay`.

**Rules.** Every open order counts for the worst it can still do; nothing is sent past the exposure limit or the budget; a report the table does not allow is an error; an order is kept until a terminal report; the three gateways hold the same table; no allocation per message.

**Acceptance tests.** `code/firm/ordergw/`: the table identical in three languages; both journals replayed to the committed summaries in Python, C++ (from fields and from the wire) and Rust; property tests over random runs against the matching engine; the exposure limit held against the engine; the naive gateway’s double position reproduced and refused by worst-case accounting; the race fraction within a point of the formula; the live session with cancel on disconnect, replay and a zero-break reconciliation.

**Stretch.** Mass cancel and mass quotes; a budget reserved for cancels; the session layer in C++ on Book 10’s `exchsim_client.hpp`; [FIX sessions](https://one-course.com/books/quant/13/en/chapter/15-protocols-i-fix#def-ll-protocols-i-fix-session) through the engine of chapter 15.

Sources and further reading

- CME Group Client Systems Wiki, *Cancel on Disconnect* , and *Drop Copy 4.0 Service for iLink* .
- FIX Trading Community, FIX Latest, code set *OrdStatus* .

## 21.8 Exercises

**Exercise 21.1 ★.**

An order is pending cancel. A fill for its whole remaining quantity arrives, then the venue’s rejection of the cancel as too late. Give the order’s state after each report, and say what the gateway should count as its remaining quantity in between.

**Solution of Exercise 21.1.**

The fill takes it from *pending cancel* to *filled*; the “too late” answer leaves it *filled*. Before the fill arrived, the gateway counted the whole remaining quantity (the cancel might fail); after it, nothing remains.

**Exercise 21.2 ★.**

A [token bucket](#def-ll-order-gateway-and-order-management-bucket) gains 300 tokens a second and holds at most 2; it is full at time 0. Messages are sent at 0, 0, 0, 1, 4 and 5 milliseconds. Which go through?

**Solution of Exercise 21.2.**

At 0: two tokens, so the first two go through and the third is refused. At $1\,\mathrm{m}\mathrm{s}$ the bucket holds 0.3 of a token: refused. At $4\,\mathrm{m}\mathrm{s}$, 1.2: it goes through, leaving 0.2. At $5\,\mathrm{m}\mathrm{s}$, 0.5: refused. So through, through, refused, refused, through, refused.

**Exercise 21.3 ★.**

The long limit is 500 lots. The position is 200 lots long, and two buy orders of 100 lots are live. The strategy asks to replace one of them with a buy of 300. Does the gateway send the replace? What is the largest new quantity it would send?

**Solution of Exercise 21.3.**

The worst long exposure is $200 + 100 + 100 = 400$. The replace would raise one order’s worst remaining quantity from 100 to 300, adding 200, for 600 above the limit of 500: refused. The largest new quantity that fits adds 100: a buy of 200.

**Exercise 21.4 ★★.**

Fills arrive at 50 per second on a resting order, the cancel takes $300\,\text{µ}\mathrm{s}$ to reach the venue and reports take $50\,\text{µ}\mathrm{s}$ to come back. What fraction of cancels race a fill? How many races should a strategy that cancels 20 000 orders a day expect?

**Solution of Exercise 21.4.**

$1 - e^{-50 \times 350 \times 10^{-6}} = 1 - e^{-0.0175} \approx 1.7\%$, so about 347 races in 20 000 cancels (if the orders have rested at least $350\,\text{µ}\mathrm{s}$ when cancelled).

**Exercise 21.5 ★★.**

The [drop copy](#def-ll-order-gateway-and-order-management-dropcopy) shows two executions of 60 lots for order 17, both with match number 88 412; the gateway shows order 17 filled for 60. Which side is wrong, and what should the reconciliation have done?

**Solution of Exercise 21.5.**

The two executions have the same match number: they are one execution delivered twice, and the gateway, with 60, is right. The reconciliation should key executions by their match number and count each once; a duplicate is itself worth reporting, since the [drop copy](#def-ll-order-gateway-and-order-management-dropcopy) is supposed to be the independent record.

**Exercise 21.6 ★★.**

A gateway recomputes the exposure by scanning a hash map of its open orders before each new order. With a thousand open orders, what does that add to each order according to [Figure 21.4](#fig-ll-order-gateway-and-order-management-exposure), and how does it compare with the gateway’s own work per message?

**Solution of Exercise 21.6.**

About $1.6\,\text{µ}\mathrm{s}$ per new order, some three hundred times the gateway’s own work of 4 to $5\,\mathrm{n}\mathrm{s}$ per message, and more than the rest of the order’s path through the machine. Running sums remove it.

**Exercise 21.7 ★★★.**

*Coding.* Add mass cancel to the gateway: one request that cancels every open order of a side. Which states does each order enter, what does the exposure count until the answers arrive, and how do you test it against Book 10’s engine?

**Solution of Exercise 21.7.**

Every open order of the side enters *pending cancel* when the request is sent (one request, one throttle token); the exposure keeps counting each in full until its own cancellation, or a fill, arrives. Test against the engine with fills crossing the mass cancel: every order must end filled or cancelled, the [drop copy](#def-ll-order-gateway-and-order-management-dropcopy) must reconcile, and the position must stay under the limit throughout.

**Exercise 21.8 ★★★.**

*Find the flaw.* “To keep the order table small, our gateway frees an order’s slot as soon as it sends the cancel: the strategy does not want the order any more.”

**Solution of Exercise 21.8.**

A fill that crossed the cancel arrives for an order the gateway no longer knows: it is ignored (or treated as an error), so the position is wrong and the exposure understated, which is exactly the double position of the chapter’s opening. An order stays in the table until a terminal report (filled, cancelled, rejected), and the table is sized for it.

## 21.9 Problem: The Cancel That Crossed a Fill

**Problem 21.1.**

Weekend problem — the window in which a cancel can lose

A market maker quotes buy orders of 100 lots, holds each for a millisecond on average and then cancels it. Fills arrive at $\lambda = 200$ per second on a resting order, reports take $d = 20\,\text{µ}\mathrm{s}$ to come back, and the cancel takes $L$ to reach the venue. Use [Figure 21.3](#fig-ll-order-gateway-and-order-management-races) (`races.csv`).

**Part I — The window.**

1. Explain why a fill at the venue between $c - d$ and $c + L$ races a cancel sent at $c$ , and why fills outside that interval do not.
2. Derive $P(\text{race}) = 1 - e^{-\lambda(L + d)}$ for an order that has rested at least $L + d$ when the cancel is sent.
3. Compute it for $L = 10\,\text{µ}\mathrm{s}$ , $100\,\text{µ}\mathrm{s}$ and $1\,\mathrm{m}\mathrm{s}$ .
4. Show that it is approximately $\lambda (L + d)$ when that product is small, and say how small.

**Part II — The measurement.**

5. Read the simulated fractions at the same three latencies from the figure’s data.
6. Why does the simple formula overestimate them at $1\,\mathrm{m}\mathrm{s}$ ?
7. For an order held a time $h$ , the window is $\min(h, L + d)$ . What happens to the race probability as $L$ grows much larger than the holding time?
8. The committed cancel journal (latency $100\,\text{µ}\mathrm{s}$ ) has 6 races in 245 cancels. Is that consistent with the figure?

**Part III — The damage.**

9. A naive gateway counts an order as gone when it sends the cancel, and the strategy sends a replacement at once. What position can the firm reach after one race, with a limit of 100 lots?
10. With $n$ races whose replacements are all still working, what is the worst over-position?
11. What does worst-case accounting do instead, and what does it cost the strategy?
12. At $100\,\text{µ}\mathrm{s}$ , how many races does a strategy that cancels 40 000 times a day expect?

**Part IV — The verdict.**

13. State the *named result* : the race probability per cancel, and the worst over-position without in-flight accounting.
14. Which two quantities does a firm control to reduce races, and which does it not?
15. Why does halving the one-way latency not halve the races when $d$ is large?
16. Why is a race not an error to be reported to the venue?
17. Why must the gateway keep a cancelled order in its table until the terminal report?
18. What does a replace change in the analysis?
19. How would you check, on a production day, that the races observed match the formula?
20. In one sentence: what is the rule that makes races harmless?

**Solution of Problem 21.1.**

1. A fill before $c - d$ was reported before $c$ , so the strategy knew it when it cancelled; a fill after $c + L$ finds the order already cancelled. A fill between the two happens while the cancel is being decided or travelling, and its report arrives after the cancel was sent.
2. The window has length $L + d$ and the fills are Poisson at rate $\lambda$ : the probability of at least one is $1 - e^{-\lambda (L + d)}$ .
3. $1 - e^{-0.006} \approx 0.60\%$ , $1 - e^{-0.024} \approx 2.4\%$ and $1 - e^{-0.204} \approx 18.5\%$ .
4. $1 - e^{-x} \approx x - x^2/2$ : the relative error is about $x/2$ , under 5% when $\lambda(L + d) < 0.1$ .
5. 0.65%, 2.1% and 11.7%.
6. The orders are held a millisecond on average, about as long as the path: many are cancelled before they have rested $L + d$ at the venue, and some that would have raced were filled before the strategy decided to cancel.
7. The window stops growing at the holding time, so the probability tends to that of a fill during the holding time; the exact curve of the figure flattens (16% at $3\,\mathrm{m}\mathrm{s}$ ), where the simple formula keeps rising.
8. $6 / 245 = 2.4\%$ , against 2.2% for the exact probability: about 5.4 races expected, with a standard deviation of about 2.3, so yes.
9. The old order fills (100), the replacement is sent because the old order was counted as gone, and it fills too: 200 lots, twice the limit.
10. $n$ times the order’s quantity, $100\,n$ lots above the limit.
11. It keeps counting the old order in full until its cancellation or its fill is reported, and refuses the replacement until then: the firm never exceeds the limit, and the strategy waits one round trip before it can reuse the quantity.
12. $40\,000 \times 2.4\% \approx 950$ (about 880 with the exact probability for orders held a millisecond).
13. **Named result.** A cancel races a fill with probability $1 - e^{-\lambda (L + d)} \approx \lambda (L + d)$ for an order resting longer than $L + d$ (2.4% at $100\,\text{µ}\mathrm{s}$ , 200 fills a second and $20\,\text{µ}\mathrm{s}$ of report delay); without in-flight accounting, each race can put one more order’s quantity on the position, beyond the limit.
14. The one-way latency $L$ and the report delay $d$ (both paths through its own systems and lines) and how often it cancels; not the fill rate, which is the market.
15. The window is $L + d$ : halving $L$ halves only one of its terms.
16. Both sides followed the protocol: the venue filled a live order and then correctly answered that the cancel came too late.
17. Reports about it can still arrive, and a fill for an unknown order would be lost.
18. A replace races in the same way, and while it is pending either the old or the new quantity may fill: the gateway counts the larger.
19. Count, from the journal, the fills that arrived while their order was pending cancel, divide by the cancels, and compare with $\lambda (L + d)$ using the measured latencies and the fill rate of the day.
20. Count every order for the worst it can still do until the venue says otherwise.

## 21.10 Interview questions

**Interview question 21.1 ★ developer.**

What states does an order go through in an [order gateway](#def-ll-order-gateway-and-order-management-gateway), and why are there more than at the exchange?

**Solution of Interview question 21.1.**

Pending new, live, partially filled, pending cancel, pending replace, then filled, cancelled or rejected. The exchange knows the truth; the firm knows what it asked and what it heard back, so it needs states for “asked, not yet answered”, during which the order can still fill.

*What the interviewer is looking for: the pending states, and why they exist.*

**Interview question 21.2 ★★ developer.**

You send a cancel and receive a fill for the same order. What happened, and what should your system do?

**Solution of Interview question 21.2.**

The fill happened at the venue before the cancel arrived: a [cancel–fill race](#def-ll-order-gateway-and-order-management-inflight). Apply the fill (the order is filled or partially filled), expect the cancel to be rejected as too late or to cancel only what remains, and never treat the fill as an error or the order as gone before the terminal report.

*What the interviewer is looking for: the race explained, and worst-case accounting.*

**Interview question 21.3 ★★ developer.**

Implement a rate limiter for an exchange session. What do you do with a cancel when the budget is exhausted?

**Solution of Interview question 21.3.**

A [token bucket](#def-ll-order-gateway-and-order-management-bucket) in integer units, set below the venue’s limit, checked before each message. Cancels reduce risk, so they get a reserved part of the budget or priority over new orders; a refused cancel must be retried at once and reported to the strategy.

*What the interviewer is looking for: a bucket, and cancels treated differently.*

**Interview question 21.4 ★★ developer, trader.**

Your connection to the exchange drops for two seconds. What do you know about your orders, and what do you do when you reconnect?

**Solution of Interview question 21.4.**

Nothing certain: orders may have filled or been cancelled (by cancel on disconnect, if enabled). Count them all in full, stop new orders, log in again asking for the reports missed, apply the replay, reconcile with the [drop copy](#def-ll-order-gateway-and-order-management-dropcopy), and only then resume.

*What the interviewer is looking for: worst case while blind, replay by sequence number, reconciliation before resuming.*

**Interview question 21.5 ★★ developer.**

How do you check a position limit in constant time with thousands of open orders?

**Solution of Interview question 21.5.**

Keep running sums of the worst remaining quantity per side (and per instrument), updated whenever an order changes: remove its old contribution, apply the change, add the new one. The check is one addition and one comparison.

*What the interviewer is looking for: incremental aggregates rather than scans.*

**Interview question 21.6 ★★★ developer, researcher.**

Design the order management layer for a market maker on three venues. What is the source of truth for positions, and how do you know it is right?

**Solution of Interview question 21.6.**

One gateway per venue session with its state machine, throttle and worst-case exposure; a firm-wide aggregation of exposures above them for limits across venues; the venue’s reports as the source of truth for fills, checked continuously against drop copies and daily against the broker; every break investigated, never overwritten.

*What the interviewer is looking for: per-venue truth from reports, independent copies, and reconciliation.*
