Low-Latency Software · Technology
22Pre-Trade Risk and Kill Switches
On 1 August 2012, one of the eight servers of a broker-dealer’s order router, Knight Capital Americas, still carried an old piece of code that a new deployment should have replaced. For 212 incoming orders it sent child orders to the market “without regard to the number of share executions” already received, in the words of the regulator’s order: 4 million executions in 154 stocks, more than 397 million shares, in about 45 minutes, and a loss of $460 million (One Quant Book 7, chapter 21, tells the deployment story, and One Quant Book 11, chapter 27, the risk policy it produced). This chapter reads the same order as an engineer would: as a list of requirements. Every paragraph that says what the firm lacked (a control that compared orders leaving the router with those that entered it, a position limit that counted outstanding orders, a firm-wide threshold that stopped orders automatically, a way to halt the router on its own misbehaviour) is a check to implement, on the path of every order, in a few tens of nanoseconds. The chapter builds that gate, then replays a runaway of the 2012 shape through it with one check at a time, and finds that most of the checks a firm would think of first would not have stopped it.
22.1 The requirements document
Definition 22.1 (Risk gate, fail-closed design)
A risk gate is the component on the path of every order, between the code that decides to send it and the order gateway, that checks it against the firm’s limits and refuses it if any is exceeded; nothing reaches the venue without passing it. A fail-closed design is one whose failures stop the activity it protects: a risk gate that has lost its limits, its reference prices or its view of the firm’s position refuses orders rather than passing them.
The regulator’s order (Release 34-70694) is precise about what was missing. The firm had controls before orders reached the router, but none after it: nothing compared “orders leaving SMARS with those that entered it”, and there were no procedures “to halt SMARS’s operations in response to its own aberrant activity”. It had a price cap, 9.5% from the best price when the parent order arrived, which did not help while prices moved less than that and did not apply to orders meant for the opening auction. It had position limits for some groups, but “these limits did not account for the firm’s exposure from outstanding orders”. It had no firm-wide capital thresholds linked to automated controls, and its position monitor “relied entirely on human monitoring”. The rule it was charged under, the market access rule, asks for controls that prevent the entry of orders that exceed “pre-set credit or capital thresholds” and of erroneous orders that “exceed appropriate price or size parameters, on an order-by-order basis or over a short period of time, or that indicate duplicative orders”. The engineering translation is a list: a price collar, size and notional limits per order, a position limit that counts orders in flight, a capital threshold across the firm, a rate limit, a duplicate check, and a kill switch that the gate itself can pull. The pre-trade risk check and the price collar are defined in One Quant Book 11 (chapter 27), which sets the policy; this chapter is about doing them fast and doing them right.
As of September 2026 — What the rules ask of pre-trade controls
In the United States, Rule 15c3-5 of the Securities Exchange Act (consulted September 2026) requires a broker or dealer with market access to have controls that prevent the entry of orders exceeding pre-set credit or capital thresholds in the aggregate for each customer and for itself, and of erroneous orders exceeding price or size parameters, order by order or over a short period, or indicating duplicative orders. In the European Union, Commission Delegated Regulation (EU) 2017/589 requires an investment firm to be able to cancel immediately any or all of its unexecuted orders on any or all venues (article 12); to have price collars, maximum order values, maximum order volumes and maximum message limits (article 15), and repeated automated execution throttles; and to generate real-time alerts within five seconds of the relevant event.
The gate sits in the trading process, on the strategy engine’s thread or right after it, before the order gateway (Figure 22.1): the engine’s order becomes the gateway’s message only if the gate says so. Everything else in the figure feeds it. The limits come from a risk service, as whole snapshots, with a heartbeat; the reference prices come from the book builder; the fills and cancellations come back from the gateway, so that the gate’s view of the position and of the orders in flight is the gateway’s; and the kill switch can be pulled by the gate itself, by a person on the risk desk, or by the venue.
22.2 Checks in nanoseconds
Definition 22.2 (Worst-case exposure)
The worst-case exposure of a firm in an instrument is its position plus everything its open orders could still add if they all filled: on the long side, the position plus the remaining quantity of every open buy order, including those whose acknowledgement or cancellation has not arrived; on the short side, the same with sell orders. A limit is respected in the worst case if a new order fits under it on top of this exposure.
A risk gate is checked on every order, so it must cost little, and cost the same whether it accepts or refuses: a gate that is fast on the common path and slow on the rare one is slowest exactly when a runaway is sending orders. The build’s gate is organised for that. The limits are held in dense tables indexed by instrument (the venue’s locate code) and by strategy, not in maps keyed by name. Whatever can be computed before the order arrives is: the collar is kept as two prices, recomputed when the reference price or the limits change, so that the collar check is one comparison; the exposures are running sums, updated by the reports as in chapter 21, so that the position check is an addition and a comparison. And every check is computed, without a branch, into one bit of a mask (Listing 22.1); the order passes if the mask is zero, and a refusal reports the lowest bit set, while the audit record keeps the whole mask, because the second reason matters when the first is fixed.
const std::int64_t long_ = pos_[instr] + open_buy_[instr] + qty;
const std::int64_t short_ = -pos_[instr] + open_sell_[instr] + qty;
std::uint32_t m = 0;
m |= std::uint32_t(firm_killed_ | desk_killed_[desk] | st.killed) << Kill;
m |= std::uint32_t(t - lim_t_ > lim_.max_age) << Stale;
m |= std::uint32_t((ref_[instr] <= 0) | (t - ref_t_[instr] > lim_.ref_max_age)) << Ref;
m |= std::uint32_t(buy ? price > hi_[instr] : price < lo_[instr]) << Collar;
m |= std::uint32_t(qty > il.max_qty) << Qty;
m |= std::uint32_t(notional > il.max_notional) << Notional;
m |= std::uint32_t(buy & (long_ > il.max_long)) << Long;
m |= std::uint32_t(!buy & (short_ > il.max_short)) << Short;
m |= std::uint32_t(st.open >= sl.max_open) << Open;
m |= std::uint32_t((desk_gross_[desk] + notional > lim_.desk_gross[desk]) |
(firm_gross_ + notional > lim_.firm_gross)) << Gross;
m |= std::uint32_t(tokens < kOne) << Throttle;
m |= std::uint32_t(dup) << Dup;
if (m) return {kCodes[__builtin_ctz(m)], m};
Figure 22.2 measures it. With every order accepted and the duplicate check off, a check costs about at the median; the duplicate check, which compares the order with the strategy’s sixteen most recent ones, adds about . The first version of that check compared five fields per remembered order and cost more than all the other checks together; packing quantity and price into one word and instrument and side into another (the wire’s quantity and price are 32-bit fields, so the packing is exact) brought it to this. On the fixture’s stream, where every kind of refusal occurs and each check follows the parsing of its input line, a check costs about at the median, the same for accepted and refused orders: the mask makes the refusals no cheaper and no dearer.
bench_risk.py.Tens of nanoseconds are small next to the path of an order through the kernel and the network (chapter 1), and there is no reason to make them smaller by checking less. The temptation runs the other way: to move the checks out of the trading process, into a separate risk process that sees the orders on a ring and answers. That costs a round trip between cores, about through the rings of chapter 12, eight times the check itself, and, worse, it creates a second failure to handle: if the risk process stops answering, the trading process must either wait or send unchecked. A gate in the trading process, with limits it can check alone, fails in one way only.
22.3 Limits in a hierarchy, updated live
Limits are set at several levels: the firm has a capital threshold, each desk a share of it, each strategy a rate, a number of open orders and its own share, each instrument a collar, a size and a position limit (the risk hierarchy of One Quant Book 6, chapter 29, and the policy of Book 11, chapter 27). An order is checked at every level it belongs to, and the gate keeps one running exposure per level. The limits change during the day, when a trader asks for more or the risk desk takes some away, and the gate applies changes the way the strategy engine applies parameters (chapter 20): as whole, versioned snapshots, between two orders, never in the middle of a check, so that no order is checked against half an old snapshot and half a new one.
The gate fails closed. A snapshot carries a maximum age: the risk service sends heartbeats, and if the last snapshot or heartbeat is older than its age, every order is refused as stale. The build’s fixture includes a risk service that goes quiet for three seconds with a maximum age of two: from the moment its last heartbeat is two seconds old until the next one arrives, the gate refuses every order, 50 of them. The same holds for the reference prices behind the collar: an instrument whose reference is older than its maximum age cannot be traded, because a collar around a stale price is no collar. Failing closed has a cost, since a strategy stops when a monitoring service fails, and the firm must decide which services the trading depends on. The rule of the build is that anything a check needs is such a service.
22.4 Kill switches: block, cancel, flatten
Definition 22.3 (Mass cancel)
A mass cancel is a request that cancels many orders at once, all the open orders of a session, of an instrument or of a side, either as one message the venue offers (the simulator’s M message) or as a cancel for every open order sent together; the orders remain in flight, and may still fill, until each cancellation is confirmed.
A kill switch stops a node of the hierarchy (a strategy, a desk, the firm) and can do three things, in increasing order of consequence. Block refuses every new order under the node: nothing more goes out, and nothing already out is touched. Cancel also sends a mass cancel for the node’s open orders; until each cancellation is confirmed, the orders count in the exposure, and those that fill in the meantime are races of the kind of chapter 21. Flatten also sends orders that close the positions, at the reference price, which the collar admits; it is the most dangerous of the three, since it trades, into a market the runaway may have moved, and the build returns those orders for a person to confirm rather than sending them. Who pulls the switch is a policy question (Book 11): the gate pulls it on a breach of a hard limit, the risk desk can pull it by hand, and many venues offer their own. Unkilling is always a person’s decision.
def kill(self, t, level, name="", action="block"):
"""block: refuse every new order under the node; cancel: also return the open orders to cancel (a mass
cancel); flatten: also return the orders that close the node's positions (at the reference price, which the
collar admits). Positions are kept per instrument for the firm, so flatten closes the firm's positions."""
self.killed.add((level, name if level != "firm" else ""))
self.audit.append((t, 0, "K", BIT["K"]))
if action == "block":
return []
cancels = sorted(cl for cl, o in self.orders.items() if self._under(level, name, o))
if action == "cancel":
return cancels
return cancels + [(instr, "S" if p > 0 else "B", abs(p), self.ref.get(instr, (0, 0))[0])
for instr, p in sorted(self.pos.items()) if p]
22.5 Where the checks sit
Checks can sit in four places, and a firm uses several. In the trading process, as here: the cheapest and the only place that sees every order before it leaves, with the firm’s own view of the orders in flight. In a separate risk process or appliance on the path: independent of the trading code, at the cost of a hop, and the only protection against a bug in the trading process itself. At the broker, when the firm reaches the venue through a broker’s sponsored access (One Quant Book 1, chapter 4): the broker is responsible under the market access rule for the orders it lets through, whatever the firm checks. And at the venue, whose collars, limits and kill switches protect the market rather than the firm. The 2012 router had checks upstream and a monitor downstream, and nothing on the path of its own output. A useful test of any architecture is to ask, for each component that generates orders, which check sees its output before the venue does.
22.6 Tutorial: a runaway, one check at a time
firm_riskgate_runaway.py rebuilds the 2012 shape against Book 10’s matching engine. The runaway sends buy child orders of 100 shares in bursts of fifteen, a microsecond apart, every : 1 500 a second on average, the rate of 4 million executions in 45 minutes, of about 99 shares each. It prices them five ticks above the last trade and never looks at its fills. A liquidity provider keeps a ladder of sell orders and moves it up a tick every ten fills, so the runaway pushes the price up as it buys. Table 22.1 and Figure 22.3 give five seconds of it, with each check switched on in turn.
| check switched on | orders sent | shares bought | first refusal | carried over 45 min |
|---|---|---|---|---|
| none | 7 500 | 750 000 | — | 405 million |
| collar, 5% around the last trade | 7 500 | 750 000 | — | 405 million |
| size and notional per order | 7 500 | 750 000 | — | 405 million |
| throttle, 500 orders a second | 2 545 | 254 500 | 135 million | |
| duplicates within | 55 | 5 500 | first repeat | 3.0 million |
| position 20 000, fills only | 210 | 21 000 | 21 000 | |
| position 20 000, in flight | 200 | 20 000 | 20 000 | |
| capital threshold, $2.5 million | 249 | 24 900 | 24 900 |
fig_runaway.py (deterministic).fig_runaway.py.The table is the chapter’s argument. The checks that look at each order alone (collar, size, notional) never trip: every child order is small, and priced near a market that the runaway itself is moving, so a collar around the last trade moves with it (the price rose 7.5% in the five seconds, more than the collar, and the collar never noticed). A throttle slows the runaway to its rate and does not stop it: 135 million shares in 45 minutes is still a disaster. The duplicate check nearly stops this runaway, because its child orders repeat, and would miss one that varied its quantity. What stops it is a limit on what the firm holds, counted in the worst case: the position limit at 20 000 shares after 200 orders, the capital threshold at 24 900. Counting only fills overshoots by the orders in flight, here one burst of ten orders, 1 000 shares; with more orders in flight, or a slower report path, it overshoots by more. The kill switch changes what happens next: without it, the runaway keeps sending and the gate keeps refusing, 7 300 refusals in five seconds; with it, the strategy is blocked at the first breach and its open orders are cancelled.
22.7 Build: the risk gate
Purpose. The check on the path of every order, between the strategy engine (chapter 20) and the order gateway (chapter 21); the policy it enforces is One Quant Book 11’s (firm.riskctl, chapter 27), and its audit records go to the binary log (chapter 23).
Interface. Python reference firm_riskgate: Limits.from_dict, from_riskctl, Gate(limits, t) with check, set_limits, heartbeat, set_reference, on_fill, on_done, kill, unkill, audit; replay, decisions; firm_riskgate_runaway.run against Book 10’s engine. C++20 firm::risk: parse_limits, Gate with the same operations and Decision{code, mask}. Rust firm_riskgate: Gate, parse_limits, replay.
Rules. Every check computed into a mask, no branch on the outcome; limits as whole snapshots between orders; stale limits or reference prices refuse (fail closed); exposure counts every order in flight; no allocation per check; a kill is undone only by a person.
Acceptance tests. code/firm/riskgate/: the fixture’s 4 874 orders replayed to the same decisions in Python, C++ and Rust, with every check refusing at least once; each check at its boundary; fail closed on stale limits and references; the kill switch’s three actions at firm, desk and strategy level; the running sums equal to a recount after random streams; zero allocations; the runaway stopped by the position limit and the capital threshold, and not by the collar; Book 11’s firm.riskctl limits, exported for the gate, replayed in Python and C++.
Stretch. A loss limit on realised and unrealised profit (exercise 7); per-venue message limits; limits shared by several gateways through shared memory, with one writer; Book 11’s monitors (loss, message rate) feeding the kill switch.
Sources and further reading
- U.S. Securities and Exchange Commission, Release No. 34-70694, In the Matter of Knight Capital Americas LLC, 2013.
- 17 CFR 240.15c3-5, Risk management controls for brokers or dealers with market access.
- Commission Delegated Regulation (EU) 2017/589 of 19 July 2016 (RTS 6).
22.8 Exercises
Exercise 22.1 ★
The reference price is 25.0000 (250 000 in units of ) and the collar is 500 basis points. Which of buy orders at 26.2500 and 26.2600, and sell orders at 23.7500 and 23.7400, pass?
Solution
Solution of Exercise 22.1.
The band is : buys up to 262 500 and sells down to 237 500 pass. The buy at 26.2500 and the sell at 23.7500 pass (they are on the boundary); the buy at 26.2600 and the sell at 23.7400 are refused.
Exercise 22.2 ★
The long limit is 5 000 shares. The position is 2 000 long, and buy orders for 1 500 and 1 000 are open, one of them not yet acknowledged. Is a buy of 600 accepted? What is the largest buy that is?
Solution
Solution of Exercise 22.2.
The worst-case long exposure is , whether or not an order has been acknowledged; a buy of 600 would make 5 100: refused. The largest buy accepted is 500.
Exercise 22.3 ★
From the regulator’s totals, compute the incident’s average number of executions a second and of shares per execution.
Solution
Solution of Exercise 22.3.
executions a second, and shares per execution: small orders, many of them.
Exercise 22.4 ★★
A burst of fifteen identical child orders leaves a microsecond apart. The duplicate window is . How many of the fifteen pass, and how would the runaway have to change to pass them all?
Solution
Solution of Exercise 22.4.
One: the other fourteen repeat it within the window. A runaway that varied its quantity by a share, or its price by a tick, from one order to the next would pass them all; the duplicate check catches loops that repeat themselves, not loops in general.
Exercise 22.5 ★★
The limits’ maximum age is two seconds and heartbeats come every . The risk service stops at time . Until when are orders accepted, and why is “until the service comes back” the wrong answer?
Solution
Solution of Exercise 22.5.
Until two seconds after the last heartbeat, which came at most before : between and seconds. A gate that kept trusting its last limits until the service returned would fail open: the limits could have been tightened, or the service could have stopped because something is wrong.
Exercise 22.6 ★★
A check costs about on the fixture’s stream. What fraction of a tick-to-trade budget of is that, and what would checking in a separate process on another core add?
Solution
Solution of Exercise 22.6.
About 3% of the budget. A separate process on another core needs a round trip through rings, about at the median (chapter 12), eight times the check, plus a decision about what to do when it does not answer.
Exercise 22.7 ★★★
Coding. Add a loss limit per strategy: realised profit from fills plus unrealised profit at the reference price, with a kill (block and cancel) when it falls below the limit. Where is the loss updated, and why is it not a pre-trade check like the others?
Solution
Solution of Exercise 22.7.
The loss is updated on every fill (realised) and every reference price (unrealised), in the gate’s report path, not in the check; when it falls below the limit the gate kills the strategy (block and cancel). It is not a check on the order, since an order does not change the loss; it is a monitor whose action is a kill. Test it with the runaway against a falling market.
Exercise 22.8 ★★★
Find the flaw. “Our risk checks run in their own process, which reads a copy of every order the gateway sends and sends a kill as soon as it sees a breach. It keeps the checks out of the latency path.”
Solution
Solution of Exercise 22.8.
It is a post-trade monitor, not a pre-trade check: the orders it sees are already on their way, and the kill arrives after a hop, while the runaway keeps sending (1 500 orders a second leave in the time a person reads an alert). If the process stops, nothing is checked and nothing says so: the design fails open. The check belongs on the path, before the gateway; a copy for monitoring is a second line.
22.9 Problem: Forty-Five Minutes
Problem 22.1
Weekend problem — which checks would have stopped the 2012 runaway
Use Table 22.1 (runaway.csv) and the regulator’s totals: 4 million executions, more than 397 million shares, about 45 minutes.
Part I — The rates.
- Compute the incident’s execution rate and check that the runaway’s 1 500 orders a second match it.
- How many shares does the unchecked runaway buy in five seconds, and in 45 minutes at the same rate? Compare with the incident.
- In the simulation, the price rose by a tick every ten fills. By how much did it rise in five seconds?
- Why did the collar never trip, although the price moved by more than the collar?
Part II — The checks that do not stop it.
- Why do the size and notional limits per order never trip?
- The throttle trips after . Why so late, and what is the runaway’s rate afterwards?
- Carry the throttled rate over 45 minutes.
- The duplicate check lets 55 orders through in five seconds. Which ones, and what small change to the runaway would defeat it?
Part III — The checks that do.
- How many orders does the position limit of 20 000 shares, counted in flight, let through, and when does it trip?
- The limit that counts only fills ends at 21 000. Explain the difference.
- In what circumstances would that difference be large?
- The capital threshold is $2.5 million. At about $100 a share, why does it stop the runaway near 24 900 shares?
Part IV — The verdict.
- State the named result: orders and shares before each check trips, carried over 45 minutes, against the incident’s totals.
- What does the kill switch add to the position limit in this simulation?
- The runaway traded 154 stocks. What does that change for a per-instrument position limit, and which check is immune to it?
- Which of the missing controls listed in the regulator’s order corresponds to each row of the table?
- Why must the gate fail closed?
- Why is a post-trade monitor, however fast, not a substitute for the gate?
- What false positives would the duplicate check raise on an ordinary day?
- In one sentence: what kind of limit stops a runaway?
Solution
Solution of Problem 22.1.
- executions in seconds: about 1 481 a second, so 1 500 matches.
- 750 000 shares in five seconds; at 1 500 orders of 100 shares a second, million in 45 minutes, against the incident’s 397 million.
- 7 500 fills make 749 ticks of 0.01 on a price of 100.00: about 7.5%.
- The collar is centred on the last trade, which the runaway moves; each child order is priced five ticks above it, well inside a 5% band. A collar catches prices far from the market, not a market pushed by the firm’s own orders.
- Each child order is 100 shares, about $10 000: far below any size or notional limit set for ordinary orders.
- The bucket starts full with 50 tokens, which the first bursts spend in about ; afterwards the runaway sends at the bucket’s rate, 500 orders a second.
- million shares.
- One order per window: once have passed, the same order is no longer a duplicate, so about eleven a second go through. A runaway that changed its quantity by one share each time would defeat the check.
- 200 orders, 20 000 shares, with the first refusal after .
- Counting only fills ignores the orders in flight: the last burst leaves before the fills of its first orders return (a round trip of against a burst of ), so ten more orders, 1 000 shares, go out.
- With more orders in flight: larger bursts, larger orders, a higher rate, or a slow report path (the incident’s position monitor lagged in high volume).
- $2.5 million at about $100 a share is about 25 000 shares; the notional is counted at the order’s price, five ticks above a rising last trade, so the threshold is reached a little earlier, at 24 900.
- Named result. At the incident’s rate of about 1 500 orders a second: with no check, or a collar, or size and notional limits per order, 7 500 orders and 750 000 shares in five seconds, 405 million shares carried over 45 minutes (the incident: 397 million); with a throttle at 500 a second, 135 million; with a duplicate check, 3.0 million; with a position limit of 20 000 counted in flight, 200 orders and 20 000 shares after (21 000 counting fills only); with a capital threshold of $2.5 million, 249 orders and 24 900 shares.
- It blocks the strategy instead of refusing its orders one by one (7 300 refusals in five seconds without it), cancels its open orders, and stops it in every instrument at once.
- A per-instrument limit stops the runaway in each instrument separately: across 154 instruments, 154 times the limit. A firm-wide capital threshold, or a kill switch at the first breach, is immune to that.
- The collar row is the 9.5% price cap; the position rows are the position limits that did not count outstanding orders (fills only) and the fix (in flight); the capital row is the missing firm-wide threshold; the kill switch is the missing procedure to halt the router; the duplicate and throttle rows are controls on the router’s own output.
- Because the moment a gate cannot check is the moment something is wrong; passing orders then gives no protection when it is needed.
- It sees executions after they happen, through a path that can lag, and a person must act on it: at 1 500 orders a second, every second of delay is 150 000 shares.
- Strategies that legitimately repeat orders: a quote refreshed at the same price and size, an order re-sent after a cancel, slices of an execution algorithm, retries of immediate-or-cancel orders.
- A limit on what the firm holds or could hold, counted in the worst case and linked to an automatic kill.
22.10 Interview questions
Interview question 22.1 ★ developer
List the pre-trade checks you would put on every order, and say which of them you would expect to catch a runaway algorithm.
Solution
Solution of Interview question 22.1.
Price collar, maximum size and notional per order, position limits counting orders in flight, capital or gross notional thresholds at desk and firm level, order and message rates, duplicate detection, open-order counts, stale-data checks, and a kill switch. The ones that catch a runaway are those on accumulated exposure (position, capital) and on rates; per-order checks catch fat fingers.
What the interviewer is looking for: the list, and the distinction between per-order and accumulated checks.
Interview question 22.2 ★★ developer
How do you make a risk check cost a few nanoseconds, and cost the same whether it accepts or refuses?
Solution
Solution of Interview question 22.2.
Precompute everything (collar bands, limits per instrument in dense arrays indexed by a small integer), keep exposures as running sums updated by reports, compute every condition into a bit mask without branches, and refuse on a nonzero mask. Measure it on both outcomes.
What the interviewer is looking for: precomputation, running sums, branch-free evaluation, measurement.
Interview question 22.3 ★★ developer, risk
Your risk service stops sending limits. What does the trading system do, and why?
Solution
Solution of Interview question 22.3.
It keeps trading on the last limits only for their stated maximum age, then refuses every order (fail closed) and alerts. The limits may have been tightened, and the silence may be a symptom of a larger failure.
What the interviewer is looking for: fail closed, with a bounded grace period.
Interview question 22.4 ★★ developer, trader
What should a kill switch do, and who should be able to pull it and undo it?
Solution
Solution of Interview question 22.4.
Block new orders, cancel open ones (the cancels are in flight until confirmed), optionally propose flattening orders for a person to confirm. The gate itself on a hard breach, the risk desk and the trader can pull it; only an authorised person, after understanding the cause, undoes it.
What the interviewer is looking for: block, cancel, flatten; automatic pull, manual release.
Interview question 22.5 ★★ developer
How do you change a strategy’s limits during the day without stopping it?
Solution
Solution of Interview question 22.5.
Publish a complete, versioned snapshot; the gate installs it between two orders, never during a check, and records the version with every decision. Checks against the new limits apply to the exposure already accumulated.
What the interviewer is looking for: atomic snapshots between events, versioned.
Interview question 22.6 ★★★ developer, risk
Read the 2012 incident as a requirements document. Design the controls, where they sit, and how you would test that they stop a runaway.
Solution
Solution of Interview question 22.6.
Controls on the output of every order-generating component, before the venue: worst-case position and capital limits linked to an automatic kill, rate and duplicate checks, collars; fail closed; post-trade monitoring as a second line. Test by replaying a synthetic runaway through the whole path, one control at a time, and asserting the orders and position at which each trips.
What the interviewer is looking for: requirements traced to the incident, and a runaway replay as the test.