---
title: "Testing and Deploying Low-Latency Systems"
book: "Low-Latency Software"
subject: quant
language: en
chapter: 25
exercises: 8
source: https://one-course.com/books/quant/13/en/chapter/25-testing-and-deploying-low-latency-systems
---

# Chapter 25 — Testing and Deploying Low-Latency Systems

A release of a trading system can pass every functional test and still be worse than the one it replaces: a log statement that moved inside the hot loop, a container that now allocates, a lock taken on a path that used to be free. None of these changes an output, so no functional test fails; each adds tens or hundreds of nanoseconds at the 99th percentile, which is where a latency-sensitive strategy lives. Catching them takes a different kind of test, one that measures, on a machine that is never perfectly quiet, and decides whether a difference is real. This chapter assembles the tests a low-latency system needs, from the ones that check what it computes (properties, differences, random inputs, the venue’s own scripts) to the one that checks how fast it computes, and ends with the way a release reaches production: gradually, reversibly, and only onto machines that have been checked.

## 25.1 The test pyramid for a trading system

**Definition 25.1 (Property-based test, differential test, fuzz testing).**

A *property-based test* generates many random inputs and checks that a stated property holds for every one of them, rather than comparing one input with one expected output. A *differential test* runs two implementations of the same specification on the same inputs and checks that their outputs are identical. *Fuzz testing* feeds a program large numbers of malformed or mutated inputs to find those that make it crash, hang or violate its invariants.

The components of this book have been tested in these ways all along. The gateway of chapter 21 is checked by properties over random interleavings against the matching engine (no order fills more than it asked for, no position exceeds the limit); the [book builder](https://one-course.com/books/quant/13/en/chapter/19-the-order-book-builder#def-ll-the-order-book-builder-builder) of chapter 19 and the engines of chapters 20 to 24 by [differential tests](#def-ll-testing-and-deploying-low-latency-systems-tests), since the Python reference, the C++ and the Rust implementations must agree on every fixture, byte for byte or hash for hash. Property-based testing is the idea of QuickCheck (Claessen and Hughes, 2000): the property is the specification, and the generator explores inputs no one would have written by hand. [Figure 25.1](#fig-ll-testing-and-deploying-low-latency-systems-pyramid) puts the kinds of test in the order in which a release meets them: many cheap ones at the bottom, run on every change, and a few expensive ones at the top, run before production.

![The tests a low-latency trading system meets on its way to production. The lower layers check what it computes and run on every change; the upper ones check how it behaves against the venue and how fast it is, and run on candidates for release.](https://one-course.com/images/onecourse/chapters/quant-13/ll-testing-and-deploying-low-latency-systems/fig-48412630bc7a.svg)

***Figure 25.1.** The tests a low-latency trading system meets on its way to production. The lower layers check what it computes and run on every change; the upper ones check how it behaves against the venue and how fast it is, and run on candidates for release.*

## 25.2 Tests backed by the exchange simulator

Unit tests stop at the process boundary; the bugs that cost money live at it: a report that arrives before the acknowledgement of its order, a cancel that crosses a fill, a session that drops in the middle of a burst, a venue that rejects a message the firm thought valid. Book 10’s simulator makes these reproducible: the same seed gives the same interleaving, a scripted fault (a dropped session, a halt, a pause) happens at the same moment, and the matching engine applies the venue’s rules exactly. Chapter 21’s property tests and chapter 24’s [failover](https://one-course.com/books/quant/13/en/chapter/24-resilience#def-ll-resilience-failover) table are simulator-backed tests in this sense: a scenario, a set of seeds or cut points, and properties checked after each run. Replays of recorded days (chapter 23) complete them: they check that a new build takes the decisions the old one took, except where a change was intended.

Random inputs reach places scenarios do not. [Listing 25.1](#lst-ll-testing-and-deploying-low-latency-systems-trusting) is a checksum reader with a classic bug: it believes the message’s BodyLength field and reads the checksum where the field says, without checking that the message is that long. The chapter’s fuzzer mutates the golden FIX messages of chapter 15 (a byte flipped, a digit changed, a field duplicated, a tail cut off) and copies each into a buffer of exactly its size, so that AddressSanitizer sees any read past it. Against this reader it stops at the first truncated message: within 12 inputs for each of ten seeds. Against the firm’s own parser it ran a million inputs without a bad read. It did find a real bug there, through an invariant rather than a crash: two mutations, a tag 11 turned into 10 and a sequence number raised by one, left the checksum unchanged, and the parser accepted a message with a CheckSum field in the middle of its body. The Python, C++ and Rust parsers of `firm.fixengine` now refuse it, and a test keeps the message.

```cpp
inline int parse_trusting(const char* p, std::size_t n) {
    const char* a = static_cast<const char*>(std::memchr(p, firm::fix::SOH, n));
    if (!a || n < 4 || p[0] != '8') return -1;
    const std::size_t f = static_cast<std::size_t>(a - p) + 1;  // the BodyLength field
    if (f + 2 > n || p[f] != '9' || p[f + 1] != '=') return -1;
    std::size_t k = f + 2, declared = 0;
    while (k < n && p[k] >= '0' && p[k] <= '9')
        declared = declared * 10 + static_cast<std::size_t>(p[k++] - '0');
    if (k >= n || p[k] != firm::fix::SOH || declared > 4096) return -1;
    const char* ck = p + k + 1 + declared + 3;  // after "10=": never checked against n
    return (ck[0] - '0') * 100 + (ck[1] - '0') * 10 + (ck[2] - '0');
}
```

***Listing 25.1.** A checksum reader that trusts BodyLength: the fuzzer’s first truncated message makes it read past the buffer. code/low-latency/25-testing-and-deploying-low-latency-systems/cpp/ll_fuzz.hpp*

## 25.3 Exchange certification and conformance

**Definition 25.2 (Exchange certification).**

*Exchange certification* is the process by which a venue checks, before granting access to production, that a client’s system handles its protocols correctly: the client runs a script of scenarios against the venue’s test system, which verifies each message and response.

Venues certify the systems that connect to them, and regulators ask firms to test their conformance with each venue before trading and after every material change (the dated box). A certification script is a list of steps, each a message and the responses it must produce: an order accepted, a replace, a cancel, a cancel that comes too late, a reused identifier, a price off the tick grid, an immediate-or-cancel order with nothing to trade, a post-only order that would cross, a [mass cancel](https://one-course.com/books/quant/13/en/chapter/22-pre-trade-risk-and-kill-switches#def-ll-pre-trade-risk-and-kill-switches-mass), a disconnect that cancels resting orders. `firm.perfgate`’s runner plays such a script against Book 10’s matching engine ([Listing 25.2](#lst-ll-testing-and-deploying-low-latency-systems-certify)); the build’s seventeen steps pass, and a script that expects the wrong behaviour of a single step fails at that step and no other. The same runner, pointed at a new version of the simulator or of the firm’s gateway, is a conformance test in the regulator’s sense.

```python
    e.process(2, 0, nt["ctl"]["L"](2, 2, "N"))
    e.process(3, 0, nt["ctl"]["P"](0, "T", "    "))
    return e, nt


def certify(script, engine=None):
    """script: [{"step", "session", "msg": [type, fields...] or {"ctl": [...]}, "expect": [[kind, reason or ""], ...]}]
    Each step sends one message and compares the kinds (and reasons) of the reports to the session with the expected
    ones, in order. Returns (step, ok, got) per step."""
    e, nt = engine or _engine()
    out = []
    t = 1_000
    for s in script:
        t += 1_000
        if "ctl" in s:
            _, reps = e.process(t, 0, nt["ctl"][s["ctl"][0]](*s["ctl"][1:]))
        else:
            m = s["msg"]
            _, reps = e.process(t, s.get("session", 1), nt["in"][m[0]](*m[1:]))
        got = []
        for sess, r in reps:
            if sess == s.get("session", 1):
```

***Listing 25.2.** The certification runner: each step sends one message and compares the responses, in order, with the script’s. code/firm/perfgate/firm_perfgate.py*

**As of September 2026 — Conformance testing and certification.**

In the European Union (consulted September 2026), Commission Delegated Regulation (EU) 2017/589 requires firms to establish methodologies to develop and test algorithmic trading systems before deployment or substantial update; to test their conformance with a trading venue’s system when becoming a member, when connecting through sponsored access for the first time, when the venue’s systems change materially and before deploying or materially updating an algorithm, verifying that it interacts with the venue’s matching logic as intended; and to test in an environment separated from production. CME Group requires client systems that route orders to or process market data from its Globex platform to be certified with AutoCert+, its automated testing tool.

## 25.4 Performance regression gates

**Definition 25.3 (Performance regression gate).**

A *performance regression gate* is a test that runs a benchmark on the current build and on a candidate, on the same machine, and refuses the candidate if a latency statistic is worse by more than a stated tolerance with a stated confidence.

A single benchmark run says little, even on an otherwise idle laptop. [Figure 25.2](#fig-ll-testing-and-deploying-low-latency-systems-runs) shows 800 runs of the gate’s benchmark (the strategy engine on the small-tick day of chapter 19, timed per event), alternating between the current build and one with a planted regression. The 99th percentile of a run moves by tens of percent from one run to the next with the operating system’s and the hypervisor’s own activity; the regression, which adds about 18% at the median of the runs, is invisible in any single pair. The gate is built for this noise. It interleaves the runs (A, B, B, A, …), so that a slow drift of the machine affects both builds alike; it takes one statistic per run, the 99th percentile, and compares their medians across runs; it decides with a permutation test, which assumes nothing about the distribution, and reports a bootstrap interval of the rise ([Listing 25.3](#lst-ll-testing-and-deploying-low-latency-systems-compare)); and it flags a regression only if the rise is both significant and larger than a tolerance of 2%, so that a real but negligible change does not block a release.

```python
def compare(a, b, tolerance=0.02, alpha=0.05, n_perm=2000, n_boot=2000, seed=0):
    rng = random.Random(seed)
    ma, mb = statistics.median(a), statistics.median(b)
    rise = mb / ma - 1
    observed = mb - ma
    pooled, na = list(a) + list(b), len(a)
    hits = 0
    for _ in range(n_perm):
        rng.shuffle(pooled)
        if statistics.median(pooled[na:]) - statistics.median(pooled[:na]) >= observed:
            hits += 1
    p = (hits + 1) / (n_perm + 1)
    low = high = float("nan")
    if n_boot:
        boots = []
        for _ in range(n_boot):
            ra = [rng.choice(a) for _ in a]
            rb = [rng.choice(b) for _ in b]
            boots.append(statistics.median(rb) / statistics.median(ra) - 1)
        boots.sort()
        low, high = boots[int(0.05 * n_boot)], boots[int(0.95 * n_boot) - 1]
    return Verdict(rise, p, low, high, p < alpha and rise > tolerance)
```

***Listing 25.3.** The gate’s decision: the rise of the median 99th percentile, a one-sided permutation test, a bootstrap interval, and a tolerance. code/firm/perfgate/firm_perfgate.py*

![The 99th percentile of each of 800 runs of the gate’s benchmark (18 000 timed events each), alternating between the current build and one with a planted regression, on a laptop (Intel Core Ultra 7 155H, WSL2), the machine otherwise idle. Data: bench_gate.py.](https://one-course.com/images/onecourse/chapters/quant-13/ll-testing-and-deploying-low-latency-systems/fig-3aec7244f5bd.svg)

***Figure 25.2.** The 99th percentile of each of 800 runs of the gate’s benchmark (18 000 timed events each), alternating between the current build and one with a planted regression, on a laptop (Intel Core Ultra 7 155H, WSL2), the machine otherwise idle. Data: `bench_gate.py`.*

How many runs does the gate need? [Figure 25.3](#fig-ll-testing-and-deploying-low-latency-systems-power) answers from the machine’s own noise. It draws gates of $n$ runs per side from the measured pools: two disjoint draws from the current build’s runs give the false-alarm rate, which stays near the test’s level of 5% (3 to 9% over a hundred gates); two draws of which one is shifted by 5% of the median give the detection rate of a 5% rise; and the planted regression against the current build gives its detection rate. The planted regression, at about 18%, is caught 99% of the time with 40 runs per side. A 5% rise needs far more: 73% at 120 runs, 84% at 160 and 91% at 200, so about 200 runs per side for 90%; the normal approximation of `runs_needed`, from the runs’ spread of about 14%, says about 220. At some 30 milliseconds a run, that is about twelve seconds of benchmark per gate on this laptop, and on a production-like machine with isolated cores (chapter 13), whose runs vary less, it is far fewer.

![The gate’s detection and false-alarm rates against the number of runs per side, from 100 gates drawn from the measured runs of for each point; the dashed line is 90% power. Data: fig_gate.py (deterministic, from the measured pools).](https://one-course.com/images/onecourse/chapters/quant-13/ll-testing-and-deploying-low-latency-systems/fig-4adb47f839a0.svg)

***Figure 25.3.** The gate’s detection and false-alarm rates against the number of runs per side, from 100 gates drawn from the measured runs of [Figure 25.2](#fig-ll-testing-and-deploying-low-latency-systems-runs) for each point; the dashed line is 90% power. Data: `fig_gate.py` (deterministic, from the measured pools).*

## 25.5 Staged rollouts and rollback

A release that has passed the gate still meets production gradually: the canary deployment and staged rollout of One Quant Book 7 (chapter 21), with paper trading where the strategy allows it. Low-latency systems add three habits. The host is checked before the process starts: `firm.perfgate.preflight` runs the tuning audit of chapter 13 against the core placement plan and stops the deployment on any failure, because a strategy tested on isolated cores and deployed on a host whose interrupts land on them is not the strategy that was tested (this laptop fails 7 of the audit’s 11 rules, which is why its measurements carry the words “no isolated cores”). The new build runs first next to the old one, on the same inputs, in shadow: the replay of chapter 23 made live, with its outputs compared and discarded. And rollback is a deployment like any other, prepared and tested before it is needed, since the moment it is needed is the moment the team has the least time; the [failover](https://one-course.com/books/quant/13/en/chapter/24-resilience#def-ll-resilience-failover) of chapter 24 is the same lesson for machines.

## 25.6 Tutorial: fuzz, certify, gate

**Goal.** Fuzz the FIX parser under the sanitisers, run a certification script against the simulator, and gate a planted regression on this laptop. **End state:** the fuzzing results, a passing certification, and [Figure 25.3](#fig-ll-testing-and-deploying-low-latency-systems-power).

1. **Fuzz.** `ll_fuzz_test.cpp` runs 200 000 mutations through the firm’s parser and checks that every accepted message lies within its buffer and carries a correct checksum; `python bench_fuzz.py` builds the fuzzer with AddressSanitizer and UndefinedBehaviorSanitizer and runs it against both parsers.
2. **Certify.** `firm_perfgate.certify(load_script("data/cert_script.json"))` plays the seventeen steps against Book 10’s engine.
3. **Gate.** `python bench_gate.py` measures 400 interleaved pairs; `python fig_gate.py` turns them into detection and false-alarm rates.
4. **Preflight.** `firm_perfgate.preflight` on the tuning audit’s two fixture hosts: the tuned one passes, the mistuned one does not.

**What to change next.** Gate on the median of the runs’ medians instead of their 99th percentiles and compare the number of runs needed; add a step to the certification script for the venue’s throttle.

## 25.7 Build: the performance gate

**Purpose.** The last checks before production for every component of the book: its speed against the current build, its behaviour against the venue’s script, and the host it will run on.

**Interface.** Python `firm_perfgate`: `interleave`, `run_pairs(cmd_a, cmd_b, n_pairs, parse)`, `compare(a, b, tolerance, alpha)` returning a `Verdict` with `rise`, `p_value`, `low`, `high`, `regression` and `report()`; `power`, `runs_needed`; `certify(script)`, `load_script`; `preflight(root, plan)`.

**Rules.** Runs interleaved, never all of A then all of B; one statistic per run; a permutation test and a tolerance; the number of runs set from the machine’s measured noise; no deployment on a host that fails the audit.

**Acceptance tests.** `code/firm/perfgate/`: a clear rise flagged and noise or a rise under the tolerance not; false alarms near the test’s level on synthetic noise; power that grows with the number of runs; the certification script passing, and failing at exactly the altered step; the preflight passing the tuned fixture host and failing the mistuned one.

**Stretch.** A gate on the whole latency distribution (a two-sample test on pooled samples, with care for dependence within runs); certification steps for sessions and throttles; the gate in continuous integration on a dedicated machine.

Sources and further reading

- K. Claessen and J. Hughes, “QuickCheck: a lightweight tool for random testing of Haskell programs”, *ICFP* , 2000.
- Commission Delegated Regulation (EU) 2017/589 (RTS 6), articles 5 to 7.
- CME Group Client Systems Wiki, *Client Application Testing and Certification* .

## 25.8 Exercises

**Exercise 25.1 ★.**

State two properties that a test over random interleavings should check for the [order gateway](https://one-course.com/books/quant/13/en/chapter/21-order-gateway-and-order-management#def-ll-order-gateway-and-order-management-gateway) of chapter 21, and one that it cannot check without the venue’s [drop copy](https://one-course.com/books/quant/13/en/chapter/21-order-gateway-and-order-management#def-ll-order-gateway-and-order-management-dropcopy).

**Solution of Exercise 25.1.**

No order ever has more filled than it asked for; the position never exceeds the limit counted in the worst case (or: the position equals the sum of the fills). That the gateway’s view agrees with the venue’s records needs an independent record of the venue’s side, the [drop copy](https://one-course.com/books/quant/13/en/chapter/21-order-gateway-and-order-management#def-ll-order-gateway-and-order-management-dropcopy).

**Exercise 25.2 ★.**

Why does the fuzzer copy each input into a buffer of exactly its size before parsing it?

**Solution of Exercise 25.2.**

So that a read one byte past the input is a read past an allocation, which AddressSanitizer reports. Inside a larger buffer (a string’s capacity, a reused array) the same bad read would return garbage silently.

**Exercise 25.3 ★.**

The runs of A and B are interleaved A, B, B, A. What goes wrong if all the runs of A come first, then all the runs of B?

**Solution of Exercise 25.3.**

Anything that changes with time (another job starting, the processor warming up, a frequency change) then affects one build and not the other, and is measured as a difference between them. Interleaving spreads drifts over both.

**Exercise 25.4 ★★.**

With runs whose 99th percentiles vary by 14%, how many runs per side does the normal approximation ask for to detect a 10% rise with 90% power, and a 2.5% rise?

**Solution of Exercise 25.4.**

$2\,(1.645 + 1.282)^2\,(1.2533 \times 0.14 / r)^2$: about 53 runs per side for $r = 10\%$, about 840 for $r = 2.5\%$. The number grows as the inverse square of the rise.

**Exercise 25.5 ★★.**

A gate runs 20 times a day at a level of 5%. How many false alarms should the team expect in a month of 21 trading days, and what does the tolerance of 2% change?

**Solution of Exercise 25.5.**

$20 \times 21 = 420$ gates; at a level of 5% with no real change, up to 21 false alarms. The tolerance flags only rises larger than 2%, which removes many of them: the measured rate on this laptop was 3 to 9%, some 13 to 38 a month; a larger tolerance or a lower level would cut them, at the cost of power. Each false alarm costs a rerun, so the level and the tolerance are set with the team’s patience in mind.

**Exercise 25.6 ★★.**

Two mutations changed a tag from 11 to 10 and a sequence number from 2 to 3, and the checksum did not change. Explain why, and say what the FIX checksum does and does not protect against.

**Solution of Exercise 25.6.**

The checksum is the sum of the bytes modulo 256. Turning “1” into “0” lowers the sum by one, turning “2” into “3” raises it by one: the sum is unchanged. It catches most single-byte corruptions and truncations; it does not catch compensating changes, reorderings of bytes, or anything deliberate. The protection against corruption on the wire is TCP’s; the parser’s checks are about structure.

**Exercise 25.7 ★★★.**

*Coding.* Add certification steps for the venue’s throttle: configure the engine with a rate, send a burst, and check the rejections with reason `T`. What does the step need that the others did not?

**Solution of Exercise 25.7.**

A venue configured with a message rate (the simulator’s throttle), a burst of orders sent at the same time stamp, and expectations for a sequence of acceptances followed by rejections with reason `T`; the step depends on time, so the runner must control the time stamps and the engine’s configuration, which the other steps did not need.

**Exercise 25.8 ★★★.**

*Find the flaw.* “Our performance test runs the benchmark once on the release candidate and fails if the 99th percentile is more than 5% above last week’s number, which we keep in a file.”

**Solution of Exercise 25.8.**

One run of each, a week apart: the machine, its load, its temperature and the rest of the system differ between the two, and on this laptop a single run’s 99th percentile moves by far more than 5%. The gate would fail at random and pass real regressions. Run both builds now, interleaved, many times, and compare them with a test.

## 25.9 Problem: A Few Hundred Nanoseconds

**Problem 25.1.**

Weekend problem — how many runs a gate needs on a noisy machine

Use Figures [25.2](#fig-ll-testing-and-deploying-low-latency-systems-runs) and [25.3](#fig-ll-testing-and-deploying-low-latency-systems-power) (`measured_gate_runs.csv`, `gate.csv`, `gate_summary.csv`).

**Part I — The noise.**

1. What is the median 99th percentile of the current build’s runs, and how much do runs vary (the robust coefficient of variation)?
2. What does the planted regression add, at the median of the runs?
3. Why is the 99th percentile of a run so much noisier than its median?
4. Why is a single A/B pair useless on this machine?

**Part II — The gate.**

5. Why compare medians across runs rather than means?
6. What does the permutation test assume, and what does it not assume?
7. What is the false-alarm rate at every $n$ , and why is it near the test’s level?
8. How many runs per side detect the planted regression 99% of the time?

**Part III — Power.**

9. Read the detection rate of a 5% rise at 80, 120 and 160 runs per side.
10. Interpolate the number of runs for 90% power.
11. Compare with the normal approximation, and explain the difference.
12. How long does a gate of that size take on this laptop, at about 30 milliseconds a run?

**Part IV — The verdict.**

13. State the *named result* : the number of interleaved runs the gate needs to detect a 5% rise in the 99th percentile with 90% power at this machine’s measured noise.
14. What would isolated cores (chapter 13) change, and why?
15. Why is a tolerance needed on top of the test?
16. What would a gate on the median latency need, and why is it not enough?
17. How would you make the gate a routine step of continuous integration without making every change wait?
18. What kinds of regression can no benchmark of the engine alone catch?
19. Why keep the pools of runs as committed data?
20. In one sentence: what makes a performance gate trustworthy?

**Solution of Problem 25.1.**

1. About $208\,\mathrm{n}\mathrm{s}$ , with a robust coefficient of variation of about 14%.
2. About 18% ( $245\,\mathrm{n}\mathrm{s}$ against 208): the planted wait of $11.4\,\mathrm{n}\mathrm{s}$ cost more than itself, because the time stamp reads of the wait take time too.
3. It is set by the slowest 1% of the run’s events, those hit by interruptions and the system’s own work, so it depends on what the rest of the machine did during that run.
4. Its difference varies by about 20% from pair to pair, four times the rise to detect.
5. Medians ignore the runs spoiled by a burst of other work, which would pull a mean.
6. That, without a real difference, the labels A and B could be swapped (the interleaving makes this plausible); not that the runs are normal, or that A and B have the same spread.
7. 3 to 9% over a hundred gates, around the test’s level of 5% (a hundred gates estimate a rate of 5% to within a few points); the tolerance removes the significant rises below 2%.
8. 40.
9. 57%, 73% and 84%.
10. About 200 runs per side, where the rate reaches 91%.
11. About 220, close: the approximation assumes normal runs and uses the median’s efficiency for them; the measured runs are concentrated with a few far outliers, for which the median does a little better than that.
12. About $2 \times 200 \times 30\,\mathrm{m}\mathrm{s} \approx 12\,\mathrm{s}$ .
13. **Named result.** On this laptop, whose runs’ 99th percentiles vary by about 14% even when it is otherwise idle, the gate needs about 200 interleaved runs per side (about 12 seconds of benchmark) to detect a 5% rise with 90% power, at a false-alarm rate near the test’s 5% (the normal approximation says about 220); an 18% regression needs 40.
14. Less variation from run to run; the runs needed fall as the square of the spread, to about 27 at a spread of 5%.
15. With enough runs, any change becomes significant; the tolerance says which changes matter.
16. Far fewer runs, since medians are stable; but it misses what the 99th percentile catches: a rare slow path, such as an occasional allocation or a log line on one branch.
17. Run it on candidates for merge and nightly, on a dedicated machine, with the number of runs set by the machine’s measured noise; keep the functional tests on every change.
18. Changes in the network, the kernel, the interaction with other processes, the hardware, and data that differ from the benchmark’s day.
19. To recompute the verdict and the power analysis with other statistics or thresholds, and to compare machines.
20. Enough interleaved runs for the machine’s measured noise, judged by a test that assumes little, and a tolerance.

## 25.10 Interview questions

**Interview question 25.1 ★ developer.**

What is a [property-based test](#def-ll-testing-and-deploying-low-latency-systems-tests)? Give an example for an order book.

**Solution of Interview question 25.1.**

A test that generates random inputs and checks a property for all of them. For an order book: after any sequence of adds, modifications and cancels, the best bid is below the best ask, the quantity at each level equals the sum of its orders, and removing every order leaves an empty book.

*What the interviewer is looking for: a property, a generator, and shrinking of failures.*

**Interview question 25.2 ★★ developer.**

How would you fuzz a market-data parser, and what would you check besides crashes?

**Solution of Interview question 25.2.**

Mutate recorded packets (bytes, lengths, sequence numbers, truncations), feed them in exact-size buffers under sanitisers, and check invariants besides crashes: every accepted message within its buffer, sequence numbers consistent, the book valid after each message, and agreement with a reference parser.

*What the interviewer is looking for: sanitisers plus invariants and a differential reference.*

**Interview question 25.3 ★★ developer.**

Your benchmark says the new build is 3% slower at the 99th percentile. How do you decide whether to believe it?

**Solution of Interview question 25.3.**

Ask how it was measured: interleaved runs of both builds on the same machine, enough of them for its noise, a test with a confidence interval. Rerun with more runs if needed; 3% may be within the noise, or below the tolerance even if real.

*What the interviewer is looking for: interleaving, noise, a test and a tolerance.*

**Interview question 25.4 ★★ developer.**

What does an exchange’s certification test, and what does it not?

**Solution of Interview question 25.4.**

That the client’s messages and its handling of the venue’s responses follow the protocol, step by step, on the venue’s test system. Not the firm’s strategy, its risk controls beyond what the script exercises, its latency, or its behaviour under load and failures in production.

*What the interviewer is looking for: protocol conformance, not correctness or performance.*

**Interview question 25.5 ★★ developer, trader.**

How would you roll out a new version of a strategy engine, and how would you roll it back?

**Solution of Interview question 25.5.**

Preflight the host; run the new build in shadow on live inputs and compare its outputs; start it on a small share (a canary), with tight limits, then widen; roll back by redeploying the previous build, a procedure prepared and tested in advance, with the journal and positions carried over.

*What the interviewer is looking for: shadow, canary, prepared rollback.*

**Interview question 25.6 ★★★ developer.**

Design the test and release pipeline for a low-latency trading system, from a commit to production, and say what each stage catches.

**Solution of Interview question 25.6.**

On every change: unit, property and [differential tests](#def-ll-testing-and-deploying-low-latency-systems-tests), and fuzzing with sanitisers. On candidates: simulator scenarios and replays of recorded days, certification scripts, and the performance gate on a dedicated machine. On release: host preflight, shadow running, canary, staged rollout, prepared rollback. Each stage catches what the one before cannot see.

*What the interviewer is looking for: the pyramid, and what each layer catches.*
