---
title: "The Strategy Engine"
book: "Low-Latency Software"
subject: quant
language: en
chapter: 20
exercises: 8
source: https://one-course.com/books/quant/13/en/chapter/20-the-strategy-engine
---

# Chapter 20 — The Strategy Engine

Run the same quoting strategy twice on the same recorded day and compare what it sent. If one component stamps its orders with the wall clock, the two runs never agree, a hundred reruns give a hundred different outputs. If another keeps its quotes in a hash set whose seed is drawn at start-up and walks them in the set’s order, 93 reruns out of a hundred differ. If two threads feed its queue, 49 do. Remove all three and the hundred runs agree to the last bit, and from then on every bug seen in production can be replayed on a desk. This chapter builds the component that runs the firm’s strategies: a single thread that processes market events, order reports, timers and parameter changes one at a time, each to completion, in an order fixed by rules rather than by the operating system, with time injected rather than read, and with nothing on its hot path that can wait, allocate or vary.

## 20.1 The event loop

**Definition 20.1 (Event loop, run-to-completion processing).**

An *event loop* is a thread that repeatedly takes the next event from its inputs and hands it to the code that handles it. *Run-to-completion processing* handles each event entirely, including everything it causes, before the next event is taken: no handler is interrupted by another, so a handler never sees a state that another has half changed.

The engine’s inputs are the rings of chapter 12: market events from the [feed handler](https://one-course.com/books/quant/13/en/chapter/18-the-feed-handler#def-ll-the-feed-handler-handler) (through the [book builder](https://one-course.com/books/quant/13/en/chapter/19-the-order-book-builder#def-ll-the-order-book-builder-builder) of chapter 19), reports from the order gateway (chapter 21), and control messages; its own timers come from a timing wheel. Its outputs go to the gateway’s ring and to the log. One thread owns all the strategy’s state, so nothing needs a lock, and the order in which events are handled is decided by one rule: before an event with time $t$ is handled, every timer due at or before $t$ fires, in order of time and then of identifier; parameter changes effective at or before $t$ are applied; then the event. [Figure 20.1](#fig-ll-the-strategy-engine-loop) draws it, and [Listing 20.1](#lst-ll-the-strategy-engine-run) is the whole loop.

![The strategy engine. Every input reaches one thread, which handles each event to completion in an order fixed by rule: the timers due by the event’s time, then the parameters effective by it, then the event. Time is whatever clock the engine was given: a recorded day’s in replay, the machine’s live.](https://one-course.com/images/onecourse/chapters/quant-13/ll-the-strategy-engine/fig-acc5b9515049.svg)

***Figure 20.1.** The strategy engine. Every input reaches one thread, which handles each event to completion in an order fixed by rule: the timers due by the event’s time, then the parameters effective by it, then the event. Time is whatever clock the engine was given: a recorded day’s in replay, the machine’s live.*

```cpp
    void run(const std::vector<Event>& market, const std::vector<Snapshot>& changes = {}) {
        std::size_t c = 0;
        for (const auto& e : market) {
            // run to completion: everything due before this event happens first, in a fixed order
            wheel_.advance(e.ts, [&](const TimerWheel::Timer& t) { on_timer(t); });
            while (c < changes.size() && changes[c].t <= e.ts) p_ = changes[c++].p;
            on_market(e);
        }
    }
```

***Listing 20.1.** The event loop: timers first, then parameters, then the event, each handled to completion. code/firm/stratengine/cpp/firm_stratengine.hpp*

The design is old and deliberately plain. The retail exchange whose architecture Martin Fowler described in 2011 ran its business logic on one thread, in memory, with the state “entirely derivable by processing the input events” kept in a durable store; and because “the business logic is deterministic”, its outputs did not even need to be stored. A strategy engine gains the same two things: speed, because one thread with its data in its own caches does not coordinate with anyone, and replay, because the same inputs in the same order produce the same outputs.

## 20.2 Single-threaded determinism

A program is reproducible bit for bit (the bitwise reproducibility of One Quant Book 4, chapter 25) only as a whole, and any one of its parts can break it. [Figure 20.2](#fig-ll-the-strategy-engine-reruns) counts the damage on the build’s small quoting engine, run a hundred times as a new process on the same 6 000 recorded events, with one source of nondeterminism switched on at a time.

- *A read of the wall clock* , here to stamp each action and to date its quotes: every run reads different times, so a hundred runs give a hundred outputs. Any decision that compares the real time with something (a quote’s age, a timeout) makes the difference reach the orders themselves.
- *Iteration over a hash container with a per-process seed* : the engine keeps its ten quotes in a hash set whose hash is seeded at start-up and requotes them in the set’s order, so the identifiers and the order of the actions follow the seed. Rust’s standard hash map does exactly this by design (its algorithm “is randomly seeded”); 93 of the hundred runs differ.
- *Two threads feeding one queue* , each carrying half of the events: the order in which the consumer sees them depends on scheduling, the book passes through states that never existed, and 49 of the hundred runs differ.

Switched off, all three give one output in a hundred runs; so does the firm’s engine. The rules that achieve it are few. Time comes from the events, through an [injected clock](#def-ll-the-strategy-engine-wheel) (next section). Containers that are iterated have a defined order: arrays, ordered maps, or hash maps whose order does not decide anything. One thread owns the state; other threads only move data through rings, and the merge of several inputs follows a rule (by time, then by source, then by sequence), not the arrival order. Floating-point sums are done in a fixed order (chapter 8), and the example strategy stays in integers, which is also why the C++ and Rust engines of the build produce the same output hash.

![Distinct output hashes in a hundred reruns (each a new process) of the same 6 000 recorded events, for the firm’s engine and for the small engine of ll_nondet.hpp with each source of nondeterminism switched on. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2. Data: bench_engine.py.](https://one-course.com/images/onecourse/chapters/quant-13/ll-the-strategy-engine/fig-1d834a956604.svg)

***Figure 20.2.** Distinct output hashes in a hundred reruns (each a new process) of the same 6 000 recorded events, for the firm’s engine and for the small engine of `ll_nondet.hpp` with each source of nondeterminism switched on. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2. Data: `bench_engine.py`.*

## 20.3 Timers and the injected clock

**Definition 20.2 (Timer wheel, injected clock).**

A *timer wheel* is a circular array of slots, each covering an interval of time (its granularity): a timer is filed in the slot of its expiry time modulo the wheel’s span, and advancing the wheel visits the slots of the time that has passed and fires the timers due in them; scheduling costs a constant time whatever the number of timers. An *injected clock* is a clock passed to the code that needs the time instead of being read by it: the same code runs on the simulated clock of a recorded day (One Quant Book 7, chapter 17) or on the machine’s real clock.

A strategy is full of timers: an order that must be acknowledged within a delay, a quote that must be pulled if the market data stops, a refresh every few seconds, a [heartbeat](https://one-course.com/books/quant/13/en/chapter/15-protocols-i-fix#def-ll-protocols-i-fix-session). A priority queue would do, at a logarithmic cost per timer; the hashed timing wheel of Varghese and Lauck (1987) does it in constant time, which matters when every order schedules and cancels timers. The build’s wheel has 1 024 slots of $100\,\text{µ}\mathrm{s}$: a timer goes to slot $\lfloor t/100\,\mu\text{s}\rfloor \bmod 1\,024$, and a timer further away than the wheel’s span of about $0.1\,\mathrm{s}$ simply stays in its slot until its time comes round. Two details make it deterministic: the timers due in a slot fire sorted by time and then by identifier, never in insertion order; and a timer scheduled while another fires, if due in the same slot, fires in the same pass.

![A hashed timing wheel of 16 slots (the build’s has 1 024 of 100\, µ s). Scheduling puts a timer in the slot of its expiry time; advancing visits the slots passed and fires only the timers due, sorted; a timer further than one revolution stays in its slot until its time comes round.](https://one-course.com/images/onecourse/chapters/quant-13/ll-the-strategy-engine/fig-54b106141414.svg)

***Figure 20.3.** A hashed timing wheel of 16 slots (the build’s has 1 024 of $100\,\text{µ}\mathrm{s}$). Scheduling puts a timer in the slot of its expiry time; advancing visits the slots passed and fires only the timers due, sorted; a timer further than one revolution stays in its slot until its time comes round.*

```cpp
    void advance(std::uint64_t to, F&& fire) {
        if (!started_) { tick_ = to / gran_; started_ = true; return; }
        const std::uint64_t last = to / gran_;
        for (; tick_ <= last; ++tick_) {
            if (pending_ == 0) { tick_ = last; break; }
            auto& s = slots_[tick_ & mask_];
            for (;;) {   // timers scheduled while firing may land in this slot again
                due_.clear();
                const std::uint64_t limit = std::min(to, (tick_ + 1) * gran_ - 1);
                for (std::size_t i = 0; i < s.size();)
                    if (s[i].t <= limit) { due_.push_back(s[i]); s[i] = s.back(); s.pop_back(); }
                    else ++i;
                if (due_.empty()) break;
                std::sort(due_.begin(), due_.end(), [](const Timer& a, const Timer& b) { return a.t != b.t ? a.t < b.t : a.id < b.id; });
                pending_ -= due_.size();
                for (const auto& d : due_) fire(d);
            }
            if (tick_ == last) break;
        }
    }
```

***Listing 20.2.** Advancing the wheel: visit the slots passed, fire the timers due in time and identifier order. code/firm/stratengine/cpp/firm_stratengine.hpp*

The wheel knows no clock. It is advanced to the time of each event, and the engine’s code asks an [injected clock](#def-ll-the-strategy-engine-wheel) for the time: the simulated clock returns the event’s time, the real clock reads `CLOCK_MONOTONIC` (chapter 5) and maps it onto the day. Replaying a day with the simulated clock gives, event for event, what the live engine would have done with the same inputs at the same times; the real clock is used live, and only there.

```cpp
struct SimClock {
    std::uint64_t now(std::uint64_t event_t) const { return event_t; }
};
struct RealClock {
    std::uint64_t t0_event = 0, t0_real = 0;
    std::uint64_t now(std::uint64_t event_t) {
        timespec ts{};
        clock_gettime(CLOCK_MONOTONIC, &ts);
        const std::uint64_t r = static_cast<std::uint64_t>(ts.tv_sec) * 1'000'000'000ULL + static_cast<std::uint64_t>(ts.tv_nsec);
        if (t0_real == 0) { t0_real = r; t0_event = event_t; }
        return t0_event + (r - t0_real);
    }
};
```

***Listing 20.3.** The two clocks the engine can be given. code/firm/stratengine/cpp/firm_stratengine.hpp*

## 20.4 Configuration and parameters at run time

**Definition 20.3 (Parameter snapshot).**

A *parameter snapshot* is a complete, versioned set of a strategy’s parameters that takes effect at a stated time, between two events and never during one; every action the strategy takes records the version of the parameters it was taken under.

Traders change parameters during the day: a wider spread before a number, a smaller size in a thin market. Changing a field of a live structure from another thread while the engine reads it is a [data race](https://one-course.com/books/quant/13/en/chapter/9-rust-for-low-latency#def-ll-rust-for-low-latency-race) (chapter 11), and even a correct atomic update of one field leaves a moment where half the parameters are new. The build applies whole snapshots, in the loop, between events: the control path publishes a snapshot with an effective time on the engine’s control ring, and the loop installs it before the first event at or after that time. The version stamped on every action says which parameters produced it, and in replay the snapshots are inputs like any other event: the build’s test applies one half-way through the day and checks that every action before it is unchanged and every later one carries the new version.

## 20.5 Hot-path discipline

The hot path of chapter 1 runs through this thread, and it is disciplined by rules the earlier chapters have measured. No allocation per event (chapter 6): the build’s test counts heap allocations while the engine replays its fixture and finds none, because the action log, the wheel’s slots and the book are sized at start-up. No system call: the simulated clock is a field, and even the real clock is read through the vDSO (chapter 5). No lock and no waiting: inputs and outputs are rings, and a full output ring is an error to report, not a reason to wait. No formatting of text: the log receives binary records (chapter 23). No unbounded work in a handler: the wheel’s work is bounded by the slots passed, the book’s by its structure.

The result on the laptop: the engine takes about $53\,\mathrm{n}\mathrm{s}$ per event at the median on the large-tick stream of chapter 18, book update included, and about $57\,\mathrm{n}\mathrm{s}$ on the small-tick one where it requotes at almost every event; the p99 is about $130\,\mathrm{n}\mathrm{s}$ and $200\,\mathrm{n}\mathrm{s}$. A decision in tens of nanoseconds leaves the [latency budget](https://one-course.com/books/quant/13/en/chapter/1-where-latency-comes-from#def-ll-where-latency-comes-from-budget) to the network and the gateway.

## 20.6 Tutorial: replay a day, then break it

**Goal.** Replay a recorded input through the engine to the same output every time, then switch on each source of nondeterminism and count the outcomes. **End state:** [Figure 20.2](#fig-ll-the-strategy-engine-reruns), the engine’s time per event, and green tests in C++ and Rust.

1. **Replay.** The C++ test runs the engine twice on the [book builder](https://one-course.com/books/quant/13/en/chapter/19-the-order-book-builder#def-ll-the-order-book-builder-builder) ’s small-tick fixture and compares the output hashes, the actions and the profit, then checks them against `data/expected.txt` ; the Rust engine, written against the same rules, reproduces the same hash.
2. **Change the parameters half-way** with a snapshot and check that only the actions after it change.
3. **Count the zero:** the test counts heap allocations during the replay.
4. **Break it** with `python bench_engine.py` : a hundred new processes per configuration of `ll_nondet` , each printing its output hash, then the distinct hashes counted.

**What to change next.** Replace the small engine’s seeded hash set by one with a fixed seed and count again; feed the two threads’ events through a merge by timestamp instead of a shared queue.

## 20.7 Build: the strategy engine

**Purpose.** The thread on which the firm’s strategies run: it consumes the [book builder](https://one-course.com/books/quant/13/en/chapter/19-the-order-book-builder#def-ll-the-order-book-builder-builder)’s view (chapter 19), sends orders through the gateway (chapter 21) past the risk gate (chapter 22), and logs everything it does (chapter 23).

**Interface.** C++20 `firm::strat`: `Engine<Clock>(tick, Params, clock)` with `run(events, snapshots)`, `actions`, `hash`, `position`, `cash`, `pnl()`; `TimerWheel(granularity, slots)` with `schedule` and `advance`; `SimClock`, `RealClock`; `Params`, `Snapshot`, `Action`. Rust `firm_stratengine`: `Engine`, `TimerWheel`, `Params`, `Action`, `read_events`.

**Rules.** One thread owns the state; each event is handled to completion; timers due by an event fire before it, in (time, identifier) order; parameter snapshots apply between events; time comes from the [injected clock](#def-ll-the-strategy-engine-wheel); the example strategy computes in integers; no allocation per event.

**Acceptance tests.** `code/firm/stratengine/`: two replays with identical actions, hash and profit, equal to the committed hash; the Rust engine equal to it; a snapshot changing only what follows it; zero allocations during the replay; the wheel firing in order across slots and revolutions.

**Stretch.** A hierarchical wheel for long timers; order reports from the simulator of One Quant Book 10 in place of the venue stub; several strategies on one engine with a fixed scheduling order.

Sources and further reading

- G. Varghese and T. Lauck, “Hashed and hierarchical timing wheels: data structures for the efficient implementation of a timer facility”, *Proceedings of the 11th ACM Symposium on Operating Systems Principles* , 1987.
- M. Fowler, *The LMAX Architecture* , 2011.
- The Rust standard library, `std::collections::HashMap` .

## 20.8 Exercises

**Exercise 20.1 ★.**

A wheel has 1 024 slots of $100\,\text{µ}\mathrm{s}$. In which slot does a timer due at 34 200.012345 seconds after midnight go, and how many revolutions away is a timer due in 5 seconds?

**Solution of Exercise 20.1.**

$34\,200.012345~\text{s}$ is $342\,000\,123$ intervals of $100\,\text{µ}\mathrm{s}$, and $342\,000\,123 \bmod 1\,024 = 507$. The wheel spans $1\,024 \times 100\,\text{µ}\mathrm{s} = 102.4\,\mathrm{m}\mathrm{s}$, so a timer 5 seconds away is about 48.8 revolutions away: it waits in its slot while the wheel passes it 48 times.

**Exercise 20.2 ★.**

Three timers are due in the same slot at 120, 121 and 120 microseconds with identifiers 7, 3 and 9. In what order do they fire?

**Solution of Exercise 20.2.**

By time, then identifier: (120, 7), (120, 9), (121, 3).

**Exercise 20.3 ★.**

Why is the time of an event the right “now” for a strategy in replay, and what must be true of the live system for replay to match it?

**Solution of Exercise 20.3.**

In replay the event’s time is when the live engine saw it, so decisions depend on the same numbers. It matches the live run only if the live engine also took its time from the events (or from a clock whose readings are recorded with the inputs), and if the recorded inputs are exactly what the live engine received, in the same order.

**Exercise 20.4 ★★.**

If an outcome depends on which of $k$ equally likely orders a hash set produces, how many distinct outcomes should a hundred reruns show on average? What does the measured 93 suggest about $k$?

**Solution of Exercise 20.4.**

$k\,(1 - (1 - 1/k)^{100})$ on average. It equals 93 for $k \approx 700$: the outcome depends on several hundred possible orders (fewer than a hundred would cap the count at $k$; unequal probabilities need more).

**Exercise 20.5 ★★.**

A trader changes the spread from two ticks to four at 10:30:00.000. An event arrives at 10:29:59.999 and the next at 10:30:00.004. Under which parameters is each handled, and what does the log show?

**Solution of Exercise 20.5.**

The first under the old parameters (two ticks), the second under the new (four): the snapshot is installed before the first event at or after its effective time. The log shows the version with every action: version 1 before, version 2 after.

**Exercise 20.6 ★★.**

The engine takes $53\,\mathrm{n}\mathrm{s}$ per event at the median. At the open, events arrive at 2 million a second for 50 milliseconds. Does it keep up, and what is its utilisation?

**Solution of Exercise 20.6.**

Yes: $2 \times 10^6 \times 53\,\mathrm{n}\mathrm{s} \approx 0.11$, a utilisation of about 11%, so no queue builds up (the tail of $130\,\mathrm{n}\mathrm{s}$ at the p99 does not change that).

**Exercise 20.7 ★★★.**

*Coding.* Add a refresh timer to the engine: every quote older than $5\,\mathrm{s}$ is re-sent at the same price. Check that two replays still agree and that the Rust engine, changed the same way, still produces the C++ hash.

**Solution of Exercise 20.7.**

Schedule a refresh timer with each quote, due $5\,\mathrm{s}$ later, with the quote’s identifier in the timer’s; when it fires and the quote is unchanged, emit a replace at the same price and schedule the next refresh. Both engines must use the same identifiers and the same ordering rule for the replay hashes to agree, and the committed hash is regenerated once from the C++ engine.

**Exercise 20.8 ★★★.**

*Find the flaw.* “Our engine is deterministic: it is single-threaded and uses no randomness. It stamps each order with `std::chrono::system_clock::now()` for the audit trail.”

**Solution of Exercise 20.8.**

The stamp differs on every run, so the output is never bitwise reproducible, and any code that ever compares it with anything (a timeout, an age) makes the orders differ too; the system clock can also jump when it is corrected. Take the audit timestamp from the event’s time in the engine, or add it outside the engine where messages leave the machine (chapter 23), and keep it out of the replayed comparison.

## 20.9 Problem: Two Runs, Two Profits

**Problem 20.1.**

Weekend problem — counting the ways a program can disagree with itself

A research team replays the same recorded day through its quoting engine twice and gets two profits. Use the measurements of [Figure 20.2](#fig-ll-the-strategy-engine-reruns) (`measured_nondet.csv`).

**Part I — The sources.**

1. List the three sources measured and the number of distinct outputs each gave in a hundred runs.
2. Why does the wall clock give a hundred?
3. Why does the seeded hash set give fewer than a hundred?
4. Why does the two-thread feed give fewer still, and what would make it give more?

**Part II — Finding them.**

5. How would you find the first action at which two runs differ?
6. What does each source leave as a clue at that point?
7. Why must each rerun be a new process to see the seeded hash’s effect?
8. Which of the three sources would a single run inside a debugger never show?

**Part III — The fixes.**

9. Fix the wall clock.
10. Fix the hash order without giving up hash maps.
11. Fix the two-thread feed.
12. After the fixes, how many distinct outputs should a hundred reruns give, and how many did the firm’s engine give?

**Part IV — The verdict.**

13. State the *named result* : the number of distinct outcomes in a hundred reruns with each source switched on, and after the fixes.
14. Why is “the profit differs by a little” not a reason to accept nondeterminism in a backtest?
15. What else can break bitwise reproducibility between two machines, even with all three fixes?
16. Why does integer arithmetic in the strategy make the C++ and Rust engines agree?
17. What should the continuous-integration pipeline run to keep determinism?
18. What does determinism give the production support team?
19. When is a nondeterministic component acceptable in a trading system?
20. In one sentence: what makes an engine replayable?

**Solution of Problem 20.1.**

1. A wall-clock read: 100 distinct outputs; a seeded hash order: 93; two threads on one queue: 49.
2. Every run reads different times, and the times are in the output.
3. The outputs depend on which order the set yields, and the seed maps many seeds to the same order: several hundred orders, repeated now and then in a hundred runs.
4. Many interleavings end in the same output when the threads happen to alternate as they did before; more contention or slower consumers would give more.
5. Log the output of both runs with sequence numbers and compare them record by record (or bisect on the hash of prefixes).
6. The wall clock: identical decisions with different time stamps; the hash order: the same actions in a different order or with different identifiers; the threads: a book state that the input never produced.
7. The seed is drawn at start-up, once per process: reruns inside one process share it.
8. The thread interleaving, which a debugger changes by stopping threads.
9. Take time from the events through an [injected clock](#def-ll-the-strategy-engine-wheel) , and stamp outside the engine.
10. Keep hash maps for lookups only, and iterate in a defined order (an array of the keys, an ordered map, or a sort) when the order decides anything.
11. Merge the two inputs by a rule (time, then source, then sequence) in one thread before the engine, or give each input its own ring and let the engine merge.
12. One; the firm’s engine gave one.
13. **Named result.** In a hundred reruns of the same 6 000 events: 100 distinct outcomes with a wall-clock read, 93 with a seeded hash order, 49 with two threads on one queue, 100 with all three, and one after the fixes (and one for the firm’s engine).
14. A backtest that is not reproducible cannot be debugged or compared: a small difference can hide a large bug, and two versions of a strategy cannot be told apart from noise.
15. Floating-point reductions in a different order, different compiler flags or libraries (chapter 8), different input data (a missed packet in a recording), and uninitialised memory.
16. Integer operations are exact and defined the same way in both languages, while floating-point results depend on the order and contraction of operations.
17. Replay a recorded day twice and compare hashes; replay it on each build of the engine and compare with the committed hash; run the replay under a sanitiser.
18. Any incident can be replayed exactly on a desk, with a debugger, as often as needed.
19. Outside the decision path: monitoring, logging of timings, operational tools whose outputs are not part of the replayed state.
20. One thread, ordered inputs, time taken from the events, and nothing iterated in an order that varies.

## 20.10 Interview questions

**Interview question 20.1 ★ developer.**

Why run a strategy on a single thread? What does it cost?

**Solution of Interview question 20.1.**

No locks, no coordination, the state in one core’s caches, and a single order of events that makes the engine deterministic and replayable. The cost: one core’s throughput is the limit, a slow handler delays every other event, and work that could run in parallel (analytics, logging) must be moved to other threads through rings.

*What the interviewer is looking for: determinism and locality against a single core’s capacity.*

**Interview question 20.2 ★★ developer.**

How would you implement timers for thousands of orders, each with its own timeouts?

**Solution of Interview question 20.2.**

A hashed timing wheel: constant-time scheduling and cancellation, slots visited as time advances, a hierarchical wheel or revolution counts for long timers; fire due timers in a defined order (time, then identifier).

*What the interviewer is looking for: the wheel, and ordering for determinism.*

**Interview question 20.3 ★★ developer.**

Your backtest and your production run disagree on the same day. Where do you look?

**Solution of Interview question 20.3.**

Inputs first (are the recorded events what production received, in the same order?), then time (does any code read the clock instead of the event’s time?), then order (threads, hash iteration), then numerics (floating-point order, flags), then parameters (the versions in force). Compare the two outputs record by record to find the first difference.

*What the interviewer is looking for: inputs, time, order, numerics, parameters.*

**Interview question 20.4 ★★ developer.**

How do you change a strategy’s parameters while it runs, safely?

**Solution of Interview question 20.4.**

Publish complete, versioned snapshots on a control ring with an effective time; the engine installs one between events, never during one, and stamps every action with the version in force; in replay the snapshots are inputs.

*What the interviewer is looking for: whole snapshots between events, and versioning.*

**Interview question 20.5 ★★ developer.**

What must never happen on the hot path of a strategy engine, and how do you check?

**Solution of Interview question 20.5.**

Allocation, system calls, locks, waiting, text formatting, unbounded loops and anything that varies from run to run. Check with allocation counters in tests, tracing of system calls, replay hashes, and measured tail latencies per event.

*What the interviewer is looking for: a list and a way to check each item.*

**Interview question 20.6 ★★★ developer, researcher.**

Design a strategy engine that is both live and replayable bit for bit. What are the rules, and how are they enforced?

**Solution of Interview question 20.6.**

One thread per engine fed by rings; inputs merged by a rule and recorded; time injected (simulated in replay, real live); timers on a wheel with ordered firing; parameters as versioned snapshots between events; no hash-order or floating-point dependence in decisions; enforced by replay tests with committed hashes, in continuous integration and after every release.

*What the interviewer is looking for: rules plus tests that enforce them.*
