Quantitative Finance · Book 13 · Technology

Low-Latency Software

Low-Latency Software · Technology

18The Feed Handler

At the open the message rate jumps: on the simulator’s busy instrument of this chapter, from about 570 messages a second at midday to more than 9 000 in the first seconds; on the consolidated feed of U.S. options, OPRA published a peak of 63.9 million messages in one second of July 2026 and 350 000 in its busiest millisecond. A handler that falls behind lets its socket buffer fill; the kernel drops what does not fit; line A loses a packet that line B has also lost; and the firm’s book is stale until the handler has recovered, by retransmission if the gap is small and recent, from the next snapshot if not. This chapter builds the first component of Part IV’s trading system, the feed handler: it decodes and arbitrates the two lines of chapter 16, detects gaps and recovers from them, normalises every message into the firm’s event record and publishes it on the rings of chapter 12, and says at every moment whether the book it feeds can be trusted.

18.1 Decode and arbitrate

Definition 18.1 (Feed handler)

A feed handler is the process that receives a venue’s market-data feed (One Quant Book 1, chapter 4) and turns it into the firm’s internal events: it reads the redundant lines, arbitrates them into one sequenced stream, detects and recovers lost messages, normalises each message into the firm’s own record (Book 1, chapter 28), and publishes the records to the components that build books and make decisions.

The build’s feed handler. Packets from both lines are merged by sequence number; a gap holds later messages and waits briefly for the other line, then asks the retransmission server, then falls back to the next snapshot; every message leaves as one fixed-size event on the firm’s ring. While a gap is open, the stale-book flag is up.
Figure 18.1. The build’s feed handler. Packets from both lines are merged by sequence number; a gap holds later messages and waits briefly for the other line, then asks the retransmission server, then falls back to the next snapshot; every message leaves as one fixed-size event on the firm’s ring. While a gap is open, the stale-book flag is up.

The handler of the build is written first in Python, as the reference, then in C++20 and Rust, and all three publish exactly the same events (compared by a 64-bit hash of their 48-byte records) on the same inputs. Its time is the arrival time recorded with each packet, so that a recorded day replays exactly: the same inputs give the same events, the same staleness intervals and the same counters. Arbitration is chapter 16’s; what the handler adds is what to do while a sequence number is missing.

18.2 Detect gaps and recover

A missing sequence number is usually not lost, only late: it is on its way on the other line, or reordered by a switch. The handler therefore holds the messages that arrive after it and waits a short time, here 0.5 ms0.5\,\mathrm{m}\mathrm{s}, for the missing ones. If they come, the held messages are released in order. If not, the handler asks the retransmission server for the range (chapter 16; the simulator answers up to 1 000 messages per request from its last 100 000), and applies the answer when it arrives, a round trip later. If the server cannot help, because the gap is older than its window or the server is unreachable, the handler waits for the next snapshot of the instrument, loads it, discards the held messages it already reflects and continues from the sequence number it gives, the procedure CME’s snapshot messages support with their last incremental sequence number processed.

    void incremental(std::uint64_t t, const std::uint8_t* p) {
        ++c.packets;
        const std::uint64_t seq = be64(p + 10), cnt = be16(p + 18);
        if (cnt != 0 && cnt != 0xFFFF && seq + cnt <= next_) {
            ++c.duplicates;
            return;
        }
        blocks(p, [&](std::uint64_t s, const std::uint8_t* m) {
            if (s < next_) return;
            if (s == next_ && mode == Mode::live) {
                publish(normalise(m, s));
                ++c.messages;
                ++next_;
            } else {
                pending_.emplace(s, m);
            }
        });
        if (!pending_.empty() && mode == Mode::live) {
            mode = Mode::await_line;
            since_ = t;
            deadline_ = t + timeout_;
            ++c.gaps;
        }
        drain(t);
    }
Listing 18.1. Each packet: drop it if every message is old, deliver the next expected message, hold the others; a held message opens a gap and starts its timer. code/firm/feedhandler/cpp/firm_feedhandler.hpp

The snapshot path is the delicate one: the snapshot must be at least as recent as the book, the held messages it already reflects must go, and a hole after it must be detected again.

    def _apply_snapshot(self, t):
        upto = be(self._snap[0], 11, 8)
        if upto + 1 < self.next:               # older than what the book already holds: wait for the next cycle
            self._snap = []
            return
        for m in self._snap:
            self._publish(normalise(m, upto))
        self._snap = []
        self.next = upto + 1
        for s in [s for s in self.pending if s <= upto]:
            del self.pending[s]
        self._drain(t)
        if self.pending and self.mode != "live":             # a hole after the snapshot: detect it again
            self.mode, self._deadline = "await_line", t + self.gap_timeout_ns
Listing 18.2. Loading a snapshot in the Python reference. code/firm/feedhandler/firm_feedhandler.py

Definition 18.2 (Stale-book flag, recovery latency)

The stale-book flag is the handler’s public statement that the book built from its events may differ from the venue’s, raised when a gap is detected and lowered when it is closed. The recovery latency of a gap is the time from its detection to the moment the book is correct again: the other line’s delay, a retransmission round trip, or the wait for and the loading of a snapshot.

Figure 18.2 shows the handler at work on a simulated minute from the open, both lines impaired: 0.2% of packets lost independently on each line, bursts of loss, jitter that reorders packets, and an outage on each line overlapping for 15 milliseconds at 40 seconds. Out of 4 232 gaps, 4 224 are filled by the other line within the half millisecond; the retransmission server closes the other eight within 2.7 ms2.7\,\mathrm{m}\mathrm{s}. Without the server, seven gaps need a snapshot, and with a snapshot every second the book stays stale for 0.3 to 1 second each time, five seconds of the minute in all. Staleness is then not a property of the network but of the recovery path.

A simulated minute from the open. Top: messages published per 100 ms, sixteen times busier at the open than at midday. Bottom: the gaps that the other line could not fill, with the retransmission server (each closed within 3\, m s: a tick) and with snapshot recovery only (a snapshot a second: a bar from detection to recovery); the 4 224 gaps filled by the other line are too short to draw. Data: fig_stale.py (deterministic: the simulator and the handler are).
Figure 18.2. A simulated minute from the open. Top: messages published per 100 ms, sixteen times busier at the open than at midday. Bottom: the gaps that the other line could not fill, with the retransmission server (each closed within 3 ms3\,\mathrm{m}\mathrm{s}: a tick) and with snapshot recovery only (a snapshot a second: a bar from detection to recovery); the 4 224 gaps filled by the other line are too short to draw. Data: fig_stale.py (deterministic: the simulator and the handler are).

18.3 Normalise and publish

Every message leaves the handler as one 48-byte event: its kind, side, instrument, sequence number, timestamp, order reference, second reference (the new reference of a replace, the match number of an execution, the sequence number of a snapshot), price and quantity. A normalised record has three virtues downstream. Its size is fixed, so it fits a ring slot (chapter 12) and a consumer reads it without decoding. It is the same for every venue the firm reads, so the book builder (chapter 19) and the strategies (chapter 20) are written once. And it carries the venue’s sequence number, so a consumer can tell exactly which events it has seen, and a snapshot’s events can say which sequence number they reflect. The build’s test publishes a whole run on an SPSC ring of 64-byte slots and reads it back, hash for hash.

On the laptop the C++ handler takes about 39 ns39\,\mathrm{n}\mathrm{s} per packet at the median, 160 ns160\,\mathrm{n}\mathrm{s} at the p99 and 290 ns290\,\mathrm{n}\mathrm{s} at the p99.9, measured over the minute of Figure 18.2 with the clock read around each input; the rare slow ones are gaps, whose held messages go through an ordered map, and retransmission answers. At 9 000 messages a second the handler is busy about 0.04% of the time: the burst of this simulated instrument is nothing for it. The real bursts are the ones of Box 18.1.

As of September 2026 — How bursty is a feed?

OPRA’s published operating metrics for the consolidated U.S. options feed (August 2026 report) give, for July 2026, a peak of 63.9 million messages in one second, 7.9 million in 100 milliseconds, 0.89 million in 10 milliseconds and 0.35 million in one millisecond: the busiest millisecond ran at 350 million a second, five and a half times the busiest second’s rate. OPRA’s capacity notice says why subscribers plan on 10-millisecond peaks (they “reflect system utilization during bursts of traffic”), and that taking both redundant streams doubles the bandwidth.

18.4 Bursts, buffers and slow consumers

Definition 18.3 (Message burst, receive buffer overrun)

A message burst is a period in which a feed’s rate exceeds the rate its receiver sustains. A receive buffer overrun is the loss of datagrams that arrive while the socket’s receive buffer is full: the kernel drops them silently, and the handler sees only a gap.

During a burst the difference between the arrival rate and the handler’s service rate accumulates in the socket’s receive buffer, the M/M/1 queue of One Quant Book 4, chapter 8, with a hard wall: whatever does not fit is dropped. The buffer’s size is set with SO_RCVBUF, and two details of Linux decide what it holds. The kernel “doubles this value (to allow space for bookkeeping overhead)” and returns the doubled value (socket(7)), capped by net.core.rmem_max, 4 MiB on the laptop, so that asking for 4 MiB grants 8. And the buffer is charged with each datagram’s whole memory footprint, not its payload. Figure 18.3 measures what a stalled reader keeps: 50 000 datagrams sent on loopback while nobody reads, then the socket drained. Each 64-byte datagram costs 960 bytes of buffer and each 1 000-byte datagram 2 305: the largest buffer the laptop grants holds 8 738 small datagrams, a fortieth of the messages of OPRA’s busiest millisecond.

How many datagrams a UDP socket keeps for a reader that is not reading, by receive-buffer size: 50 000 datagrams sent on loopback during the stall. The kernel grants twice the size asked and charges each datagram its full memory footprint. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2 (kernel 6.18); rmem_max 4 MiB. Data: bench_feed.py.
Figure 18.3. How many datagrams a UDP socket keeps for a reader that is not reading, by receive-buffer size: 50 000 datagrams sent on loopback during the stall. The kernel grants twice the size asked and charges each datagram its full memory footprint. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2 (kernel 6.18); rmem_max 4 MiB. Data: bench_feed.py.

The handler is in turn a producer. Its consumers read the ring at their own pace, and a consumer slower than the feed (a slow consumer, chapter 12) must not slow the handler down: the broadcast ring never waits, and a lapped consumer resynchronises. The two defences against bursts are therefore the same at both ends of the handler: enough buffer for the burst the data says to expect, and a way to recover when even that is not enough.

18.5 Monitoring a feed

A handler that is wrong silently is worse than one that stops. The build counts packets, messages, duplicates, gaps, gaps filled by the other line, retransmission requests and snapshot recoveries, and keeps every staleness interval with its start and end. Four of these numbers deserve alarms: the rate of gaps on each line (a line going bad, while the other hides it), the gaps filled by neither line (a common cause, chapter 16), any snapshot recovery (a book that was wrong for up to a snapshot interval), and the receive queue’s high-water mark (a burst closer to the buffer than expected). The stale-book flag itself goes to every consumer with each event, so that a strategy never quotes on a book that its handler knows to be wrong.

18.6 Tutorial: the handler on impaired lines

Goal. Run the handler in three languages on the same impaired lines, recover by retransmission and by snapshot, and check the output against a clean run. End state: Figures 18.2 and 18.3, the handling-time quantiles, and green tests in Python, C++ and Rust.

  1. Make the fixture. make_feed_fixtures.py runs Book 10’s simulator for three seconds with both lines impaired and writes the lines, the lossless stream (which the retransmission server holds) and the snapshot channel.
  2. Three runs, three languages. Clean, impaired with the server, impaired without it: the Python reference writes the expected counts, hashes and staleness, and the C++ and Rust tests reproduce them exactly.
  3. Check what recovery means. With retransmission, the impaired run publishes exactly the clean run’s events. With a snapshot, it publishes the snapshot instead of the lost messages, then exactly the clean run’s events after it, and the book at the end of the run is the same.
  4. Measure with python bench_feed.py: the minute of Figure 18.2 through the C++ handler with the clock read around each input (the script checks that it publishes what the reference publishes), and the stalled-reader experiment of Figure 18.3.

What to change next. Shorten the wait for the other line to 50 µs50\,\text{µ}\mathrm{s} and count the retransmission requests that follow; lengthen the snapshot interval to five seconds and read the staleness again.

18.7 Build: the feed handler

Purpose. The first stage of the trading system: from the venue’s lines to the firm’s events. The book builder of chapter 19 consumes its events, the strategy engine of chapter 20 its stale-book flag, and the monitoring of chapter 24 its counters.

Interface. Python reference firm_feedhandler: Handler(retx, gap_timeout_ns, rtt_ns).run(events), .events, .hash, .stale, .counters; RetxServer, merge, recorded, normalise, book. C++20 firm::feed2::Handler(retx, timeout, rtt) with run(a, b, snapshots), on_event, hash, stale, counters and optional timing hooks; RetxServer; recorded. Rust firm_feedhandler::Handler with the same.

Rules. Each sequence number is published once, in order; held messages are released only in order; a gap waits for the other line, then retransmission (one request per hole, answers capped by the server), then a snapshot; the handler is a function of its timestamped inputs; the stale-book flag is up exactly while a gap is open; no allocation per packet without a gap.

Acceptance tests. code/firm/feedhandler/: on the shared fixture, clean, retransmission and snapshot runs with identical event counts, hashes, staleness intervals and counters in Python, C++ and Rust; the retransmission run equal to the clean run event for event; the snapshot run equal to the clean run after the snapshot and in its final book; publication on firm.ring read back hash for hash; arbitration on scripted packets.

Stretch. One request for all the holes of a gap; per-instrument recovery on a multi-instrument feed; reading the lines from sockets (chapter 13’s busy polling) with the receive queue’s high-water mark as a counter.

Sources and further reading

  • Linux man page socket(7) (SO_RCVBUF); Linux kernel documentation, “Documentation for /proc/sys/net/” (rmem_max, rmem_default, netdev_max_backlog).
  • OPRA, Key Operating Metrics of U.S. Options Securities Information Processor (August 2026); SIAC, Revised OPRA Capacity Projections (September 2025).
  • CME Group Client Systems Wiki, MDP 3.0: “Market Data Snapshot – Full Recovery”.
  • The simulator’s protocols and services, code/firm/exchsim/PROTOCOL.md (One Quant Book 10, chapter 26).

18.8 Exercises

Exercise 18.1 ★

The handler expects sequence number 1 000 and receives packets starting at 1 003 (3 messages) on line A, then 1 000 (3 messages) on line B. What does it publish, in what order, and what is the staleness interval if the second packet arrives 40 µs40\,\text{µ}\mathrm{s} after the first?

Solution

Solution of Exercise 18.1.

Line A’s packet holds 1 003–1 005 and opens a gap at 1 000; line B’s brings 1 000–1 002, which are published at once, followed by the held 1 003–1 005: six events in sequence order. The book was stale from the first packet’s arrival to the second’s, 40 µs40\,\text{µ}\mathrm{s}, and the gap counts as filled by the other line.

Exercise 18.2 ★

A handler asks for a 256 KiB receive buffer. What does getsockopt return, and how many 64-byte datagrams will the buffer hold, by the measurement of Figure 18.3?

Solution

Solution of Exercise 18.2.

512 KiB (the kernel doubles the value asked). At 960 bytes per 64-byte datagram it holds 546 of them, as measured.

Exercise 18.3 ★

With snapshots every 500 milliseconds and a snapshot that takes 5 milliseconds to receive, what are the mean and worst recovery latencies of an unrecoverable gap?

Solution

Solution of Exercise 18.3.

The next snapshot comes on average half an interval later and at worst a full interval: 250+5=255 ms250 + 5 = 255\,\mathrm{m}\mathrm{s} on average, 500+5=505 ms500 + 5 = 505\,\mathrm{m}\mathrm{s} at worst.

Exercise 18.4 ★★

Why does the handler wait for the other line before asking for a retransmission? What does a shorter wait cost, and a longer one?

Solution

Solution of Exercise 18.4.

Most gaps are messages that the other line delivers a few microseconds later or that a switch reordered: in the chapter’s minute, 4 224 of 4 232. Asking at once would load the server and the network with requests for messages already on their way. A shorter wait sends more useless requests; a longer one leaves the book stale longer whenever a message really is lost on both lines.

Exercise 18.5 ★★

A burst arrives at 2 million datagrams a second for 30 milliseconds; the handler serves 1.2 million a second. How many datagrams queue, and what receive buffer (asked) holds them if each is charged 960 bytes?

Solution

Solution of Exercise 18.5.

(2−1.2)×106×0.03=24 000(2 - 1.2) \times 10^6 \times 0.03 = 24\,000 datagrams queue. The kernel must grant 24 000×960≈2324\,000 \times 960 \approx 23 MB, so about 11.5 MB must be asked, which needs rmem_max raised well above the laptop’s 4 MiB.

Exercise 18.6 ★★

After a snapshot recovery the handler’s event stream differs from the clean run’s. In what sense is the output nevertheless “identical once recovered”, and how does the build test it?

Solution

Solution of Exercise 18.6.

The snapshot replaces the lost messages by the state they led to: the book after the snapshot equals the clean run’s book at that sequence number, and every event after it is the clean run’s event. The build checks both: the events after the snapshot’s sequence number equal the clean run’s, and the final book built from each stream is the same.

Exercise 18.7 ★★★

Coding. Make the handler ask for all the holes of a gap in one request (a list of ranges), in the Python reference and the C++ handler, and check on the minute of Figure 18.2 that the worst retransmission staleness falls.

Solution

Solution of Exercise 18.7.

Collect the holes between the next expected number and the largest held one (the gaps in the held map’s keys) and send them as one request with a list of ranges; the server answers all of them in one round trip. The 2.7 ms2.7\,\mathrm{m}\mathrm{s} gap of the minute, which needed several rounds, should then close in one: the timeout plus one round trip, 0.7 ms0.7\,\mathrm{m}\mathrm{s}.

Exercise 18.8 ★★★

Find the flaw. “When our handler detects a gap it drops everything it holds and waits for the next snapshot; that way the book is always consistent.”

Solution

Solution of Exercise 18.8.

Most gaps are filled by the other line within microseconds or by a retransmission within a millisecond; waiting for a snapshot turns each of them into a stale book for up to a snapshot interval, a thousand times longer: in the chapter’s minute, more than four thousand gaps would each have waited for a snapshot. Consistency comes from holding messages until the gap is closed, not from discarding them.

18.9 Problem: The Opening Burst

Problem 18.1

Weekend problem — buffers for the open, staleness when they fail

A handler reads one line of a feed whose opening burst reaches 1.5 million datagrams a second for 20 milliseconds, one message per datagram; the handler serves one datagram in 2 µs2\,\text{µ}\mathrm{s} (decoding, book update, publication). Take the kernel’s charge per 64-byte datagram and its doubling of SO_RCVBUF from Figure 18.3, a maximum rmem_max of 4 MiB, and a snapshot every second.

Part I — The queue.

  1. What is the handler’s service rate, and its utilisation during the burst?
  2. How many datagrams are queued at the end of the burst?
  3. How long after the burst does the queue take to drain if the rate falls to 200 000 a second?
  4. What is the queueing delay of the last datagram of the burst?

Part II — The buffer.

  1. How many bytes of receive buffer must the kernel grant to hold the queue?
  2. What must be asked with SO_RCVBUF, and is it allowed on this machine?
  3. With the largest buffer allowed, how many datagrams are lost?
  4. What else could hold the burst?

Part III — When it fails.

  1. Can the retransmission server fill the loss?
  2. If not, what staleness follows, on average and at worst?
  3. During that staleness, how many datagrams arrive and must be held?
  4. What should the strategy do meanwhile?

Part IV — The verdict.

  1. State the named result: the receive buffer that absorbs the burst without loss, and the staleness a snapshot interval implies when it does lose.
  2. How much faster would the handler have to be to need no queue at all?
  3. Why does reading both lines double the problem?
  4. What does OPRA’s millisecond peak say about a single-threaded handler for a whole options feed?
  5. What would you measure in production to size the buffer?
  6. Which counter tells you that a burst came close to the limit?
  7. Why is a snapshot interval a business decision of the venue with a latency cost for the firm?
  8. In one sentence: what does a feed handler do when it cannot keep up?
Solution

Solution of Problem 18.1.

  1. 500 000 datagrams a second; the burst asks for three times that, a utilisation of 3.
  2. (1.5−0.5)×106×0.02=20 000(1.5 - 0.5) \times 10^6 \times 0.02 = 20\,000.
  3. The queue shrinks at 500 000−200 000=300 000500\,000 - 200\,000 = 300\,000 a second: about 67 ms67\,\mathrm{m}\mathrm{s}.
  4. It waits behind 20 000 others served at 2 µs2\,\text{µ}\mathrm{s}: 40 ms40\,\mathrm{m}\mathrm{s}.
  5. 20 000×960=19.220\,000 \times 960 = 19.2 MB granted.
  6. 9.6 MB, more than rmem_max: the kernel caps it at 4 MiB (8 MiB granted) unless the administrator raises the limit or the process has the privilege to force it.
  7. The 8 MiB buffer holds 8 738 datagrams: about 11 262 are lost.
  8. Memory the handler controls: a reader thread that only moves datagrams from the socket to a large ring (the work is done by another thread), or the network card’s own queues read directly (kernel bypass, chapter 13); or a faster handler.
  9. Yes if the server allows it: 11 262 messages are within its 100 000-message window, in twelve requests of at most 1 000.
  10. Half a snapshot interval on average, 0.5 s0.5\,\mathrm{s}, and a whole one at worst, 1 s1\,\mathrm{s}, plus loading.
  11. At 200 000 a second after the burst, about 100 000 on average and 200 000 at worst.
  12. Stop quoting on the instrument and cancel resting orders that depend on the book, until the flag is lowered.
  13. Named result. Absorbing the burst takes 20 000 datagrams of queue, 19.2 MB of receive buffer granted (9.6 MB asked, more than this machine’s 4 MiB limit); when it fails and only a snapshot can recover, the book is stale for 0.5 s0.5\,\mathrm{s} on average and 1 s1\,\mathrm{s} at worst with a snapshot a second.
  14. Three times faster: serving 1.5 million a second, 0.67 µs0.67\,\text{µ}\mathrm{s} a datagram.
  15. Each line has its own socket and buffer, so the memory doubles, and the handler must read twice the datagrams, half of them duplicates discarded after a comparison.
  16. At 350 million messages a second in its busiest millisecond, a feed of that size cannot be handled by one thread at tens of nanoseconds a message: it must be split across lines and cores.
  17. The distribution of arrivals per 10 milliseconds and per millisecond from captures of the busiest days, the handler’s service time under that load, and the kernel’s drop counters for the sockets.
  18. The receive queue’s high-water mark (and any drop counted by the kernel).
  19. A shorter interval costs the venue bandwidth on a channel most subscribers never read; the firm pays for a longer one in staleness after every unrecoverable gap.
  20. It queues what it can, loses the rest, says so with its stale flag, and recovers from retransmission or the next snapshot.

18.10 Interview questions

Interview question 18.1 ★ developer

What does a market-data feed handler do, from the network card to the strategy?

Solution

Solution of Interview question 18.1.

It receives the redundant lines (from the network card through the kernel or directly), decodes the packets, arbitrates by sequence number, detects and recovers gaps (other line, retransmission, snapshot), normalises each message into the firm’s record, publishes it on shared-memory rings with a staleness flag, and keeps counters for monitoring.

What the interviewer is looking for: the whole chain, with recovery and a staleness signal.

Interview question 18.2 ★★ developer

A gap appears on line A. Walk through what the handler does until the book is correct again.

Solution

Solution of Interview question 18.2.

Hold the later messages and raise the stale flag; wait briefly for line B; if the gap is not filled, ask the retransmission server for the range and apply the answer in order; if the server cannot help, wait for the instrument’s next snapshot, load it, discard held messages it already reflects, apply the rest, and lower the flag.

What the interviewer is looking for: the three recovery paths in order, and the flag.

Interview question 18.3 ★★ developer

How would you size a UDP socket’s receive buffer for a feed? What does the kernel do with the value you set?

Solution

Solution of Interview question 18.3.

From the largest burst expected: the backlog (arrival rate minus service rate, times the burst’s duration) times what the kernel charges per datagram, much more than its payload. The kernel doubles the value set and caps it at rmem_max; check the granted value with getsockopt and watch drop counters.

What the interviewer is looking for: sizing from data, the doubling, the cap and the per-datagram charge.

Interview question 18.4 ★★ developer

How do you know, downstream, that the book you are reading is stale?

Solution

Solution of Interview question 18.4.

The handler publishes its stale-book flag with the events (and the sequence numbers let a consumer see a snapshot reset). A consumer that keeps its own book checks the flag before using it, and checks invariants such as an uncrossed book.

What the interviewer is looking for: an explicit flag propagated with the data.

Interview question 18.5 ★★ developer

Why normalise every venue’s messages into one internal record? What does it cost?

Solution

Solution of Interview question 18.5.

Downstream code (books, strategies, logs) is written once for every venue, records have a fixed size for rings and logs, and sequence numbers travel with them. The cost is a copy per message and fields that some venues do not fill; a venue’s specifics must be carried in the record or they are lost.

What the interviewer is looking for: one downstream interface against a copy and lost detail.

Interview question 18.6 ★★★ developer, researcher

How would you test a feed handler’s recovery paths before production, and prove that it recovers to the right book?

Solution

Solution of Interview question 18.6.

Replay recorded lines with injected losses, reorderings, outages on one and both lines, and bursts, against a retransmission server and a snapshot channel that can be made to fail; compare the output with a clean run (identical events after retransmission; identical book and later events after a snapshot), in every language the handler is written in, as the build does.

What the interviewer is looking for: deterministic replay with injected faults and a clean reference.

Terms defined in this chapter

See all 2333 terms in the glossary