Quantitative Finance · Book 13 · Technology

Low-Latency Software

Low-Latency Software · Technology

16Protocols II: Binary Exchange Protocols

A futures exchange sends every packet of its market data twice, over two separate networks, as binary messages at fixed offsets that a program reads without parsing. CME Group’s documentation tells its clients to “process both the Incremental Feed A and Incremental Feed B due to the unreliable nature of UDP transport”, to keep each sequence number from whichever line brings it first, and to treat a sequence number that neither brought as a loss on both lines, which calls for recovery. The feed handler’s first job is therefore not to understand the market but to keep the first copy of everything and forget the second. This chapter reads the binary protocols of the exchange simulator of One Quant Book 10, which follow the public Nasdaq specifications: order-by-order messages at fixed offsets, packets that carry sequence numbers over multicast, snapshot and retransmission channels for recovery, and order entry over a sequenced session. It generates the codecs from a schema instead of writing them by hand, arbitrates two lossy lines, and measures what decoding costs.

16.1 Order-by-order feeds and their messages

A venue’s direct feed (One Quant Book 1, chapter 9) publishes every event of its book: an order added, executed, reduced, deleted, replaced, a trade against hidden liquidity, an auction’s cross. This is the market-by-order view of One Quant Book 1, chapter 19, from which a subscriber rebuilds the book to any depth. Nasdaq’s TotalView-ITCH is the model the simulator follows: each message starts with a one-byte type, then fields at fixed offsets, with “all integer fields … big endian (network byte order)”. An Add Order is 36 bytes: the type at offset 0, a two-byte stock locate code at 1 (a small integer “employed with the intent of serving as an array index”), a tracking number at 3, a six-byte timestamp in nanoseconds since midnight at 5, the order reference at 11, the side at 19, the shares at 20, the symbol at 24 and the price at 32, an integer in ten-thousandths (Figure 16.1). Nothing needs to be searched: the price of an Add Order is always the four bytes at offset 32.

A MoldUDP64 packet (session, the sequence number of its first message, a message count, then length-prefixed message blocks) and the 36-byte Add Order message it carries, with the byte offset of each field. The layout is the one of Nasdaq’s specifications, which the simulator of One Quant Book 10 follows.
Figure 16.1. A MoldUDP64 packet (session, the sequence number of its first message, a message count, then length-prefixed message blocks) and the 36-byte Add Order message it carries, with the byte offset of each field. The layout is the one of Nasdaq’s specifications, which the simulator of One Quant Book 10 follows.

Definition 16.1 (Flyweight codec)

A flyweight codec reads and writes a message in place: a small object that holds only a pointer to the message’s bytes and decodes a field from its fixed offset when the field is asked for, so that decoding a message copies nothing and costs nothing for the fields not read.

The build generates one flyweight per message from Book 10’s machine-readable schema, a reader and a writer, and a dispatcher that calls a visitor with the right type for a message’s first byte (Listing 16.1). Book 1’s decoder of chapter 28 did the opposite: it copied every message into a common structure and called a std::function for it. Both are correct; Figure 16.3 measures the difference.

// feed A: 36 bytes
struct FeedA {
    static constexpr char kType = 'A';
    static constexpr std::size_t kLength = 36;
    const std::uint8_t* p;
    std::uint16_t locate() const { return static_cast<std::uint16_t>(be16(p + 1)); }
    std::uint16_t tracking() const { return static_cast<std::uint16_t>(be16(p + 3)); }
    std::uint64_t ts() const { return be(p + 5, 6); }
    std::uint64_t ref() const { return be64(p + 11); }
    char side() const { return static_cast<char>(p[19]); }
    std::uint32_t shares() const { return static_cast<std::uint32_t>(be32(p + 20)); }
    std::string_view stock() const { return {reinterpret_cast<const char*>(p + 24), 8}; }
    std::uint32_t price() const { return static_cast<std::uint32_t>(be32(p + 32)); }
Listing 16.1. The generated flyweight of the Add Order message: a pointer, and a byte swap at a fixed offset for each field read. code/firm/wirecodec/cpp/firm_wirecodec.hpp

16.2 Packets, sequence numbers and multicast

Definition 16.2 (Multicast, message gap)

Multicast is IP delivery of one packet to every host that has joined a group address: the sender transmits once, and the network copies the packet towards each subscriber. A message gap is a range of sequence numbers that a receiver has not received when a later one arrives.

Market data is published over UDP multicast because one packet then reaches every subscriber at the same time, whatever their number: MoldUDP64, the packet protocol of the simulator’s feed, is described by its specification as a “lightweight protocol layer built on top of UDP” in which “each outbound packet is transmitted only once regardless of the number of listeners”. UDP delivers without acknowledgement, retransmission or ordering, so the protocol above it carries what the receiver needs to notice a loss: each packet’s header gives the sequence number of its first message and the number of messages it carries (Figure 16.1). A packet whose first number is above the next expected one reveals a gap of the difference. When there is no data, the publisher sends heartbeats, packets with a count of zero carrying the next sequence number, “typically … once per second”, so that a loss is noticed even in a quiet market; a count of 65 535 marks the end of the session. The sequence numbers are those of One Quant Book 1, chapter 28, on a wire.

16.3 Incremental and snapshot channels, retransmission and recovery

Definition 16.3 (Incremental channel, snapshot channel)

An incremental channel carries the feed’s messages, each a change to the book, in sequence. A snapshot channel publishes, at a regular interval, the full state of each instrument’s book together with the incremental sequence number it reflects, so that a receiver that has lost messages, or just started, can rebuild the book without replaying the day.

A gap leaves the receiver two ways back. Retransmission asks a server for the missing sequence numbers: MoldUDP64 has a request packet for this, the simulator answers up to 1 000 messages per request from a window of its last 100 000, and a small gap is closed in a round trip. A gap beyond the window, or a receiver starting in the middle of the day, needs the snapshot channel: the simulator publishes, for each instrument, a begin marker with the incremental sequence number the snapshot reflects, the trading state, every displayed order in priority order, and an end marker with a checksum. The receiver keeps buffering incremental messages, waits for the next snapshot, loads it, discards the buffered messages the snapshot already reflects and applies the rest (the snapshot-and-delta pattern of One Quant Book 3, chapter 26, and the snapshot recovery of One Quant Book 10, chapter 26). CME’s snapshot messages carry the same synchronisation field, LastMsgSeqNumProcessed, “used to synchronize the snapshot loop with the real-time feed”. The recovery time is set by the snapshot interval, and Proposition 16.5 says how often it is needed.

16.4 Arbitration of redundant lines

Definition 16.4 (Line arbitration)

Line arbitration merges the identical packets of redundant feed lines (One Quant Book 10, chapter 26) into one stream: each sequence number is processed from whichever line delivers it first and discarded when the other line delivers it again, and a gap is declared only when a later sequence number arrives before any line has delivered the missing ones.

    std::uint64_t on_packet(const std::uint8_t* p, std::uint64_t& first) {
        const firm::wire::MoldHeader h{p};
        const std::uint64_t seq = h.seq();
        const std::uint64_t count = h.count() == 0xFFFF ? 0 : h.count();
        if (count != 0 && seq + count <= next_) {
            ++c.duplicates;
            return 0;
        }
        if (seq > next_) {
            ++c.gaps;
            c.missing += seq - next_;
        }
        first = seq > next_ ? seq : next_;
        const std::uint64_t fresh = seq + count - first;
        c.messages += fresh;
        ++c.packets;
        if (seq + count > next_) next_ = seq + count;
        return fresh;
    }
Listing 16.2. The arbiter: a duplicate is dropped, a jump forward is a gap lost on both lines, everything else is new. code/low-latency/16-protocols-ii-binary-exchange-protocols/cpp/ll_arb.hpp

Redundancy multiplies probabilities only when the lines fail independently, which is the point of sending them over separate networks, and the reason the simulator draws each line’s losses from its own random stream.

Proposition 16.5 (Losing a packet on both lines)

Let each line lose a given packet with probability pp, independently of the other, and let a common cause (a shared switch, the publisher itself) lose it on both with probability cc, independently of the lines’ own losses. The arbitrated stream loses the packet with probability c+(1−c)p2c + (1 - c)p^2. If each line’s losses follow a two-state Gilbert–Elliott process that moves from the good to the bad state with probability gg per packet, back with probability hh, and loses a packet with probability ℓ\ell in the bad state and none in the good one, the stationary loss rate of a line is p=ℓ g/(g+h)p = \ell\, g/(g+h).

Proof. The packet is lost on the arbitrated stream if the common cause strikes, or if it does not and both lines lose it; the latter events are independent, with probability p2p^2. For the two-state chain, the stationary probability π\pi of the bad state balances the flows between states, πh=(1−π)g\pi h = (1-\pi) g, so π=g/(g+h)\pi = g/(g+h), and a packet is lost with probability ℓπ\ell\pi. ∎

Burstiness does not change a line’s average loss rate, but it does change what a receiver sees: losses arrive in runs, several packets at once, too many for a quick retransmission at the worst moment. And a common cause, however rare, dominates: at p=10−4p = 10^{-4}, independent losses on both lines cost 10−810^{-8} per packet, and a shared failure of 10−610^{-6} is a hundred times more.

Figure 16.2 runs the arithmetic on the simulator. Sixty seconds of a busy instrument, about 32 000 packets, with each line losing 0.1% of packets independently plus Gilbert–Elliott bursts, and a 60-millisecond outage on each line, the two overlapping by 25 milliseconds. Line A shows 56 gaps and 160 missing messages, line B 65 gaps and 186; the arbitrated stream shows two: the 17 messages published while both lines were out, and one packet lost on line B while line A was down.

Gaps on each line and on the arbitrated stream over 60 seconds of the simulator’s feed (about 32 000 packets), with 0.1% independent loss per line, Gilbert–Elliott bursts, and outages of 60 milliseconds on each line overlapping by 25. Each mark is one gap. Data: fig_gaps.py (deterministic: the simulator is seeded).
Figure 16.2. Gaps on each line and on the arbitrated stream over 60 seconds of the simulator’s feed (about 32 000 packets), with 0.1% independent loss per line, Gilbert–Elliott bursts, and outages of 60 milliseconds on each line overlapping by 25. Each mark is one gap. Data: fig_gaps.py (deterministic: the simulator is seeded).

Arbitration costs little: the build’s arbiter handles a packet in about 1.3 ns1.3\,\mathrm{n}\mathrm{s}, a comparison of two sequence numbers. What costs is the design around it: two network interfaces, two paths through two sets of switches, and a feed handler that reads both without letting one line’s delay hold up the other (chapter 18).

As of September 2026 — CME market data

CME Group’s MDP 3.0 market data (client systems documentation, consulted September 2026) is sent on Incremental Feeds A and B over UDP, and CME “strongly recommends that client systems process both”; any packet “can arrive first on either feed”. Its messages use Simple Binary Encoding, little-endian, “optimized for low latency of encoding and decoding”; recovery uses snapshot loops whose messages carry the last incremental sequence number processed.

16.5 Binary order entry and simple binary encoding

Definition 16.6 (Simple binary encoding)

Simple binary encoding (SBE) is the FIX Trading Community’s standard for binary messages: fields of native binary types at fixed positions described by an XML message schema, each message preceded by a small header that gives the length of its fixed block, a template identifier, the schema’s identifier and its version; the byte order is a property of the schema, little-endian by default.

Order entry is binary too at most venues where speed matters. The simulator’s order-entry protocol (One Quant Book 10, chapter 26) runs over SoupBinTCP, which Nasdaq’s specification calls “a lightweight point-to-point protocol, built on top of TCP/IP sockets” that delivers the server’s sequenced messages in order “even across underlying TCP/IP socket connection failures”: its sequenced messages carry no explicit number, both sides count them, and a client that logs in again asks for the next number it wants. It plays the part of the FIX session of chapter 15 at a fraction of the cost: a new order is a 38-byte message with its fields at fixed offsets.

SBE is the FIX Trading Community’s answer to the same needs, standardised in 2016: its specification states a “preference for fixed positions and fixed length fields, supporting direct access to data”. The build derives an SBE-style schema for the simulator’s feed messages from the same JSON schema, little-endian, with the eight-byte message header and natural field sizes (the six-byte timestamp widened to eight), and generates a second set of flyweights from that XML alone.

  <sbe:message name="FeedA" id="65" blockLength="37">
    <field name="locate" id="1" type="uint16" offset="0" />
    <field name="tracking" id="2" type="uint16" offset="2" />
    <field name="ts" id="3" type="uint64" offset="4" />
    <field name="ref" id="4" type="uint64" offset="12" />
    <field name="side" id="5" type="char" offset="20" />
    <field name="shares" id="6" type="uint32" offset="21" />
    <field name="stock" id="7" type="alpha8" offset="25" />
    <field name="price" id="8" type="uint32" offset="33" />
  </sbe:message>
Listing 16.3. The Add Order message in the generated SBE-style schema. code/firm/wirecodec/sbe/feed_sbe.xml

Figure 16.3 compares the decoders on 60 seconds of line A. Book 1’s decoder, which copies each message and calls a std::function, takes about 11 ns11\,\mathrm{n}\mathrm{s} per message; the generated flyweights about 5 ns5\,\mathrm{n}\mathrm{s}, whichever the byte order: a byte swap is one instruction and costs nothing measurable here, while the copy and the indirect call cost half. Framing whole packets and dispatching every message costs about 7 ns7\,\mathrm{n}\mathrm{s} per message in C++ and 8 ns8\,\mathrm{n}\mathrm{s} in Rust. At a million messages a second, decoding is under 1% of a core: the protocol was designed for that.

Decoding the order messages of 60 seconds of the simulator’s line A (about 32 000 messages), median of 21 passes: Book 1’s decoder (a copy and a std::function call per message), the generated flyweights in both byte orders, and, in the last two bars, every packet framed and every message dispatched. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, no isolated cores; GCC 11 at -O2, Rust 1.97 release. Data: bench_wire.py.
Figure 16.3. Decoding the order messages of 60 seconds of the simulator’s line A (about 32 000 messages), median of 21 passes: Book 1’s decoder (a copy and a std::function call per message), the generated flyweights in both byte orders, and, in the last two bars, every packet framed and every message dispatched. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, no isolated cores; GCC 11 at -O2, Rust 1.97 release. Data: bench_wire.py.

16.6 Tutorial: generated codecs and two lossy lines

Goal. Generate the codecs from the simulator’s schema, hold them to Book 10’s golden messages, record two lossy lines, arbitrate them and measure. End state: Figures 16.2 and 16.3, and green tests in Python, C++ and Rust.

  1. Generate with python gen_wirecodec.py: the C++ and Rust flyweights from code/firm/exchsim/schema.json, the SBE-style XML schema, and the little-endian flyweights from that XML. The test runs the generator again and requires no difference.
  2. Check against the golden messages. The C++ test decodes every message of Book 10’s golden fixtures (feed, order entry in both directions, control), compares each with its row in the golden CSV files, re-encodes it with the writer and compares the bytes, and sends every feed message through the SBE-style layout and back; the Rust test decodes the same messages and the recorded MoldUDP64 packets.
  3. Record two lossy lines. ll_lines.record(60, dir) runs the simulator with the losses of Figure 16.2 and writes both lines in the recorded-file format.
  4. Arbitrate with Listing 16.2; bench_wire.py checks the C++ counts against the Python reference, then measures the decoders.

What to change next. Remove the outage overlap and check that the arbitrated stream has no gap; put the two lines behind one simulated switch (the same outage on both) and watch every gap pass through.

16.7 Build: the wire codecs

Purpose. Every byte the firm reads from or writes to the simulator: the feed handler (chapter 18) decodes the feed with it, the order gateway (chapter 21) encodes and decodes order entry, and the capture tools (chapter 23) read recorded files.

Interface. Generated by gen_wirecodec.py from code/firm/exchsim/schema.json: C++20 firm::wire with a reader Feed<T>, In<T>, Out<T>, Ctl<T> and a writer Writer per message (FeedAWriter and so on), dispatch_feed, dispatch_in, dispatch_out, dispatch_ctl, MoldHeader, for_each_block, csv and rewrite; the SBE-style firm::sbe readers and writers and to_sbe; Rust firm_wirecodec with the same readers, csv_feed, csv_in, csv_out, csv_ctl, MoldHeader and for_each_block.

Rules. No hand-written layout: every offset comes from the schema; generated files are committed and regenerated without difference; readers copy nothing and allocate nothing; a message whose length does not match its type is refused by the dispatcher.

Acceptance tests. code/firm/wirecodec/: every message of Book 10’s golden fixtures decoded to its golden values and re-encoded to the same bytes in C++, decoded in Rust; every feed message round-tripped through the SBE-style layout; the recorded MoldUDP64 fixture framed; the generator deterministic and its outputs current.

Stretch. Generate the SoupBinTCP frames and a MoldUDP64 retransmission request; generate from the XML schema in Rust as well; a codec for a real venue’s published schema.

Sources and further reading

  • Nasdaq, TotalView-ITCH 5.0, MoldUDP64 and SoupBinTCP specifications.
  • CME Group Client Systems Wiki, MDP 3.0: “Incremental Feed Arbitration”, “Simple Binary Encoding”, “Market Data Snapshot – Full Recovery”.
  • FIX Trading Community, Simple Binary Encoding, version 1.0 (GitHub).
  • E. N. Gilbert, “Capacity of a burst-noise channel”, Bell System Technical Journal 39(5), 1960; E. O. Elliott, “Estimates of error rates for codes on burst-noise channels”, Bell System Technical Journal 42(5), 1963.
  • The protocol description of the simulator, code/firm/exchsim/PROTOCOL.md (One Quant Book 10, chapter 26).

16.8 Exercises

Exercise 16.1 ★

A MoldUDP64 packet has sequence number 1 000 and count 3; the next packet on the same line has sequence number 1 007. Which messages are missing, and what does a heartbeat with sequence number 1 010 then tell you?

Solution

Solution of Exercise 16.1.

The first packet carries 1 000–1 002, so the next expected number is 1 003: messages 1 003–1 006 are missing (a gap of four). The heartbeat’s number is the next expected one, so the packet at 1 007 carried three messages (1 007–1 009) and nothing more was lost after it.

Exercise 16.2 ★

Decode by hand the price of an Add Order whose bytes 32–35 are 00 0F 42 A4. What is it in currency units?

Solution

Solution of Exercise 16.2.

0x000F42A4=1 000 100\mathtt{0x000F42A4} = 1\,000\,100 ten-thousandths: 100.01.

Exercise 16.3 ★

Each line loses 0.2% of packets independently. What fraction does the arbitrated stream lose? And if both lines pass through one switch that drops 0.01% of packets?

Solution

Solution of Exercise 16.3.

0.0022=4×10−60.002^2 = 4 \times 10^{-6} with independent lines. With the shared switch, 10−4+(1−10−4)×4×10−6≈1.04×10−410^{-4} + (1 - 10^{-4}) \times 4 \times 10^{-6} \approx 1.04 \times 10^{-4}: the switch alone accounts for 96% of the losses.

Exercise 16.4 ★★

Compute the stationary loss rate of the simulator’s Gilbert–Elliott parameters (good to bad 0.0005, bad to good 0.1, 80% lost in the bad state) and the mean length of a bad period in packets. Why does burstiness matter for retransmission?

Solution

Solution of Exercise 16.4.

The bad state’s stationary probability is 0.0005/0.1005≈0.50%0.0005/0.1005 \approx 0.50\%, so the line loses 0.8×0.50%≈0.40%0.8 \times 0.50\% \approx 0.40\% of packets; a bad period lasts 1/0.1=101/0.1 = 10 packets on average. Losses therefore come in runs of several packets at once: a retransmission request covers them, but at the moment they happen, which is often a burst of traffic, and one request may not be enough.

Exercise 16.5 ★★

Why is market data sent over UDP multicast rather than TCP, and what does a subscriber give up?

Solution

Solution of Exercise 16.5.

One multicast packet reaches every subscriber at once, with no per-subscriber connection, acknowledgement or retransmission to slow the publisher, and no subscriber can hold the others back. A subscriber gives up reliable, ordered delivery: it must detect gaps from sequence numbers, arbitrate redundant lines and recover from retransmission or snapshots itself.

Exercise 16.6 ★★

The SBE-style layout widens the six-byte timestamp to eight bytes. What does that cost in bytes for an Add Order, and what does it save?

Solution

Solution of Exercise 16.6.

The message becomes 37 bytes of fields plus an 8-byte header, 45 against 36: 25% more bandwidth and cache for Add Orders. What it saves is the assembly of a six-byte integer; the measurement shows that this saving is not measurable, while the header buys versioning and a length that lets old readers skip fields they do not know.

Exercise 16.7 ★★★

Coding. Extend gen_wirecodec.py to generate the SoupBinTCP frames of the schema’s soup_client and soup_server protocols, and test them on Book 10’s fixture_soup.bin.

Solution

Solution of Exercise 16.7.

Add soup_client and soup_server to the protocols the generator walks, generating for the variable-length messages (U, S, +) a reader that returns the payload as a view; the frame is a big-endian u16 length that counts the type byte. The test reads fixture_soup.bin frame by frame and compares each with fixture_soup.csv.

Exercise 16.8 ★★★

Find the flaw. “Our feed handler reads line A; if it sees a gap it switches to line B until the next gap.”

Solution

Solution of Exercise 16.8.

It does not arbitrate: while it reads one line, the other’s packets are ignored, so each switch loses whatever arrived first on the other line, and a gap on line A is only filled if line B’s copies are still to come. The handler must read both lines all the time and keep, per sequence number, the first copy from either.

16.9 Problem: Two Copies of Everything

Problem 16.1

Weekend problem — how often redundancy fails, and what recovery costs

A feed publishes 50 000 packets a second over a 6.5-hour session on two lines. Assume each line loses a packet with probability 10−410^{-4}, independently; retransmission serves gaps of up to 1 000 messages in a round trip of 0.5 ms0.5\,\mathrm{m}\mathrm{s}; a snapshot of the instrument arrives every second and takes 20 ms20\,\mathrm{m}\mathrm{s} to receive; applying a buffered message costs 8 ns8\,\mathrm{n}\mathrm{s} (a little above the measured cost of framing and dispatching one).

Part I — Independent loss.

  1. How many packets does a session carry, and how many does one line lose?
  2. What is the probability that the arbitrated stream loses a packet, and how many losses is that per session?
  3. In the simulation of Figure 16.2, why are there any arbitrated gaps at all?
  4. What does each arbitrated loss cost with retransmission?

Part II — Correlation.

  1. A shared switch drops 10−610^{-6} of packets on both lines. What is the arbitrated loss rate now (Proposition 16.5), and how many losses per session?
  2. What stationary loss rate do the simulator’s burst parameters give each line, and what arbitrated rate if the lines are independent?
  3. Why do bursts make retransmission less useful even when the average loss rate is unchanged?
  4. Name three common causes a firm can remove, and one it cannot.

Part III — Recovery.

  1. When is retransmission not enough?
  2. How long does recovery from the snapshot channel take on average, and at worst?
  3. How many incremental messages arrive meanwhile, and how long does applying them take?
  4. Which buffered messages must be discarded after the snapshot, and how does the handler know?

Part IV — The verdict.

  1. State the named result: the arbitrated loss probability with independent and with correlated losses, and the recovery time from the snapshot channel.
  2. What should a strategy do with its view of the book during a recovery?
  3. Would a snapshot every 100 milliseconds help? What does it cost the venue and the subscriber?
  4. Why must the arbiter never wait for the slower line?
  5. How would you test the recovery path before production?
  6. What should be logged for every gap?
  7. What does the measured cost of decoding say about where the latency of a feed handler lies?
  8. In one sentence: what does a second line buy?
Solution

Solution of Problem 16.1.

  1. 50 000×23 400=1.17×10950\,000 \times 23\,400 = 1.17 \times 10^9 packets; one line loses about 117 000 of them.
  2. 10−810^{-8}, about 11.7 losses a session.
  3. Because the losses were not independent: both lines were out together for 25 milliseconds (17 messages), and line B lost a packet while line A was down.
  4. One retransmission round trip, 0.5 ms0.5\,\mathrm{m}\mathrm{s}, during which the book cannot be trusted.
  5. 10−6+(1−10−6)×10−8≈1.01×10−610^{-6} + (1 - 10^{-6}) \times 10^{-8} \approx 1.01 \times 10^{-6}: about 1 182 losses a session, a hundred times more.
  6. About 0.40% per line, and 1.6×10−51.6 \times 10^{-5} arbitrated if the lines are independent.
  7. Losses arrive in runs at the busiest moments; the request, the answer and the replay happen while the market moves, and a run longer than the retransmission limit needs a snapshot.
  8. A shared switch, a shared network card or server, a shared cross-connect or carrier, shared power; it cannot remove the publisher’s own failures, which are common to both lines.
  9. When the gap exceeds what the server keeps or serves (1 000 messages per request, a window of 100 000), when the server is unreachable, or when the handler starts late.
  10. Half an interval plus the snapshot on average, 0.5+0.02=0.52 s0.5 + 0.02 = 0.52\,\mathrm{s}; a full interval at worst, 1.02 s1.02\,\mathrm{s}.
  11. About 26 000 messages on average (51 000 at worst), applied in about 0.21 ms0.21\,\mathrm{m}\mathrm{s} (0.41 ms0.41\,\mathrm{m}\mathrm{s}).
  12. Those the snapshot already reflects: sequence numbers up to the one the snapshot’s begin marker gives (CME’s LastMsgSeqNumProcessed).
  13. Named result. With independent losses of 10−410^{-4} per line the arbitrated stream loses 10−810^{-8} per packet, about 12 packets a session at 50 000 a second; a common cause of 10−610^{-6} makes it 1.01×10−61.01 \times 10^{-6}, about 1 180 a session. Recovery from a one-second snapshot cycle takes 0.52 s0.52\,\mathrm{s} on average and 1.02 s1.02\,\mathrm{s} at worst, against 0.5 ms0.5\,\mathrm{m}\mathrm{s} for a retransmission.
  14. Mark the book stale, stop quoting on it (or quote only defensively) and cancel what depends on it until it is rebuilt.
  15. It cuts the average wait to 50 ms50\,\mathrm{m}\mathrm{s}, at the price of ten times the snapshot bandwidth for every subscriber, most of whom never need it.
  16. Waiting for the slower line adds that line’s delay to every message; the point of two lines is to take the faster copy.
  17. Replay recorded lines with losses injected, one line at a time and both together (as the simulator does), and check that the rebuilt book equals the book of a lossless run.
  18. The line, the sequence range, the time, whether it was recovered and how, and how long the book was stale.
  19. Almost nowhere in decoding: a few nanoseconds per message; the latency is in the network, the kernel and the book.
  20. Protection from independent failures of one path, not from failures common to both.

16.10 Interview questions

Interview question 16.1 ★ developer

Why do exchanges publish market data over UDP multicast, and how does a receiver detect a lost packet?

Solution

Solution of Interview question 16.1.

One packet reaches every subscriber at once, and the publisher never waits for anyone. Each packet carries the sequence number of its first message and a count; a packet that starts above the next expected number reveals a gap, and heartbeats carry the next number so that losses are seen even when nothing trades.

What the interviewer is looking for: fairness and scale; sequence numbers and heartbeats.

Interview question 16.2 ★★ developer

Implement A/B line arbitration. What state do you keep, and what do you do on a gap?

Solution

Solution of Interview question 16.2.

Keep the next expected sequence number. For each packet from either line: if it ends before that number, drop it; if it starts after it, record a gap (lost on both) and start recovery; otherwise deliver the new messages and advance. Read both lines without waiting for either; buffer what arrives during recovery.

What the interviewer is looking for: one counter, duplicates dropped, gaps recovered.

Interview question 16.3 ★★ developer

What is a flyweight decoder, and why is it faster than decoding into structs?

Solution

Solution of Interview question 16.3.

An object holding only a pointer to the message, whose accessors read each field from its fixed offset when asked. Nothing is copied, nothing allocated, unread fields cost nothing, and there is no indirect call: the measured cost is half that of a decoder that copies into a struct and calls a callback.

What the interviewer is looking for: views over bytes and lazy field access.

Interview question 16.4 ★★ developer

You joined the feed in the middle of the day. How do you build a correct book?

Solution

Solution of Interview question 16.4.

Subscribe to the incremental lines and buffer; wait for a snapshot of each instrument; load it; drop buffered messages up to the sequence number the snapshot reflects; apply the rest in order; then process live. Verify with the snapshot’s checksum where there is one.

What the interviewer is looking for: buffer, snapshot, sequence number alignment.

Interview question 16.5 ★★ developer

Big-endian or little-endian on the wire: does it matter for latency? What does?

Solution

Solution of Interview question 16.5.

Hardly: a byte swap is one instruction, and the chapter’s measurement shows no difference between big- and little-endian flyweights. What matters is not copying, not allocating, not calling through pointers, and fixed offsets that need no search.

What the interviewer is looking for: measure; the costs are elsewhere.

Interview question 16.6 ★★★ developer, researcher

Your two lines are “redundant” but you see simultaneous gaps on both several times a day. What do you investigate?

Solution

Solution of Interview question 16.6.

A common cause: shared switches, cards, cables, carriers or power; the publisher itself (both lines losing the same range at the source); the receiving host (both lines’ interrupts on one overloaded core, a full socket buffer, a slow handler). Compare gap times and ranges across lines and with other subscribers, and check the host’s drop counters.

What the interviewer is looking for: independence as the assumption to test, including the receiving host.

Terms defined in this chapter

See all 2333 terms in the glossary