---
title: "Capturing and Monitoring the Network"
book: "Networks, Hardware and Trading Infrastructure"
subject: quant
language: en
chapter: 5
exercises: 8
source: https://one-course.com/books/quant/14/en/chapter/5-capturing-and-monitoring-the-network
---

# Chapter 5 — Capturing and Monitoring the Network

In this chapter’s simulated cage the strategy’s own log adds up, at the median, to $1.5\,\text{µ}\mathrm{s}$ between decoding a packet and handing the order it caused to the network card. The capture taken at the cage’s taps says that the order’s frame left the order handoff $4.4\,\text{µ}\mathrm{s}$ after the triggering packet’s first byte reached the market-data handoff. The difference is the part of the path the program cannot see: the fibre and the layer-1 devices, the card’s receive path into memory, and its transmit path out. At the 99th percentile the capture says $38\,\text{µ}\mathrm{s}$, and no log written by the program could have said which part of the path owns it.

This chapter is about the firm watching its own traffic from the outside. It copies every frame at the handoffs (chapter 2’s taps), timestamps it against the clocks of chapter 4, stores it, and answers from the capture alone the questions the firm asks every day: how long did we take, what did we lose, when did the port fill.

## 5.1 Taps, capture appliances and capture files

**Definition 5.1 (Packet capture, capture appliance).**

A *packet capture* is a record of the frames that crossed a point of a network, each with the time it crossed. A *capture appliance* is a dedicated device that receives copies of traffic from taps, timestamps each frame against a disciplined clock, and writes the frames to storage at the [line rate](https://one-course.com/books/quant/14/en/chapter/1-networking-for-trading#def-nw-networking-for-trading-ser) of its inputs without losing any.

**Definition 5.2 (Packet capture file, timestamp trailer).**

A *packet capture file* stores a capture as a sequence of records, each a timestamp, a length and the frame’s bytes; the libpcap format and its successor pcapng are the common ones, with nanosecond timestamps in their newer variants. A *timestamp trailer* is a block of bytes that a tapping or switching device appends to the end of a copied frame, carrying the time at which the device saw the frame and where it saw it (device and port), so that the timestamp travels with the frame to wherever it is stored.

The timestamp that matters is the one taken where the frame crossed the tap, not the one the storage server writes when the copy reaches it: between the two lie an aggregation switch, a queue and the server’s own card. A trailer fixes that. It is taken on the wire, in the frame’s first bytes, and it rides the copy through any later aggregation; the storage server keeps it and can ignore its own time. [Box 5.1](#dat-nw-capturing-and-monitoring-the-network-trailers) describes one vendor’s published layout; the book’s `firm.wirecap` uses a 16-byte trailer of its own in the same spirit.

![Watching a cage from outside. Taps on every handoff copy both directions; a timestamping device appends a trailer to each copy (time, device, port) against the grandmaster’s time and aggregates the copies; the capture appliance stores and analyses them. No program in the cage is involved.](https://one-course.com/images/onecourse/chapters/quant-14/nw-capturing-and-monitoring-the-network/fig-6c3e7f10de38.svg)

***Figure 5.1.** Watching a cage from outside. Taps on every handoff copy both directions; a timestamping device appends a trailer to each copy (time, device, port) against the grandmaster’s time and aggregates the copies; the [capture appliance](#def-nw-capturing-and-monitoring-the-network-capture) stores and analyses them. No program in the cage is involved.*

![The timestamp trailer of firm.wirecap: the device and port that saw the frame and the full UTC time in seconds and nanoseconds, appended after the frame. Vendors’ trailers carry the same information in their own layouts.](https://one-course.com/images/onecourse/chapters/quant-14/nw-capturing-and-monitoring-the-network/fig-56c39bbd9573.svg)

***Figure 5.2.** The [timestamp trailer](#def-nw-capturing-and-monitoring-the-network-trailer) of `firm.wirecap`: the device and port that saw the frame and the full UTC time in seconds and nanoseconds, appended after the frame. Vendors’ trailers carry the same information in their own layouts.*

**As of September 2026 — A published trailer.**

Arista documents that its 7130 devices insert a self-contained 16-byte timestamp header at the end of each copied [Ethernet frame](https://one-course.com/books/quant/14/en/chapter/1-networking-for-trading#def-nw-networking-for-trading-frame), recompute its check sequence, and carry 64 bits of seconds and nanoseconds, so that “full UTC time” needs no keyframes; the header names the device and port, and a flag says whether the original frame’s check sequence was valid. Timestamps are taken “on the first byte of a packet on the ingress port”, with the front-panel delay removed, and the document gives $\pm$75 picoseconds of accuracy when the device is synchronised by a pulse per second. Other devices overwrite a field of the frame with a shorter timestamp and send periodic keyframes to relate it to UTC.

```python
def write_pcap(path, records):
    with open(path, "wb") as f:
        f.write(struct.pack("<IHHiIII", MAGIC_NS, 2, 4, 0, 0, 65535, LINKTYPE_ETHERNET))
        for t, fr in records:
            f.write(struct.pack("<IIII", t // 1_000_000_000, t % 1_000_000_000, len(fr), len(fr)) + fr)


def read_pcap(path):
    b = open(path, "rb").read()
    magic, _, _, _, _, _, link = struct.unpack_from("<IHHiIII", b, 0)
    if magic != MAGIC_NS or link != LINKTYPE_ETHERNET:
        raise ValueError("not a nanosecond Ethernet pcap written by this library")
    out, off = [], 24
    while off + 16 <= len(b):
        sec, ns, incl, _ = struct.unpack_from("<IIII", b, off)
        out.append((sec * 1_000_000_000 + ns, b[off + 16:off + 16 + incl]))
        off += 16 + incl
    return out
```

***Listing 5.1.** Writing and reading a nanosecond pcap file: a 24-byte file header, then per frame a 16-byte record header and the bytes. code/firm/wirecap/firm_wirecap.py*

## 5.2 Measuring wire-to-wire latency

Wire-to-wire latency (One Quant Book 13, chapter 1) is measured on the wire: from the first byte of a market-data frame at the firm’s handoff to the first byte of the order it caused at the order handoff. From a capture it is a subtraction, once each order has been paired with its cause.

**Definition 5.3 (Trigger packet).**

The *trigger packet* of an order is the market-data packet whose content caused the firm to send it. Its capture time is the start of the order’s wire-to-wire latency.

The pairing is the hard part. A capture sees two streams of packets, and nothing on the wire says which packet caused which order. The naive rule, “the latest packet with a triggering message before the order”, works when triggers are rare and the firm is fast; it fails exactly when latency is worth measuring, in bursts, when several triggers precede an order by less than the firm’s own latency. The reliable rule is to carry the cause in the order: the firm puts the [trigger packet](#def-nw-capturing-and-monitoring-the-network-trigger)’s sequence number (or a hash of it) into a field of the order it sends, usually its client order identifier, and the capture reads it back.

**Method 5.4 (Wire-to-wire from a capture).**

1. Tap both market-data lines and the order handoff; timestamp all three against the same clock, on the wire.
2. For each market-data packet keep the time of the first copy seen on either line: the strategy acted on whichever arrived first.
3. Have the strategy carry the trigger’s sequence number in every order; for each order, subtract its trigger’s time from its own.
4. Report the distribution, not the mean, and break it down against a budget of the path’s stages, so that a slow percentile can be charged to a stage.

The chapter’s tutorial runs this method on a simulated second: 99 055 market-data packets on two lines, each copy lost independently with probability 0.2%, and a strategy that answers 30% of the packets carrying a trade message with an order, 9 558 orders in all. The orders’ timing is drawn from `firm.wirepath`, the extension of One Quant Book 13’s tick-to-trade path to the wire ([Figure 5.3](#fig-nw-capturing-and-monitoring-the-network-budget)). The capture recovers every order’s latency exactly by the sequence carried in the order; the naive rule pairs 47% of them with the wrong packet.

![The wire-to-wire budget of firm.wirepath: the layer-1 cage of chapter 2 (network in and out), the card’s receive and transmit paths (model values), and the six in-process stages measured by One Quant Book 13 on a laptop (Intel Core Ultra 7 155H, WSL2). The last bar is their composition, a simulation. Data: fig_capture.py.](https://one-course.com/images/onecourse/chapters/quant-14/nw-capturing-and-monitoring-the-network/fig-483a58316e20.svg)

***Figure 5.3.** The wire-to-wire budget of `firm.wirepath`: the layer-1 cage of chapter 2 (network in and out), the card’s receive and transmit paths (model values), and the six in-process stages measured by One Quant Book 13 on a laptop (Intel Core Ultra 7 155H, WSL2). The last bar is their composition, a simulation. Data: `fig_capture.py`.*

The budget says where the time goes. At the median the card’s two passes ($0.8\,\text{µ}\mathrm{s}$ each, model values) and the ring between the feed handler and the strategy ($0.76\,\text{µ}\mathrm{s}$, measured) are the largest stages; the fibre and devices of the cage together are $0.56\,\text{µ}\mathrm{s}$. At the 99th percentile the ring owns it, at $32\,\text{µ}\mathrm{s}$: on the laptop that measured it, a consumer that sleeps or is descheduled wakes late. Composed, the stages give a wire-to-wire median of $4.4\,\text{µ}\mathrm{s}$ and a 99th percentile of $35\,\text{µ}\mathrm{s}$; the capture of the simulated session measures $4.4\,\text{µ}\mathrm{s}$ and $38\,\text{µ}\mathrm{s}$, the difference being the sample.

![Simulation: wire-to-wire latency of 9 558 orders measured from the capture of one simulated second, by the trigger sequence each order carries (bins of equal width in log scale). The distribution has the long right tail of its slowest stage. Data: fig_capture.py.](https://one-course.com/images/onecourse/chapters/quant-14/nw-capturing-and-monitoring-the-network/fig-b5eabbc97dc0.svg)

***Figure 5.4.** Simulation: wire-to-wire latency of 9 558 orders measured from the capture of one simulated second, by the trigger sequence each order carries (bins of equal width in log scale). The distribution has the long right tail of its slowest stage. Data: `fig_capture.py`.*

```python
    return rec, truth


def analyse(records):
    feed, seqs, trig_t, orders, stamps, sizes = {}, {b"A": [], b"B": []}, [], [], [], []
    for _, fr in records:                                   # time from the trailer, not the pcap header
        frame, t, _, port = wc.strip_trailer(fr)
        p = wc.parse(frame)
        if p["proto"] == "udp":
            _, s, k, msgs = wc.mold(p["payload"])
            line = b"A" if port == 1 else b"B"
            seqs[line].extend(range(s, s + k))
            if line == b"A":
                stamps.append(t)
                sizes.append(len(frame) + 24)               # frame check sequence, preamble and gap
            if s not in feed:                               # first copy seen, either line
                feed[s] = t
                if any(m[:1] == b"E" for m in msgs):
                    trig_t.append(t)
        else:
            _, _, trig = wc.ORDER.unpack(p["payload"])
            orders.append((t, trig))
```

***Listing 5.2.** Reading the capture: each frame’s time from its trailer, the first copy of each market-data packet on either line, and the trigger sequence of each order. code/networks/05-capturing-and-monitoring-the-network/python/nw_capture.py*

## 5.3 Detecting gaps, drops and microbursts from a capture

The same capture answers the questions the feed handler of One Quant Book 13 answers from inside, and some it cannot. A gap is a sequence number missing from a line; a gap on both lines is a message the firm never had, and the capture shows whether it was lost before the handoff (missing at the tap) or after it (present at the tap, missing in the handler’s counters). In the simulated second, line A shows 158 gaps and line B 198, and no message is missing from both: the handler recovered everything from the other line. Bursts come from the same timestamps: the busiest 100 microseconds of line A carried $1.24\,\mathrm{G}\mathrm{bit}/\mathrm{s}$, nine times its mean of $0.14\,\mathrm{G}\mathrm{bit}/\mathrm{s}$, the same shape as chapter 1’s feeds at a smaller scale.

A capture is the firm’s evidence as well as its instrument. It shows what the firm actually sent and when, which is what a venue, a client or a regulator asks about after an incident; European rules ask firms to identify the exact point at which each timestamp is applied, and a tap on the handoff is the one point that no software change can move.

## 5.4 Monitoring in production: counters, telemetry and alarms

**Definition 5.5 (Streaming telemetry).**

*Streaming telemetry* is the continuous export by network devices of their counters and state (per-port traffic, discards, queue occupancy, optical levels, clock state) to a collector, pushed at a chosen interval or on change, instead of polled by the collector.

A capture is detailed and expensive: it stores every frame, and a busy cage fills disks in hours. Monitoring runs continuously on summaries: every device’s counters, streamed every second or on change; every server’s socket and ring drops; every handler’s gap and staleness counters; every clock’s offset. The alarms that matter in a trading network are few and specific:

- **discards on an [egress queue](https://one-course.com/books/quant/14/en/chapter/1-networking-for-trading#def-nw-networking-for-trading-egress)** , which only [microbursts](https://one-course.com/books/quant/14/en/chapter/1-networking-for-trading#def-nw-networking-for-trading-burst) produce on an underused link (chapter 1);
- **gaps on one line** persisting, a line degrading while the other hides it;
- **gaps on both lines** , a common cause, and any snapshot recovery;
- **a clock out of lock or in [holdover](https://one-course.com/books/quant/14/en/chapter/4-time-synchronisation#def-nw-time-synchronisation-holdover)** , and offsets approaching the budget (chapter 4);
- **wire-to-wire percentiles** drifting from the baseline, computed from the capture every minute.

Each alarm names a place on the path, and the capture of the minutes around it is kept for the investigation.

## 5.5 Tutorial: a second of the cage, from the taps

**Goal.** Capture a simulated second at the cage’s taps, write it as a pcap file, and measure from it alone what the firm’s programs cannot see. **End state:** Figures [5.3](#fig-nw-capturing-and-monitoring-the-network-budget) and [5.4](#fig-nw-capturing-and-monitoring-the-network-hist), the pairing error of the naive rule, and the gap and burst counts quoted in the text.

1. **The budget.** `nw_capture.stages()` takes the cage of chapter 2, the card’s model stages and Book 13’s measured stages; `budget_rows()` composes them with `firm.wirepath` .
2. **The session.** `session()` builds the feed’s frames on both lines with independent losses, the orders with their trigger sequence, and appends a trailer to each copy.
3. **The file.** `firm.wirecap.write_pcap(path, records)` writes it ( [Listing 5.1](#lst-nw-capturing-and-monitoring-the-network-pcap) ); any tool that reads nanosecond pcap files opens it.
4. **The analysis.** `analyse(read_pcap(path))` pairs orders with their triggers, counts gaps per line and messages lost on both, and finds the busiest window ( [Listing 5.2](#lst-nw-capturing-and-monitoring-the-network-analyse) ); `python fig_capture.py` writes the charts’ data.

**What to change next.** Raise the loss to 5% per line and count the messages lost on both; make the strategy four times faster and measure how the naive rule’s error changes.

## 5.6 Build: capture tools and the wire-to-wire path

**Purpose.** Two components: `firm.wirecap`, the firm’s capture toolkit, and `firm.wirepath`, the tick-to-trade path of One Quant Book 13 extended to the wire, which chapters 7 and 29 use for their latency tables.

**Interface.** `firm_wirecap`: `udp_frame`, `tcp_frame`, `add_trailer`, `strip_trailer`, `write_pcap`, `read_pcap`, `write_pcapng`, `read_pcapng`, `parse`, `mold`, `ORDER`, `match_orders`, `match_nearest`, `gaps`, `bursts`; C++20 reader `cpp/firm_wirecap.hpp`. `firm_wirepath`: `Stage(name, p50, p99, kind)`, `book13_software()`, `wire_stages(cage, …)`, `card_stages`, `hardware_stage(cycles, mhz)`, `compose`, `report`, composing with `firm.latbudget`.

**Rules.** Timestamps are integers of nanoseconds since midnight UTC; the trailer’s time wins over any storage time; the IPv4 checksum is correct; an order is paired only by the sequence it carries; Book 13’s measurements are read, never copied or changed.

**Acceptance tests.** `code/firm/wirecap/tests/`: frames parse back, checksums hold, trailers round-trip, pcap and pcapng round-trip, a wrong magic is refused, matching, gaps and bursts on hand-made data, and the shared fixture read identically in Python and C++. `code/firm/wirepath/tests/`: the lognormal fit hits its median and 99th percentile, Book 13’s stages are read, a cage’s path gives the network stage, and the report’s rows.

**Stretch.** Per-port burst detection with the switch’s buffer as the threshold; a capture-based order-book rebuild compared, message by message, with the handler’s.

Sources and further reading

- `pcap-savefile(5)` (tcpdump/libpcap); IETF, *PCAP Now Generic (pcapng) Capture File Format* , draft-ietf-opsawg-pcapng-04.
- Arista, *An Overview of Arista Ethernet Capture Timestamps* ; OpenConfig, *gNMI specification* .
- Commission Delegated Regulation (EU) 2017/574 (RTS 25), article 4.

## 5.7 Exercises

**Exercise 5.1 ★.**

A nanosecond pcap record header gives 34 200 seconds and 1 500 000 321 nanoseconds. Is that a valid record, and what time of day should the reader report?

**Solution of Exercise 5.1.**

No: the nanoseconds field must be below $10^9$. A lenient reader would carry the second and report 34 201 s and 500 000 321 ns, 09:30:01.500 000 321; a strict one should reject the record, because the writer is broken.

**Exercise 5.2 ★.**

A [trigger packet](#def-nw-capturing-and-monitoring-the-network-trigger) reaches the tap on line A at 09:30:00.000 012 345 and on line B $3\,\text{µ}\mathrm{s}$ later; the order leaves at 09:30:00.000 016 771. What is the wire-to-wire latency? If the copy on line A had been lost, from which copy would you measure, and what would change?

**Solution of Exercise 5.2.**

$16\,771 - 12\,345 = 4426\,\mathrm{n}\mathrm{s}$, from the first copy (line A). With line A’s copy lost the strategy saw the packet on line B $3\,\text{µ}\mathrm{s}$ later and its order would have left about $3\,\text{µ}\mathrm{s}$ later: measured from the B copy the wire-to-wire latency is the same, and the $3\,\text{µ}\mathrm{s}$ belong to the lost line.

**Exercise 5.3 ★.**

A 10 Gb/s link is tapped in both directions and written to disk with trailers; each frame is stored with a 16-byte record header and its 16-byte trailer. At a sustained $4\,\mathrm{G}\mathrm{bit}/\mathrm{s}$ each way with 200-byte frames, how many bytes a second go to disk?

**Solution of Exercise 5.3.**

Each direction carries $4 \times 10^9 / (220 \times 8) \approx 2.27$ million frames a second (200 bytes plus 20 of preamble and gap on the wire); each is stored as $200 + 16 + 16 = 232$ bytes: about 527 MB a second per direction, 1.05 GB a second for the link.

**Exercise 5.4 ★★.**

Why does the naive pairing rule fail more often the more frequent the triggers are and the slower the firm is?

**Solution of Exercise 5.4.**

The rule assigns an order to the latest trigger before it. It is wrong whenever another trigger arrives between the true trigger and the order, which happens with a probability that grows with the trigger rate times the firm’s latency: frequent triggers, or a slow firm, put more triggers in the window.

**Exercise 5.5 ★★.**

Line A shows 158 gaps and line B 198 in the tutorial’s second, and none is on both. With 0.2% independent loss per copy and 99 055 packets, how many packets would you expect to be lost on both lines?

**Solution of Exercise 5.5.**

$99\,055 \times 0.002^2 \approx 0.4$: seeing none is what independence predicts. A packet lost on both lines is a common cause.

**Exercise 5.6 ★★.**

In [Figure 5.3](#fig-nw-capturing-and-monitoring-the-network-budget), which stage would you attack first to cut the median, and which to cut the 99th percentile? What would each cost?

**Solution of Exercise 5.6.**

The median: the card’s two passes (1.6 $\text{µ}\mathrm{s}$ together, a better card or kernel bypass on a tuned host) and the ring (a busy-polling consumer on an isolated core). The 99th percentile: the ring, whose tail comes from scheduling; isolating its cores costs cores and tuning (One Quant Book 13, chapter 13), not money.

**Exercise 5.7 ★★★.**

*Coding.* Make every stage of `firm.wirepath` four times faster except the network’s, run the session again and measure the naive rule’s pairing error. Explain the direction of the change.

**Solution of Exercise 5.7.**

The naive rule’s error falls from 47% to 25%: with a four times faster firm, fewer triggers arrive between a trigger and its order. It does not vanish, because [trigger packets](#def-nw-capturing-and-monitoring-the-network-trigger) still come in bursts closer together than the firm’s latency.

**Exercise 5.8 ★★★.**

*Find the flaw.* “We measure our wire-to-wire latency by timestamping the market-data packet in the handler with the card’s hardware timestamp and the order with the card’s transmit timestamp. That is as good as a tap.”

**Solution of Exercise 5.8.**

The card’s timestamps miss everything outside the card: the cage’s fibre and devices on both sides, which the tap sees. They also rely on the card’s clock and on how the card defines its receive and transmit instants. It is a good in-host measurement; it is not the firm’s wire-to-wire latency at the handoffs.

## 5.8 Problem: The Two Microseconds Nobody Logged

**Problem 5.1.**

Weekend problem — a latency budget read from the taps

A firm’s strategy logs its in-process time and reports a median of $1.5\,\text{µ}\mathrm{s}$. The capture at the taps (the chapter’s simulated second) reports a wire-to-wire median of $4.4\,\text{µ}\mathrm{s}$ and a 99th percentile of $38\,\text{µ}\mathrm{s}$. Use the stage medians and 99th percentiles of [Figure 5.3](#fig-nw-capturing-and-monitoring-the-network-budget): network in 263 and out 295 ns; card receive 798 (p99 1 997) and transmit 801 (p99 2 000); the six software stages 437, 756, 158, 97, 47 and 33 ns at the median.

**Part I — Medians.**

1. What is the sum of the software stages’ medians?
2. What is the sum of every stage’s median?
3. Why is the sum of medians not the median of the sum?
4. How much of the wire-to-wire median lies outside the program?

**Part II — The tail.**

5. The ring’s 99th percentile is $31.9\,\text{µ}\mathrm{s}$ . What share of the wire-to-wire 99th percentile is that?
6. Why can a single slow stage own the tail?
7. What did Book 13 find the ring’s tail to be on the laptop?
8. What would a production host change?

**Part III — The capture.**

9. Why is the tap’s timestamp the right start for the latency, rather than the card’s?
10. The naive rule mispaired 47% of the orders. What would the 99th percentile have looked like with it?
11. What field would you use to carry the trigger in the order?
12. Which alarm would have told the firm that its tail had moved?

**Part IV — The verdict.**

13. State the *named result* : the wire-to-wire median and 99th percentile, the part outside the program, and the stage that owns the tail.
14. If the cards’ paths were halved, how much would the median fall?
15. If the ring’s tail were removed entirely, which stage would own the 99th percentile?
16. What would you need to measure the cards’ stages instead of assuming them?
17. Why does the capture belong to the firm’s evidence as well as its tools?
18. How much disk does an hour of this capture take, at the tutorial’s rates?
19. What would you keep for a month, and what for a day?
20. In one sentence: what does a tap see that a log does not?

**Solution of Problem 5.1.**

1. $437 + 756 + 158 + 97 + 47 + 33 = 1528\,\mathrm{n}\mathrm{s}$ .
2. $1\,528 + 263 + 295 + 798 + 801 = 3685\,\mathrm{n}\mathrm{s}$ .
3. The stages are skewed to the right: the median of a sum of skewed variables lies above the sum of their medians (here $4386\,\mathrm{n}\mathrm{s}$ in the budget).
4. $4.4 - 1.5 = 2.9\,\text{µ}\mathrm{s}$ : two thirds of the median.
5. $31.9 / 38.5 = 83\%$ .
6. A high percentile of the sum is dominated by the stage whose tail is heaviest; the other stages’ tails rarely coincide with it.
7. A 99.9th percentile of $3.5\,\mathrm{m}\mathrm{s}$ : the consumer descheduled on shared cores.
8. Isolated cores and a busy-polling consumer, which remove the scheduler from the ring.
9. The tap sees the frame where the firm receives it from the venue; the card’s timestamp comes after the cage’s network, so it misses part of the path, and the card’s clock is one more thing to trust.
10. Understated: the naive rule pairs slow orders with later packets, and gives a 99th percentile of $27\,\text{µ}\mathrm{s}$ instead of 38.
11. The client order identifier, or a field the venue returns unchanged.
12. Wire-to-wire percentiles computed from the capture every minute, compared with the baseline.
13. **Named result.** Wire-to-wire median $4.4\,\text{µ}\mathrm{s}$ and 99th percentile $38\,\text{µ}\mathrm{s}$ ; $2.9\,\text{µ}\mathrm{s}$ of the median lies outside the program; the ring between handler and strategy owns the tail.
14. By about $0.9\,\text{µ}\mathrm{s}$ (the budget’s median falls from 4.39 to 3.47).
15. Decode (p99 $2.8\,\text{µ}\mathrm{s}$ ), with the card’s two passes close behind at $2\,\text{µ}\mathrm{s}$ each.
16. Hardware timestamps at the card’s port on both sides, compared with the taps’, or a capture on each side of the server.
17. It is what the firm actually sent and received, and when, independent of any program’s logs: the record a venue, a client or a regulator asks for.
18. About 38 MB a second of capture, 138 GB an hour.
19. A month of summaries (latency percentiles, gaps, bursts, alarms) and of the order handoff’s frames; a day of every market-data frame, unless an incident asks for more.
20. Everything outside the program, timed by a clock the program cannot move.

## 5.9 Interview questions

**Interview question 5.1 ★ developer.**

How do you measure wire-to-wire latency? What do you need that a program’s own logs do not give you?

**Solution of Interview question 5.1.**

Tap the market-data handoff and the order handoff, timestamp both on the wire against one disciplined clock, pair each order with its [trigger packet](#def-nw-capturing-and-monitoring-the-network-trigger), and subtract. A program’s logs miss the network, the card and the kernel on both sides, and use the host’s clock.

*What the interviewer is looking for: taps, one clock, and pairing.*

**Interview question 5.2 ★★ developer.**

Where should a capture’s timestamp be taken, and why does it matter?

**Solution of Interview question 5.2.**

On the wire, at the first byte of the frame, by the tapping device, carried with the copy (a trailer). Anything later adds the aggregation network’s queueing and the storage server’s clock to every measurement.

*What the interviewer is looking for: timestamp at the tap, not at storage.*

**Interview question 5.3 ★★ developer, researcher.**

How do you pair each order with the market-data packet that caused it?

**Solution of Interview question 5.3.**

Make the strategy carry the trigger’s sequence number (or a hash) in the order, in a field the capture can read. Heuristics such as the latest trigger before the order fail in bursts, which is when latency matters.

*What the interviewer is looking for: instrumenting the order rather than guessing.*

**Interview question 5.4 ★★ developer.**

Which network alarms would you set for a trading cage, and why those?

**Solution of Interview question 5.4.**

Egress discards on ports that should never be full, gaps per line and on both lines, snapshot recoveries, clock lock and offsets, and wire-to-wire percentiles against a baseline: each points at a place on the path.

*What the interviewer is looking for: alarms tied to mechanisms, not a dashboard of everything.*

**Interview question 5.5 ★★★ developer.**

Your wire-to-wire 99th percentile doubled overnight; nothing was deployed. How do you find out why?

**Solution of Interview question 5.5.**

Decompose it: compare the capture’s stage timestamps (taps, card timestamps if kept) with the budget; check whether the market got busier (bursts, message rates), whether a clock moved ([holdover](https://one-course.com/books/quant/14/en/chapter/4-time-synchronisation#def-nw-time-synchronisation-holdover), a new asymmetry), whether a host changed (kernel, firmware, a new process on a shared core), and whether the venue changed its feed. Look at when it started to the second.

*What the interviewer is looking for: decomposition by stage and a list of non-deployment causes.*

**Interview question 5.6 ★★ developer, researcher.**

What would you use a capture for besides latency?

**Solution of Interview question 5.6.**

Gap and loss analysis, [microburst](https://one-course.com/books/quant/14/en/chapter/1-networking-for-trading#def-nw-networking-for-trading-burst) detection, rebuilding the book to check the handler, reconstructing what the firm sent during an incident, verifying venue acknowledgement times, and evidence for disputes and regulators.

*What the interviewer is looking for: the capture as evidence and as a second, independent implementation.*
