Quantitative Finance · Book 14 · Technology

Networks, Hardware and Trading Infrastructure

Networks, Hardware and Trading Infrastructure · Technology

7Programmable Hardware II: Trading Designs

The design of this chapter decides two clock cycles after the last byte of a price reaches it: 12.8 ns12.8\,\mathrm{n}\mathrm{s} at 156.25 MHz, for the first order and for the ten-thousandth, whether the market is quiet or bursting. It has no cache to miss and no thread to wake, and the order it sends leaves as a template with three fields patched in. It also cannot do much: it compares one price with one threshold, for one instrument, and everything it knows about the world was written into its registers by software a moment before. That combination (very fast, very narrow, steered by software) is what firms mean by a hybrid trading system.

This chapter builds a hardware trigger on the exchange simulator’s feed, with its risk checks and its order template, in SystemVerilog on the toolkit of chapter 6, and checks it four ways: in Verilator, in Icarus, against a C++20 cycle model and against a Python model, order for order and cycle for cycle. It then puts the trigger in the firm’s wire-to-wire path and asks what it buys.

7.1 Feed decoding in hardware

A software handler (One Quant Book 13, chapter 18) receives a whole datagram, then walks its messages. A hardware decoder walks the messages as the bytes arrive, eight a cycle, keeping only the state it needs between cycles: where it is in the packet’s header, the length of the current message, how far into it it is, and the fields it has captured so far. In the simulator’s feed, as in the exchange feed format it follows, an add-order message is 36 bytes with its price in the last four; the decoder therefore knows everything about an order the cycle its last byte arrives, while the rest of the packet (other messages, the frame’s check sequence) may still be to come.

    for (int i = 0; i < 8; i++) begin
      if (in_keep[7 - i]) begin
        b = in_data[63 - 8 * i -: 8];
        case (n_lstate)
          2'd0: if (n_ppos == 16'd19) n_lstate = 2'd1;
          2'd1: begin n_mlen[15:8] = b; n_lstate = 2'd2; end
          2'd2: begin n_mlen[7:0] = b; n_mpos = '0; n_lstate = 2'd3; end
          default: begin
            if (n_mpos == 16'd0) n_mtype = b;
            if (n_mpos == 16'd1) n_mloc[15:8] = b;
            if (n_mpos == 16'd2) n_mloc[7:0] = b;
            if (n_mpos == 16'd19) n_mside = b;
            if (n_mpos >= 16'd20 && n_mpos <= 16'd23) n_mqty = {n_mqty[23:0], b};
            if (n_mpos >= 16'd32 && n_mpos <= 16'd35) n_mprice = {n_mprice[23:0], b};
            if (n_mpos + 16'd1 == n_mlen) begin
              if (n_mtype == "A" && n_mlen == 16'd36 && n_mloc == LOC[15:0] && n_mside == "S"
                  && n_mprice <= thresh) begin
                n_trig = 1'b1;
                n_tqty = n_mqty;
                n_tprice = n_mprice;
              end
              n_lstate = 2'd1;
            end
            n_mpos = n_mpos + 16'd1;
          end
        endcase
        n_ppos = n_ppos + 16'd1;
      end
    end
Listing 7.1. The decoder’s combinational logic: a byte-serial state machine unrolled over the beat’s eight lanes. It captures the type, instrument, side, size and price, and at the message’s last byte decides whether to fire. code/firm/hwtrade/hdl/hwt_trigger.sv

Definition 7.1 (Hardware trigger, cut-through decision)

A hardware trigger is logic on a network card that watches a market-data stream for a condition set in advance by software (a price crossing a threshold, a trade on a given instrument) and, when it occurs, starts an order without the host’s involvement. A cut-through decision is a decision taken as soon as the bytes it depends on have arrived, before the frame that carries them has ended or been checked.

The unrolled state machine of Listing 7.1 is the simplest correct design, and it is slow in hardware: eight byte steps chained in one cycle make a long critical path. Faster decoders exploit the format: messages start at known offsets relative to their length prefixes, so dedicated logic can look for a message start in each lane at once, and fixed-layout messages can be matched by comparing whole words. That is the work of chapter 6’s timing closure, and it is why production decoders are specific to one venue’s format.

7.2 Book building in hardware

A trigger on single messages needs no book. A strategy that trades on the best bid or offer needs one, and a book in hardware is a different data structure from the ones of One Quant Book 13, chapter 19. Orders by reference need a hash table in block RAM or external memory; price levels need an array indexed by the price’s offset from a reference, with the best level found by a priority encoder over occupied levels; and every update must finish within the time between messages, or the design must queue them. Firms that build books in hardware keep only a few levels near the touch, for the few instruments that need them, and leave the rest to software.

7.3 Risk checks in cycles

A firm’s hardware sends orders on behalf of a broker-dealer that must keep pre-trade controls “under the direct and exclusive control” of the broker (One Quant Book 11, chapter 27), controls that reject orders exceeding “appropriate price or size parameters, on an order-by-order basis or over a short period of time”. In hardware those checks are comparisons and counters, and they cost a pipeline stage.

Definition 7.2 (Register map)

A register map is the set of addresses at which a hardware design exposes its configuration and state to software (thresholds, limits, enable and kill bits, counters), read and written over the host’s bus, and the document that defines them.

      // stage 2: risk, with the token bucket refilled every REFILL cycles
      pass <= 1'b0;
      if (since + 16'd1 == 16'(REFILL)) begin
        since <= '0;
        if (tokens < 16'(BURST)) tokens <= tokens + 16'd1;
      end else begin
        since <= since + 16'd1;
      end
      if (trig) begin
        if (kill || tqty > max_qty || tokens == 16'd0) begin
          rejects <= rejects + 16'd1;
        end else begin
          pass <= 1'b1;
          pqty <= tqty;
          pprice <= tprice;
          tokens <= tokens - 16'd1
                    + ((since + 16'd1 == 16'(REFILL) && tokens < 16'(BURST)) ? 16'd1 : 16'd0);
        end
      end
      // stage 3: the order template, patched
      order_valid <= pass && !kill;
      if (pass) begin
        order_id <= next_id;
        order_qty <= pqty;
        order_price <= pprice;
        if (!kill) next_id <= next_id + 32'd1;
      end
    end
Listing 7.2. The risk stage and the order template: a size limit, a kill bit and a token bucket, then the order’s id, size and price patched in. code/firm/hwtrade/hdl/hwt_trigger.sv

Proposition 7.3 (A token bucket’s burst)

A token bucket of capacity BB refilled with one token every RR cycles lets out at most B+⌊t/R⌋B + \lfloor t/R\rfloor orders in any interval of tt cycles. If software needs tt cycles to notice a runaway and set the kill bit, the hardware can send at most that many orders before it stops, and none after.

Proof. The bucket holds at most BB tokens at the start of the interval and gains at most one every RR cycles; each order consumes one. The kill bit masks the output stage in the cycle it is set. ∎

The proposition is the requirement document of a hardware risk layer. At 156.25 MHz, with the chapter’s bucket of 4 tokens refilled every 64 cycles, a software monitor that reacts in 20 µs20\,\text{µ}\mathrm{s} lets out at most 52 orders; one that reacts in 100 µs100\,\text{µ}\mathrm{s}, 248. The limits that matter in hardware are those that bind before software can act.

7.4 Order generation from templates

Definition 7.4 (Order template)

An order template is an order message prepared in advance, complete except for the fields that depend on the trigger (client order identifier, size, price), which the hardware patches in as it sends.

Definition 7.5 (TCP offload engine)

A TCP offload engine implements the TCP protocol (sequence numbers, acknowledgements, retransmission) in a network card’s hardware, so that the card can send and receive on a TCP connection without the host.

Order entry runs over TCP (chapter 1), so a card that sends orders by itself must own the connection’s state: its next sequence number, its acknowledgements, its retransmissions. A TCP offload engine does that; the order template then carries the TCP and IP headers too, with sequence number and checksums patched at send time. The software that owns the session (logon, heartbeats, the venue’s session layer) shares the connection with the hardware through the same engine.

As of September 2026 — A published hardware tick-to-trade figure

In a STAC-T0 benchmark reported in June 2024, an FPGA system built on an AMD Alveo UL3524 card with a TCP and UDP core (an Exegy and AMD solution) was measured at a minimum of 13.9 ns13.9\,\mathrm{n}\mathrm{s} for 507-byte frames and 14.1 ns14.1\,\mathrm{n}\mathrm{s} for 68-byte frames, from “the last bit of inbound data needed to make a trading decision to the first bit of the simulated outbound order”. The definition excludes the time for the inbound bits to arrive, which the chapter’s design, and any design, cannot avoid.

7.5 Hybrid designs: what stays in software and why

Definition 7.6 (Hybrid trading system)

A hybrid trading system splits a strategy between software, which computes its model, positions and parameters at its own pace, and hardware, which applies precomputed decisions to the market data at line rate; software writes the hardware’s registers and reads its counters and copies of what it sent.

The split follows from what each side is good at. Hardware is fast and fixed: every feature costs logic, every change costs a build of hours and a verification of days, and it cannot hold much state. Software is slow and flexible: it can model, learn, net positions across venues and change within minutes. So software decides what would be worth doing (“buy up to 700 at 100.02 or better if anyone offers”) and hardware decides when (the cycle the offer arrives). Everything that needs history, many instruments or judgement stays in software; the one comparison that must happen first moves to the card.

The chapter’s hybrid design. Three pipeline stages on the card decode the feed, check each candidate order and patch it into a template; software on the host writes the thresholds, limits and kill bit through the register map and receives copies of everything the card sends.
Figure 7.1. The chapter’s hybrid design. Three pipeline stages on the card decode the feed, check each candidate order and patch it into a template; software on the host writes the thresholds, limits and kill bit through the register map and receives copies of everything the card sends.
One packet with one add-order message (58 bytes: 20 of header, 2 of length, 36 of message) through the design, one beat of 8 bytes a cycle. The price’s last byte arrives in beat 7; the trigger is registered at the end of that cycle, the risk decision one cycle later, and the order in the cycle after: two cycles after the price, ten after the packet’s first byte.
Figure 7.2. One packet with one add-order message (58 bytes: 20 of header, 2 of length, 36 of message) through the design, one beat of 8 bytes a cycle. The price’s last byte arrives in beat 7; the trigger is registered at the end of that cycle, the risk decision one cycle later, and the order in the cycle after: two cycles after the price, ten after the packet’s first byte.

7.6 Tutorial: a trigger on the simulator’s feed, four ways

Goal. Run the trigger on two seconds of the exchange simulator’s feed in Verilator, in Icarus, in a C++20 cycle model and in a Python model, check that all four send the same orders on the same cycles, and put the design in the wire-to-wire path. End state: Table 7.1, Figure 7.3 and the green tests of firm.hwtrade.

  1. The feed. make_hwtrade_fixture.py runs One Quant Book 10’s simulator for two seconds from the open (992 packets) and stores them; beats_of cuts each into beats with two idle cycles between packets (the day compressed: the bucket counts cycles, not seconds).
  2. Four models. build_verilator, build_icarus, CycleModel and cpp/firm_hwtrade.hpp (Listing 7.3) run the same settings; the tests compare their outputs line for line.
  3. The reference. triggers() parses the same packets message by message in Python: every qualifying message must be either an order or a reject.
  4. The path. nw_hwtrade.paths() builds the hardware variant of firm.wirepath (the cage of chapter 2, the card’s transceivers, three cycles) next to chapter 5’s software path, and win_probability races each against a competitor; python fig_hwtrade.py writes the chart’s data.

What to change next. Set the bucket to one token refilled every 10 000 cycles and count the orders; move the kill bit to the cycle before the tenth order and check that exactly nine leave.

    void step(long cycle, const std::optional<Beat>& beat, std::uint32_t thresh, std::uint32_t max_qty,
              bool kill) {
        bool n_trig = false;
        std::uint32_t n_qty = 0, n_price = 0;
        if (beat) parse(*beat, thresh, n_trig, n_qty, n_price);
        if (passed_ && !kill) orders.push_back({cycle, next_id_++, pqty_, pprice_});   // stage 3
        const bool refill_now = since_ + 1 == refill_;                                     // stage 2
        std::uint16_t new_tokens = (refill_now && tokens_ < burst_) ? tokens_ + 1 : tokens_;
        since_ = refill_now ? 0 : since_ + 1;
        bool new_pass = false;
        if (trig_) {
            if (kill || tqty_ > max_qty || tokens_ == 0) {
                ++rejects;
            } else {
                new_pass = true, pqty_ = tqty_, pprice_ = tprice_;
                new_tokens = tokens_ - 1 + ((refill_now && tokens_ < burst_) ? 1 : 0);
            }
        }
        tokens_ = new_tokens, passed_ = new_pass;
        trig_ = n_trig, tqty_ = n_qty, tprice_ = n_price;                                   // stage 1
    }

    std::vector<Order> orders;
    std::uint32_t rejects = 0;
Listing 7.3. The C++20 cycle model’s clock edge: stage 3, then stage 2, then stage 1, each reading the registers the previous edge wrote, as the hardware does. code/firm/hwtrade/cpp/firm_hwtrade.hpp
Threshold (price ×104\times 10^4)Maximum sizeKill set at cycleOrdersRejects
1 000 20010 000never1085
1 000 30010 000never13327
1 000 300600never13129
1 000 30010 0003 00041119
Table 7.1. The trigger on two seconds of the simulator’s feed (instrument 1; sell-side add orders at or below the threshold), with a bucket of 4 tokens refilled every 64 cycles. Verilator, Icarus, the C++20 cycle model and the Python model give the same orders on the same cycles; orders and rejects add up to the qualifying messages the Python reference finds. Data: firm.hwtrade’s fixture.

The table shows each control at work. A lower threshold fires less often; the bucket rejects most of the extra triggers of the higher one, because they come in bursts; the size limit rejects two more; and the kill bit, set at cycle 3 000, stops every later order. In the wire-to-wire path of chapter 5 (Figure 7.3), the hardware path is the cage’s fibre, the card’s transceivers (120 ns120\,\mathrm{n}\mathrm{s} each way, a model value) and three cycles: 0.82 µs0.82\,\text{µ}\mathrm{s} at every quantile. The software path has a median of 4.4 µs4.4\,\text{µ}\mathrm{s} and a 99th percentile of 35 µs35\,\text{µ}\mathrm{s}. Against a competitor whose latency has a median of 1 µs1\,\text{µ}\mathrm{s} (a lognormal of spread 0.3, a model), the hardware path wins 75% of races and the software path one in ten thousand.

Simulation: wire-to-wire latency of the software path of chapter 5 (Book 13’s measured stages, the card and the cage) and of the hardware path (the cage, the card’s transceivers at 120\, n s each way, a model value, and three cycles at 156.25 MHz). Left: quantiles; the hardware path is a constant. Right: the probability of beating one competitor whose latency is lognormal with the given median and a spread of 0.3 (a model). Data: fig_hwtrade.py.
Figure 7.3. Simulation: wire-to-wire latency of the software path of chapter 5 (Book 13’s measured stages, the card and the cage) and of the hardware path (the cage, the card’s transceivers at 120 ns120\,\mathrm{n}\mathrm{s} each way, a model value, and three cycles at 156.25 MHz). Left: quantiles; the hardware path is a constant. Right: the probability of beating one competitor whose latency is lognormal with the given median and a spread of 0.3 (a model). Data: fig_hwtrade.py.

7.7 Build: the hardware trigger

Purpose. A complete, tested hardware trigger: feed decoding, risk checks and an order template, with the three models that make it trustworthy, and a hardware variant of the firm’s wire-to-wire path for chapter 29’s plan.

Interface. hdl/hwt_trigger.sv (parameters LOC, BURST, REFILL; registers thresh, max_qty, kill; outputs the order’s id, size, price and a reject counter); testbenches tb/hwt_tb.cpp and tb/hwt_tb.sv; C++20 firm::hwtrade::Model; Python firm_hwtrade: beats_of, write_stim, CycleModel, triggers, build_verilator, build_icarus, run_verilator, run_icarus, parse_output, wirepath_stage.

Rules. One order per qualifying message at most; nothing leaves while the kill bit is set; sizes above the limit are rejected, not clipped; the bucket never lends; the models agree with the HDL register for register.

Acceptance tests. code/firm/hwtrade/: the Python model reproduces the fixture’s four settings, and every qualifying message is an order or a reject; Verilator and Icarus print the same orders as the Python model; the C++20 model reproduces the fixture; the size limit at its boundary; nothing after the kill bit; the bucket’s rationing; two cycles from price to order.

Stretch. Keep the best offer of the instrument from add, execute and delete messages and fire on it; move the three checks into one cycle and measure the critical path it would need.

Sources and further reading

  • STAC, “STAC-T0” results with an Exegy and AMD FPGA solution (2024).
  • 17 CFR 240.15c3-5 (the market access rule).
  • Nasdaq, TotalView-ITCH 5.0 specification (Add Order layout); the simulator’s code/firm/exchsim/PROTOCOL.md.

7.8 Exercises

Exercise 7.1 ★

A packet carries one 36-byte add order behind the 20-byte header and the 2-byte length. In which beat is the price’s last byte, and how many nanoseconds after the packet’s first beat does the design emit the order at 156.25 MHz?

Solution

Solution of Exercise 7.1.

The price ends at byte 20+2+35=5720 + 2 + 35 = 57, in beat 7 (bytes 56 to 63). The order is registered two cycles later, at the end of cycle 9: ten cycles from the start of the packet’s first beat, 10×6.4=64 ns10 \times 6.4 = 64\,\mathrm{n}\mathrm{s}.

Exercise 7.2 ★

With a bucket of 4 tokens refilled every 64 cycles at 156.25 MHz, what is the sustained maximum rate of orders?

Solution

Solution of Exercise 7.2.

One token every 64 cycles: 156.25×106/64≈2.44156.25 \times 10^6 / 64 \approx 2.44 million orders a second sustained, after an initial burst of 4.

Exercise 7.3 ★

How many orders can the chapter’s design send before a software monitor that reacts in 50 µs50\,\text{µ}\mathrm{s} sets the kill bit?

Solution

Solution of Exercise 7.3.

50×156.25=7 812.550 \times 156.25 = 7\,812.5 cycles; 4+⌊7 812.5/64⌋=4+122=1264 + \lfloor 7\,812.5 / 64\rfloor = 4 + 122 = 126 orders.

Exercise 7.4 ★★

In Table 7.1, why does raising the threshold from 1 000 200 to 1 000 300 add 25 orders but 22 rejects?

Solution

Solution of Exercise 7.4.

The 47 extra qualifying messages (160 against 113) arrive in bursts, when the offers near the threshold are being placed: the bucket, already drained by the first orders of each burst, rejects 22 of them and lets 25 through.

Exercise 7.5 ★★

Why does the design reject orders above the size limit instead of sending them at the limit?

Solution

Solution of Exercise 7.5.

An order at the limit is not the order the strategy computed; clipping silently changes risk and hides a configuration or model error. Rejecting it keeps the hardware’s behaviour simple and auditable and makes the error visible in the reject counter.

Exercise 7.6 ★★

Name three decisions the chapter’s strategy leaves in software, and why each would be hard in hardware.

Solution

Solution of Exercise 7.6.

The threshold itself (a model over many inputs and history), the position and exposure across venues (state shared with other systems), and whether to trade at all (news, halts, the desk’s limits). Each needs state, arithmetic or judgement that would cost logic and weeks of verification for every change.

Exercise 7.7 ★★★

Coding. With firm.hwtrade.CycleModel, set the bucket to 1 token refilled every 10 000 cycles and count the orders at threshold 1 000 300. Then set the kill bit one cycle before the tenth order of the unrestricted run and count again.

Solution

Solution of Exercise 7.7.

One order: the fixture lasts fewer than 10 000 cycles, so the bucket never refills after its single token. With the kill bit set one cycle before the tenth order of the unrestricted run, exactly nine orders leave.

Exercise 7.8 ★★★

Find the flaw. “Our hardware fires in 13 nanoseconds, so our wire-to-wire latency is 13 nanoseconds.”

Solution

Solution of Exercise 7.8.

The 13 nanoseconds (a benchmark’s figure) runs from the last bit needed for the decision to the first bit of the order. Wire to wire adds the time for the triggering bits to arrive, the card’s transceivers, and the cage’s fibre and devices both ways: hundreds of nanoseconds in the chapter’s cage.

7.9 Problem: The Trigger That Must Not Fire Twice

Problem 7.1

Weekend problem — a hybrid path, a race and a runaway

A firm moves the decision of a latency-sensitive strategy onto the chapter’s hardware trigger. Use the chapter’s numbers: the software path’s wire-to-wire median 4.39 µs4.39\,\text{µ}\mathrm{s} and 99th percentile 35.3 µs35.3\,\text{µ}\mathrm{s}; the hardware path 0.82 µs0.82\,\text{µ}\mathrm{s} (263 ns263\,\mathrm{n}\mathrm{s} of cage in, 120 ns120\,\mathrm{n}\mathrm{s} of card in, three cycles at 156.25 MHz, 120 ns120\,\mathrm{n}\mathrm{s} of card out, 295 ns295\,\mathrm{n}\mathrm{s} of cage out); a competitor with a lognormal latency; a bucket of 4 tokens refilled every 64 cycles.

Part I — The path.

  1. Add up the hardware path’s stages.
  2. How much of it is the decision itself?
  3. What is the hardware path’s 99th percentile, and why?
  4. By how much does the hardware path beat the software path at the median, and at the 99th percentile?

Part II — The race.

  1. Against a competitor with a median of 1 µs1\,\text{µ}\mathrm{s}, what are the two paths’ probabilities of winning (the chapter’s simulation)?
  2. And against one with a median of 5 µs5\,\text{µ}\mathrm{s}?
  3. Which part of the hardware path would you shorten next, and what is left once the decision is gone?
  4. What does the race model leave out?

Part III — The runaway.

  1. A software monitor needs 20 µs20\,\text{µ}\mathrm{s} to see a runaway and set the kill bit. How many orders can leave meanwhile?
  2. With a monitor on the card itself, reacting in 16 cycles?
  3. Why does the kill bit act on stage 3 rather than only on stage 2?
  4. What limits would you add to the card, and which should stay in the broker’s layer?

Part IV — The verdict.

  1. State the named result: the two paths’ latencies, their probabilities of winning against a competitor with a median of 1 µs1\,\text{µ}\mathrm{s}, and the orders the bucket lets out before a 20 µs20\,\text{µ}\mathrm{s} kill.
  2. Is the 0.82 µs0.82\,\text{µ}\mathrm{s} path worth building for a strategy that races competitors at 5 µs5\,\text{µ}\mathrm{s}?
  3. What does the firm give up by moving the decision into hardware?
  4. How often can the threshold change, and who changes it?
  5. Why must the software receive copies of every order the card sends?
  6. How would you test the design against a recorded day before it trades?
  7. What would a failure of the register map’s write path do, and how would the design detect it?
  8. In one sentence: what does hardware decide, and what does software decide?
Solution

Solution of Problem 7.1.

  1. 263+120+19.2+120+295≈817 ns263 + 120 + 19.2 + 120 + 295 \approx 817\,\mathrm{n}\mathrm{s}: 0.82 µs0.82\,\text{µ}\mathrm{s}.
  2. Three cycles, 19.2 ns19.2\,\mathrm{n}\mathrm{s}: about 2%.
  3. 0.82 µs0.82\,\text{µ}\mathrm{s}: every stage of the model is a constant (fibre, fixed transceiver latency, fixed cycles).
  4. By 3.57 µs3.57\,\text{µ}\mathrm{s} at the median and 34.5 µs34.5\,\text{µ}\mathrm{s} at the 99th percentile.
  5. 75% for the hardware path, 0.01% for the software path.
  6. 100% and 57%.
  7. The cage’s fibre and the transceivers: what is left is physics (distance) and serialisation, chapters 9 to 14.
  8. The competitor’s correlation with the firm (both see the same bursts), queueing at the venue’s gateway, and anything but a single race.
  9. 4+⌊3 125/64⌋=524 + \lfloor 3\,125/64\rfloor = 52 orders.
  10. 4+⌊16/64⌋=44 + \lfloor 16/64\rfloor = 4 orders: the bucket’s burst.
  11. So that an order already past the risk stage is stopped in its last cycle too: the kill bit then blocks everything from the cycle it is set.
  12. On the card: size, price collar against a reference, order and notional rates, the kill bit. In the broker’s layer, under its control: credit and capital thresholds, and its own kill switch.
  13. Named result. 0.82 µs0.82\,\text{µ}\mathrm{s} wire to wire in hardware against a median of 4.39 µs4.39\,\text{µ}\mathrm{s} in software; against a competitor at 1 µs1\,\text{µ}\mathrm{s} they win 75% and 0.01% of races; a 20 µs20\,\text{µ}\mathrm{s} kill lets at most 52 orders out.
  14. Hardly: at 5 µs5\,\text{µ}\mathrm{s} the software path already wins 57% and the rest is spread across its tail; better to fix the tail.
  15. Flexibility, speed of change and much of the strategy’s expressiveness; it gains a verification burden.
  16. As often as software can compute and write it, within microseconds; the strategy’s software owns it, under the risk layer’s limits.
  17. To know its position and to reconcile with the venue’s acknowledgements: the card acts alone, and the firm must not lose track of what it did.
  18. Replay the recorded feed through the HDL in simulation and through the card in a test harness, compare both with the models, and run the venue’s certification.
  19. Stale thresholds or limits: the design should hold a version number and a heartbeat from software, and stop firing when they are stale.
  20. Hardware decides when; software decides what, and whether at all.

7.10 Interview questions

Interview question 7.1 ★ developer

How does an FPGA parse a market-data feed while the packet is still arriving?

Solution

Solution of Interview question 7.1.

It processes a fixed number of bytes per cycle with a state machine that tracks the position in the header and in each length-prefixed message, capturing fields as their bytes pass; a message’s fields are known the cycle its last byte arrives, and decisions can start before the packet ends or is checked.

What the interviewer is looking for: streaming state, not buffering, and deciding before the frame ends.

Interview question 7.2 ★★ developer

What risk checks would you put in hardware, and how do you prove they cannot be bypassed?

Solution

Solution of Interview question 7.2.

Size, price collar, rate limits (a token bucket), duplicate detection and a kill bit, each a comparison or counter in a pipeline stage that every order must cross. Prove it by construction (the only path to the output goes through the stage), by simulation of every boundary, and by keeping the broker’s controls independent of the firm’s configuration.

What the interviewer is looking for: a single path to the wire and independent controls.

Interview question 7.3 ★★ developer, trader

What stays in software in a hybrid trading system, and why?

Solution

Solution of Interview question 7.3.

Models, parameters, positions, cross-venue risk and anything needing history or judgement: software computes what would be worth doing, hardware applies it at the moment the market offers it. Hardware changes are slow and costly to verify; software changes are cheap.

What the interviewer is looking for: what against when, and the cost of change.

Interview question 7.4 ★★ developer

How do you send an order over TCP from an FPGA?

Solution

Solution of Interview question 7.4.

With a TCP offload engine on the card that owns the connection’s sequence numbers, acknowledgements and retransmissions, and an order template whose TCP and IP headers and checksums are patched at send time; the session layer (logon, heartbeats) is shared with software.

What the interviewer is looking for: connection state in hardware and a template with patched fields.

Interview question 7.5 ★★★ developer

How would you verify that the hardware trigger and its software model agree, and what would you do when they disagree?

Solution

Solution of Interview question 7.5.

Run both on the same recorded and synthetic inputs, in two simulators, and compare every output with its cycle; add message-level reference checks (each qualifying message decided once). A disagreement is a bug in one of them until shown otherwise: reduce the input to the smallest case and read the waveform.

What the interviewer is looking for: cycle-level comparison and disciplined reduction.

Interview question 7.6 ★★ developer, trader

A vendor quotes 14 nanoseconds tick to trade. What questions do you ask?

Solution

Solution of Interview question 7.6.

From which bit to which bit it is measured, on which frame sizes and message types, whether it includes the transceivers and the order’s TCP, under what load and with what risk checks enabled, and how the figure maps to my wire-to-wire path.

What the interviewer is looking for: the definition of the measurement.

Terms defined in this chapter

See all 2333 terms in the glossary