Markets I: The Ecosystem and Exchange-Traded Markets · Markets
28Reading Market Data
Before lunch an exchange has sent one firm tens of millions of messages about thousands of instruments. Each says that an order was added, reduced, executed or removed, in a few dozen bytes. None says what the best bid is: that is for the receiver to work out, by applying every message, in order, to a book it maintains itself. If one message is lost, the book is wrong until the end of the day, silently. If one trade report is a typing error, a price of 10.01 in a stock worth 100, the day’s low, its volatility, its volume-weighted average price and every signal computed from them are wrong, also silently. Every result in the rest of this series stands on a table of prices; this chapter is about how that table is made, and about the defects it inherits.
28.1 Three levels of data
Definition 28.1 (Level 1, level 2, level 3)
Level 1 data is the best bid and offer with their sizes, and the last trade. Level 2 data is the book aggregated by price: for each price level, the total size (and sometimes the number of orders), to some depth or in full: the market-by-price feed of Definition 19.9. Level 3 data is every order individually, with its identifier and every event in its life: the market-by-order feed, from which the other two can be computed and which alone gives queue positions.
Definition 28.2 (Sequence number)
A sequence number is a counter that the publisher increments with each message on a channel. A receiver that sees a jump has lost messages and must recover them, from a retransmission service or a snapshot, before its book can be trusted again.
A level 3 feed is a stream of differences. The state is on the receiver’s side, which makes the protocol compact and makes every receiver responsible for correctness. Direct feeds are published on two redundant network paths; the receiver takes whichever copy of each sequence number arrives first and notices when neither does.
As of September 2026 — One wire format
The order-by-order equity feed whose version 5.0 specification is used as this chapter’s model encodes all integers big-endian; prices as integers with four implied decimal places; timestamps as nanoseconds since midnight in six bytes; and instruments by a two-byte locate code assigned for the day, at the same offset in every message. An add-order message has 36 bytes: type, locate, tracking number, timestamp, an eight-byte order reference, side, shares, an eight-character symbol and the price. An execution has 31 bytes and refers to the order only by its reference, as do the cancel (23 bytes) and the delete (19 bytes): the price of a trade is not in the execution message, it is wherever the receiver stored the order.
28.2 Trades, quotes and condition codes
Definition 28.3 (Trade condition)
A trade condition (or sale condition) is a code attached to a trade report that says what kind of trade it was and how it may be used: a regular trade; an opening or closing auction print; an odd lot; a trade reported late or out of sequence; a trade at an average or derived price; a correction or a cancellation of an earlier report. The consolidated tapes and each exchange define which conditions update the last price, the high and low, and the volume.
A tape is not a list of prices at which one could have traded. An out-of-sequence report carries a price from minutes ago. An average-price trade matches no quote that ever existed. A cancelled trade did not happen; its cancellation arrives as a later message that a careless loader ignores. An auction print is one price for a large share of the day’s volume. Off-exchange trades (Chapter 10) are reported after the fact, seconds late, with a timestamp that may be the reporter’s. Before any statistic is computed from trades, a rule must say which conditions are in.
Method 28.4 (Cleaning a tape)
- Apply corrections and cancellations, matching them to the original reports.
- Keep or drop each condition explicitly, by purpose: all executed volume for a volume total; only regular, in-sequence trades for a last price, a high or a low; auctions separately.
- Flag prints far from the prevailing quote or from a rolling median; inspect before deleting: a flash crash is data, a decimal error is not.
- Record what was removed and why. The rule is part of the result.
28.3 Timestamps
Definition 28.5 (Exchange and receive timestamps)
The exchange timestamp is written by the publisher’s matching engine or gateway when the event occurs. The receive timestamp is written by the recipient’s capture system when the message arrives. The first orders events within one venue; only the second says what the recipient could have known, and when.
A backtest that uses exchange timestamps assumes the firm learned of each event at the instant it happened. The difference, tens of microseconds in the same building and milliseconds between cities, is the whole of some strategies’ edge (One Quant Book 10). Two venues’ exchange clocks are not the same clock: comparing their timestamps to decide which quote moved first is valid only to the accuracy of their synchronisation.
As of September 2026 — How accurate clocks must be
European Union. The regulatory technical standard on clock synchronisation requires members of a trading venue that use high-frequency algorithmic techniques to keep their business clocks within 100 microseconds of UTC and to timestamp with a granularity of a microsecond or better. United States. The consolidated audit trail plan requires exchanges to keep their clocks within 100 microseconds of the national standards institute’s time, and broker-dealers within 50 milliseconds (one second for clocks used only for manual events).
28.4 Symbology and reference data
Definition 28.6 (Symbology and reference data)
Symbology is the set of identifiers under which one instrument is known: exchange ticker, vendor codes, national and international security numbers, a feed’s numeric code for the day. Reference data is everything about the instrument that is not a price: identifiers and their history, listing venue, currency, tick size, lot size, multiplier, expiry, corporate actions, trading calendar.
A ticker is not an identifier. Tickers are reused after delistings, changed on mergers, and differ by venue and by vendor; a feed’s locate code is valid for one day. Research tables must be keyed by a permanent internal identifier, with every external symbol mapped to it by date range. The commonest silent error in equity research, after look-ahead, is a time series that is two different companies joined at the date one took the other’s ticker. The builds of chapters 8, 18 and 23 are this chapter’s reference data.
28.5 From raw feed to research table
Definition 28.7 (Normalisation)
Normalisation is the conversion of each venue’s native messages into one internal representation: common message types, integer prices in a common unit, a common clock field, permanent instrument identifiers, explicit condition flags.
The pipeline has four stages, and research should know which one its table came from. Capture: every packet, with a hardware receive timestamp, written unmodified. Decode and normalise: the build below. Book building: apply messages in sequence, detect gaps, emit level 2 and level 1 as derived streams. Tables: trades with conditions and aggressor side; quotes; bars, which are a lossy summary chosen for a purpose. Each stage must be replayable from the one before, bit for bit: when a result looks wrong, the question “is it the data?” must have an answer.
28.6 Tutorial: a binary feed into a book and a tape
Goal. Decode a framed binary order-by-order feed, maintain the book, produce level 1 and a tape with aggressor side, and measure what receive-time jitter does to message order. End state: the three data figures of this chapter, from code/firm/feed/data/sample.itch (2 020 messages, 65 842 bytes).
Framing and header. Length prefix, type, then the fields every message shares. A length that does not fit the type is an error, not a hint.
def decode(buf: bytes): """Yield messages from a framed buffer. Raises on a truncated frame or a length that does not match the message type: a decoder that guesses is worse than one that stops.""" i = 0 while i < len(buf): if i + 2 > len(buf): raise ValueError("truncated length prefix") (n,) = struct.unpack_from(">H", buf, i) body = buf[i + 2:i + 2 + n] if len(body) != n: raise ValueError("truncated message") kind = body[:1] if LENGTHS.get(kind) != n: raise ValueError(f"type {kind!r} with length {n}") locate, tracking = struct.unpack_from(">HH", body, 1) ts = int.from_bytes(body[5:11], "big") (ref,) = struct.unpack_from(">Q", body, 11)Listing 28.1. Python: frame, check, read the common header. code/firm/feed/firm_feed.py The same in C++, over a span of bytes, without copies.
inline void decode(std::span<const std::uint8_t> buf, const std::function<void(const Msg&)>& on) { std::size_t i = 0; while (i < buf.size()) { if (i + 2 > buf.size()) throw std::runtime_error("truncated length prefix"); const std::size_t n = static_cast<std::size_t>(be(buf, i, 2)); if (i + 2 + n > buf.size()) throw std::runtime_error("truncated message"); const auto body = buf.subspan(i + 2, n); Msg m; m.kind = static_cast<char>(body[0]); if (expected_length(m.kind) != n) throw std::runtime_error("length does not match type"); m.locate = static_cast<std::uint16_t>(be(body, 1, 2)); m.ts = be(body, 5, 6); m.ref = be(body, 11, 8);Listing 28.2. C++: the decode loop’s framing and header. code/firm/feed/cpp/firm_feed.hpp And in Rust, as an iterator that yields a result per message and stops after the first error.
fn next(&mut self) -> Option<Self::Item> { if self.failed || self.buf.is_empty() { return None; } let fail = |s: &mut Self, e| { s.failed = true; Some(Err(e)) }; if self.buf.len() < 2 { return fail(self, DecodeError::TruncatedPrefix); } let n = be(&self.buf[..2]) as usize; if self.buf.len() < 2 + n { return fail(self, DecodeError::TruncatedMessage); } let body = &self.buf[2..2 + n]; let kind = body.first().copied().unwrap_or(0); if expected_length(kind) != Some(n) { return fail(self, DecodeError::BadLength { kind, len: n }); } let mut m = Msg { kind, locate: be(&body[1..3]) as u16, ts: be(&body[5..11]), reference: be(&body[11..19]), ..Msg::default() };Listing 28.3. Rust: one step of the decoder. code/firm/feed/rust/src/lib.rs The book. An execution carries no price: it is looked up.
def apply(self, m: Msg) -> None: if m.ts < self.last_ts: self.errors.append(f"timestamp went backwards at ref {m.ref}") self.last_ts = max(self.last_ts, m.ts) if m.kind == "A": if m.ref in self.orders: self.errors.append(f"duplicate order {m.ref}") return self.orders[m.ref] = Order(m.side, m.shares, m.price, m.locate) lv = self._level(m.locate, m.side) lv[m.price] = lv.get(m.price, 0) + m.shares elif m.kind == "E": before = self.orders.get(m.ref) price, side = (before.price, before.side) if before else (0, "") if self._reduce(m.ref, m.shares) is not None: self.trades.append(Trade(m.ts, m.locate, price, m.shares, m.match, "S" if side == "B" else "B")) elif m.kind == "X": self._reduce(m.ref, m.shares) elif m.kind == "D": o = self.orders.get(m.ref) self._reduce(m.ref, o.shares if o else 0) elif m.kind == "P": self.trades.append(Trade(m.ts, m.locate, m.price, m.shares, m.match, ""))Listing 28.4. Applying one message to the book and the tape. code/firm/feed/firm_feed.py
What to change next. Drop one add-order message from the sample and watch the errors the book records, and how long it takes before an execution refers to the missing order. Then add a replace message (cancel and add with a new reference, losing priority) to all three decoders.
28.7 Build: the feed normaliser
Purpose. The entry point of all market data into the miniature firm: the simulated exchange of One Quant Book 10 publishes this format, the strategies of Book 11 consume the book, and the research tables of Book 7 are written from the tape.
Interface. Python: encode(msg), decode(buffer), Book.apply(msg), Book.best(locate), Book.depth(locate, side, n), Book.trades, Book.errors, summary(book, locate). C++20 (header only): firm::feed::decode(span, callback), Book, Summary. Rust: decode(&[u8]) returning an iterator of results, Book, Summary.
Rules. Big-endian integers, prices in ten-thousandths, 48-bit nanosecond timestamps, two-byte length framing. A truncated frame or a length inconsistent with the type stops decoding with an error. The book never guesses: an unknown order reference, an over-execution, a duplicate add or a timestamp that goes backwards is recorded and counted. The three implementations must produce the same summary on the shared sample.
Acceptance tests. code/firm/feed/tests/ (round trip, garbage, a book by hand, recorded errors, reproducibility of the sample); cpp/firm_feed_test.cpp and the Rust crate’s tests (the shared sample’s summary, truncation, a bad length, an unknown order).
Stretch. Sequence numbers with gap detection and A/B line arbitration; a snapshot message for recovery; a benchmark of messages per second in each language (One Quant Book 13 starts from there).
Sources and further reading
- Nasdaq, Nasdaq TotalView-ITCH 5.0 specification (revision of 28 April 2023).
- Commission Delegated Regulation (EU) 2017/574 (clock synchronisation, “RTS 25”).
- CAT NMS Plan, frequently asked questions on clock synchronisation (R1); FINRA Rule 6820.
28.8 Exercises
Exercise 28.1 ★
From the level 3 view of Figure 28.1 an execution of 300 shares arrives for order 1001, then a cancel of 100 shares for order 1002. Give levels 2 and 1 afterwards. At what price did the trade print and who was the aggressor?
Solution
Solution of Exercise 28.1.
Order 1001 disappears (300 executed of 300); order 1002 goes from 200 to 100. Level 2: bid 99.99 for 100 in 1 order, 99.98 for 100; asks unchanged. Level 1: 99.99 100 bid, 100.01 500 offered, last 99.99 300. The trade printed at 99.99, the resting order’s price, which the execution message does not carry; the resting order was a bid, so the aggressor was a seller.
Exercise 28.2 ★
A price field contains the four bytes 00 0F 41 DC (hexadecimal), big-endian, with four implied decimals. Give the price. What would a little-endian reader make of it?
Solution
Solution of Exercise 28.2.
0x000F41DC : $99.9900. Read little-endian the same bytes are : $369 525.12. Byte order errors are usually that obvious; an error in the number of implied decimals (99.99 read as 9 999) is not, in a feed that mixes instruments.
Exercise 28.3 ★
The chapter’s sample has 1 102 adds, 311 executions, 114 cancels, 473 deletes and 20 hidden trades. Verify the file’s size from the message lengths and the framing.
Solution
Solution of Exercise 28.3.
Each message costs its length plus 2 bytes of framing: bytes.
Exercise 28.4 ★★
A timestamp of six bytes counts nanoseconds since midnight. What is the largest time of day it can hold? Why is four bytes not enough, and what is gained over eight?
Solution
Solution of Exercise 28.4.
nanoseconds is 78.2 hours: more than a day, with room for sessions that cross midnight. Four bytes hold only 4.3 seconds of nanoseconds. Against eight bytes, six save two bytes on every message, a few percent of a feed of tens of millions of messages a day, at the cost of an unaligned field that every decoder must assemble by hand.
Exercise 28.5 ★★
Venue A timestamps a quote change at 10:00:00.000 150 and venue B at 10:00:00.000 190. Both venues’ clocks are within 100 microseconds of UTC. Can you say which moved first? What would you need?
Solution
Solution of Exercise 28.5.
No. The difference is 40 microseconds and each clock may be off by up to 100, in either direction: the true order is undetermined. To say which moved first one needs one clock for both: receive timestamps from a single capture point with known, stable latencies to each venue, or venues synchronised an order of magnitude more tightly than the gap being measured.
Exercise 28.6 ★★
A receiver sees sequence numbers 5 001, 5 002, 5 004 on line A and 5 001, 5 003, 5 004 on line B. What does it do? Now both lines skip 5 003: what must happen before the book is used again, and what may the strategy do meanwhile?
Solution
Solution of Exercise 28.6.
Line arbitration: it takes 5 001 from whichever line delivers first, 5 002 from A, 5 003 from B, 5 004 from either, and the book is complete. If both lines skip 5 003 the message is lost: request a retransmission, or wait for the next snapshot and rebuild, buffering later messages meanwhile. Until then the book for the affected instruments is marked stale; the strategy must stop quoting from it, since it no longer knows its own queue positions or even the best price, and may only reduce risk using a slower, independent price source.
Exercise 28.7 ★★★
Coding. With receive_times and inversions on the sample, reproduce the out-of-order percentages at jitters of 100 and 1 000 microseconds. Then halve the message rate (keep every second message) and explain the change.
Solution
Solution of Exercise 28.7.
About 1% at 100 microseconds and 11% at 1 millisecond. With every second message dropped the gaps between consecutive messages double, so a given jitter overtakes fewer neighbours and the percentages fall to roughly half at small jitters. Reordering depends on jitter relative to the inter-message time: it is worst in bursts, which is when order matters most.
Exercise 28.8 ★★★
Find the flaw. A researcher downloads one-minute bars for 3 000 tickers over fifteen years, keyed by ticker, computes each stock’s daily low as the minimum of the bars’ lows, and finds that buying stocks that fell more than 50% intraday and recovered by the close is hugely profitable. Name three data defects, from this chapter, that could produce the result.
Solution
Solution of Exercise 28.8.
(i) Bad ticks: decimal errors and cancelled trades that the vendor’s bars still contain produce “falls” of 50% or more that recover by the next bar; no one could have bought there. (ii) Conditions: late and out-of-sequence reports, or prints from another day’s price level, entering the low. (iii) Symbology: bars keyed by ticker join two companies when a ticker is reused, and unadjusted splits look like a 50% fall at the open. One may add survivorship in the ticker list, and the fact that a bar’s low says nothing about the size available at that price.
28.9 Problem: The Bad Tick
Problem 28.1
Weekend problem — fifteen prints and a benchmark
A desk must report the volume-weighted average price of a stock from 09:30:00 to 09:32:00 as the benchmark of a client order. The tape (sequence, time, price, shares, condition; the conditions are this book’s labels):
| 1 | 09:30:00 | 100.02 | 12 000 | open | 9 | 09:30:52 | 100.09 | 300 | cancelled |
| 2 | 09:30:04 | 100.05 | 300 | 10 | 09:31:03 | 100.11 | 600 | ||
| 3 | 09:30:09 | 100.04 | 50 | odd lot | 11 | 09:31:10 | 99.80 | 5 000 | late |
| 4 | 09:30:15 | 100.07 | 500 | 12 | 09:31:18 | 100.12 | 200 | ||
| 5 | 09:30:21 | 100.06 | 200 | 13 | 09:31:25 | 100.13 | 700 | ||
| 6 | 09:30:30 | 10.01 | 400 | 14 | 09:31:40 | 100.12 | 100 | ||
| 7 | 09:30:31 | 100.08 | 400 | 15 | 09:31:55 | 100.15 | 900 | ||
| 8 | 09:30:40 | 100.10 | 1 000 |
Part I — The naive number.
- Give total shares and the volume-weighted average price of all fifteen prints.
- Give the period’s low from all prints.
- By how many basis points does the naive average differ from 100.05?
- What would a volatility estimate from these prints do?
Part II — Print by print.
- Print 6: what is it, how would a filter catch it, and how would you confirm?
- Print 9: why does an outlier filter not catch it, and what does?
- Print 11 was executed at 09:24 and reported at 09:31. Should it be in the benchmark of an order that started at 09:30?
- Print 3 is an odd lot. In or out?
- Print 1 is the opening auction. The client’s order was entered at 09:30:02. In or out?
Part III — The corrected numbers.
- Excluding the erroneous, the cancelled and the late prints, give the number excluded, the shares, and the average price.
- Excluding the auction as well, give the average price of continuous trading.
- Give the corrected low and high.
- The client’s order bought 2 000 shares at an average of 100.09. Give its slippage in basis points against each of the two corrected benchmarks.
- Which benchmark is right?
Part IV — Judgement.
- The exchange later breaks print 6. How does that reach your database, and what must your loader do?
- Why should the cleaning rule be fixed before looking at the order’s result?
- A vendor’s “cleaned” bars already exclude some conditions. What must you obtain from the vendor?
- Where in the four-stage pipeline does each of the defects of Part II get handled?
- State the named result: the number of prints to exclude and the corrected volume-weighted average price.
- In one sentence: what is a trade report?
Solution
Solution of Problem 28.1.
1. 22 650 shares at 98.40. 2. 10.01. 3. basis points. 4. Explode: two consecutive returns of and dominate any estimate for months. 5. A decimal or keying error: a price ten times too small between two normal prints a second apart. A filter on distance from a rolling median or from the prevailing quote catches it; confirm against the quote at 09:30:30 (no bid near 10 existed) and against the exchange’s later cancellation. 6. Its price is normal. Only its condition, or a separate cancellation message referring to its sequence number, removes it: step 1 of Method 28.4. 7. No. It is volume of the day, and belongs in a daily total, but it was not executed during the order’s life and its price reflects 09:24; a benchmark for 09:30 to 09:32 must use the execution time, not the report time. 8. In: it is a real trade in the interval at a market price. (Whether odd lots update the consolidated last price is a rule of the tape; for a volume-weighted average they count.) 9. Out, for this client: the auction closed before the order existed and could not have been joined. It would be in for an order entered before the open. 10. Three excluded (prints 6, 9 and 11); 16 950 shares; 100.045. 11. 4 950 shares at 100.106. 12. Low 100.02 (the auction) or 100.04 in continuous trading; high 100.15. 13. Against 100.045: basis points (it paid more). Against 100.106: (it paid less). 14. The second. A benchmark should contain the trades the order could have taken part in; the auction is 71% of the first figure’s volume and was over before the order arrived. 15. As a cancel or break message later in the day, or in the next day’s corrections file, referring to the original trade by its sequence or match number. The loader must apply it to the stored trade, keep both records, and recompute anything derived. 16. Because there is always a set of exclusions that makes a given execution look good. A rule chosen after seeing the result is not a benchmark. 17. The exact list of conditions excluded, the treatment of corrections and late reports, the timestamp used for bar assignment, and whether odd lots and off-exchange trades are in. 18. Capture: nothing, everything is kept. Normalise: conditions mapped to explicit flags, corrections linked to originals. Book building: not involved for trades, but the prevailing quote it produces is what detects print 6. Tables: the purpose-specific inclusion rule and the outlier flag. 19. Three prints excluded; 100.045 (100.106 for continuous trading only). 20. A message saying that someone reported a trade, with a price, a size, a time and a set of qualifications that decide what it means.
28.10 Interview questions
Interview question 28.1 ★ developer, researcher, trader
What is the difference between level 1, level 2 and level 3 market data?
Solution
Solution of Interview question 28.1.
Level 1: best bid and offer with sizes, and last trade. Level 2: total size at each price, by level. Level 3: every order and every event on it. Each is computable from the next; only level 3 gives queue position, order lifetimes and the identity of what was executed.
What the interviewer is looking for: the one-way relation and what only level 3 provides.
Interview question 28.2 ★ developer
How do you build an order book from an order-by-order feed? What data structures?
Solution
Solution of Interview question 28.2.
A hash map from order reference to the order (side, price, remaining size) and, per instrument and side, a sorted structure from price to aggregate size (and, for queue position, the orders at that price in time order). Add: insert in both. Execute or cancel: look up by reference, reduce both, remove when empty. Delete: the same with the full size. Best prices are the ends of the sorted structure. Production books replace trees by arrays indexed by tick around the touch, and pre-allocate.
What the interviewer is looking for: lookup by reference, because executions carry no price.
Interview question 28.3 ★★ developer, researcher
Your feed handler detects a gap in sequence numbers. What happens next?
Solution
Solution of Interview question 28.3.
First check the other line. If both missed it: mark the affected books stale, tell the strategies, buffer what follows, request retransmission or take the next snapshot, rebuild, replay the buffer, clear the flag. Log the gap with times. Trading on a stale book is forbidden; pulling quotes is the default.
What the interviewer is looking for: the stale flag reaching the strategy.
Interview question 28.4 ★★ researcher, mle
Which timestamp do you use in a backtest, and what goes wrong if you choose the other?
Solution
Solution of Interview question 28.4.
The receive timestamp, at the point where the strategy would have seen the data, plus the strategy’s own latency to act. Exchange timestamps give the backtest information before it could have arrived: it trades on quotes that were already gone, and cross-venue signals look profitable because one venue’s clock runs ahead of the other’s. Exchange time remains right for ordering events within one venue.
What the interviewer is looking for: look-ahead through the clock.
Interview question 28.5 ★★ researcher, mle
You are given a trades file. What do you check before computing anything?
Solution
Solution of Interview question 28.5.
Identifiers: what keys the file, and are they stable through time. Time: which clock, which zone, execution or report time. Conditions: which are present, which were removed. Corrections: applied or not. Units: currency, price scale, adjusted or raw for splits. Duplicates and gaps: per day counts against an independent total. Outliers against quotes. Coverage: which venues, on- and off-exchange. Then ten minutes looking at one day by eye.
What the interviewer is looking for: conditions, corrections, identifiers, clock.
Interview question 28.6 ★★★ developer
Design the storage of a firm’s captured market data so that any research result can be reproduced five years later.
Solution
Solution of Interview question 28.6.
Store the raw capture, unmodified, with hardware receive timestamps, compressed and checksummed, by venue, line and day: it is the only ground truth. Version the decoder and the reference data; record with every derived table the capture files, code version and parameters that produced it. Derived data (books, bars, tables) are caches that can be deleted and rebuilt. Reference data is kept as of each date, never overwritten. Test reproducibility by rebuilding a random old day every week and comparing hashes.
What the interviewer is looking for: raw capture as ground truth; point-in-time reference data; provenance on every table.