---
title: "Crypto Data and Infrastructure"
book: "Markets III: Commodities, Energy and Crypto"
subject: quant
language: en
chapter: 26
exercises: 8
source: https://one-course.com/books/quant/3/en/chapter/26-crypto-data-and-infrastructure
---

# Chapter 26 — Crypto Data and Infrastructure

A purchased history of a crypto order book shows, on a quiet afternoon, the best bid above the best offer for two seconds. Nothing happened in the market. One update of the venue’s websocket feed was lost on the way to the recorder, the recorder applied the next ones anyway, and every system downstream, a backtest, a research notebook, a model trained on the data, treated the crossed book as real. Crypto data looks easy to get: the venues publish their books for free, the chains are public. It is hard to get right. This chapter covers the two sources of a crypto firm’s data, the chains and the venues: running nodes and reading what contracts emit, rebuilding venues’ books from snapshots and deltas and checking them, understanding what each timestamp measures, and knowing where the matching engines physically are.

## 26.1 Running nodes

**Definition 26.1 (Full node, archive node, RPC endpoint).**

A *full node* is a program that downloads and verifies every block of a chain and keeps its current state, and the recent history needed to serve it. An *archive node* is a full node that also keeps every historical state, so that it can answer questions about any past block. An *RPC endpoint* is the interface through which a program queries a node and submits transactions to it.

A trading firm reads the chain through nodes: its own, or a provider’s. Its own nodes cost hardware and operations but give it the chain’s data as soon as its node sees it, no [rate limits](https://one-course.com/books/quant/3/en/chapter/15-centralised-exchanges#def-m3-centralised-exchanges-ratelimit) and no dependence on a third party; a provider’s endpoints are convenient and are shared with everyone else, throttled, and one more party between the firm and the chain. For latency-sensitive work ([Chapter 22](https://one-course.com/books/quant/3/en/chapter/22-maximal-extractable-value#ch-m3-maximal-extractable-value)) the firm runs its own nodes, close to the other nodes it needs to hear from first.

**As of September 2026 — Ethereum nodes.**

Ethereum’s developer documentation states that [full nodes](#def-m3-crypto-data-and-infrastructure-node) verify every block but keep only relatively recent state (typically the last 128 blocks), regenerating older data on demand; that [archive nodes](#def-m3-crypto-data-and-infrastructure-node) never delete downloaded data and represent terabytes of storage; and that one client, Erigon, can perform a full archive sync using around 2 TB of disk in under three days. It lists five execution clients (Geth, Nethermind, Besu, Erigon, Reth) written in four languages.

## 26.2 On-chain data: logs and indexers

**Definition 26.2 (Event log, indexer).**

An *event log* is a record a [smart contract](https://one-course.com/books/quant/3/en/chapter/14-blockchains-for-traders#def-m3-blockchains-for-traders-contract) emits during a transaction (a swap, a transfer, a liquidation), stored in the transaction’s receipt with indexed fields that can be searched. An *indexer* is a system that reads every block from a node, decodes the logs and state changes of the contracts it cares about, and stores them in a database that can be queried by time, address or event type.

The chain is a database with an unhelpful interface: it answers “what is the state now” and “what happened in block $n$”, but not “every swap in this pool last month”. The firm’s [indexer](#def-m3-crypto-data-and-infrastructure-log) turns the one into the other. Its traps are those of any data pipeline with a twist: blocks can be reorganised away ([Chapter 14](https://one-course.com/books/quant/3/en/chapter/14-blockchains-for-traders#ch-m3-blockchains-for-traders)), so recent data must carry its block hash and be revised when a reorganisation happens; contract upgrades change what logs mean; and a [token](https://one-course.com/books/quant/3/en/chapter/14-blockchains-for-traders#def-m3-blockchains-for-traders-key)’s decimals, not its log, say what a raw amount is worth.

![A crypto firm’s two data pipelines: blocks from its own nodes through an indexer, and venues’ feeds through a book builder that checks sequence and checksums, both into one normalised store that keeps the source’s timestamp and the time of receipt. Schematic.](https://one-course.com/images/onecourse/chapters/quant-3/m3-crypto-data-and-infrastructure/fig-77d293ae9aa3.svg)

***Figure 26.1.** A crypto firm’s two data pipelines: blocks from its own nodes through an [indexer](#def-m3-crypto-data-and-infrastructure-log), and venues’ feeds through a book builder that checks sequence and checksums, both into one normalised store that keeps the source’s timestamp and the time of receipt. Schematic.*

## 26.3 Venue market data: snapshots, deltas and checksums

**Definition 26.3 (Snapshot-and-delta feed, book checksum).**

A *snapshot-and-delta feed* publishes a venue’s market-by-price book (One Quant Book 1, chapter 28) as a snapshot of all levels, with an update identifier, followed by a stream of changes to individual levels numbered with the same identifiers. A *book checksum* is a value the venue computes over the top levels of its book after each change and publishes with it, so that a client can verify that its copy is identical.

**As of September 2026 — Rebuilding a book: one venue’s procedure.**

Binance’s documentation for a local order book instructs the client to buffer the depth stream, fetch a snapshot, discard buffered events whose last update id $u$ is at most the snapshot’s `lastUpdateId`, and then, for each event: ignore it if $u$ is below the book’s id; if its first update id $U$ exceeds the book’s id plus one, discard the book and restart, since events were missed; otherwise set each level to its new quantity, removing it if the quantity is zero, and set the book’s id to $u$. A single websocket connection is valid for 24 hours, and the server pings every 20 seconds.

**As of September 2026 — A book checksum.**

Kraken’s documentation for its version-2 websocket book computes the checksum over the top 10 price levels whatever the subscription depth: for asks from low to high and then bids from high to low, each price and quantity is written without its decimal point and leading zeros, the strings are concatenated, and the CRC32 of the result, as an unsigned 32-bit integer, is compared with the message’s checksum field. Verification is optional. OKX went the other way: on 23 June 2026 it deprecated the checksum of its order-book channels, still sent but fixed at zero, and told clients to check continuity with the messages’ sequence identifiers instead.

![Rebuilding a book from a snapshot-and-delta feed, following the procedure of : buffer, snapshot, discard what the snapshot already contains, apply in order, and go back to the snapshot on any gap. The book is flagged stale from the gap until the new snapshot is applied. Schematic.](https://one-course.com/images/onecourse/chapters/quant-3/m3-crypto-data-and-infrastructure/fig-ad623748d515.svg)

***Figure 26.2.** Rebuilding a book from a [snapshot-and-delta feed](#def-m3-crypto-data-and-infrastructure-feed), following the procedure of [Box 26.2](#dat-m3-crypto-data-and-infrastructure-binance): buffer, snapshot, discard what the snapshot already contains, apply in order, and go back to the snapshot on any gap. The book is flagged stale from the gap until the new snapshot is applied. Schematic.*

Update identifiers catch losses on a venue that publishes them; checksums catch what identifiers cannot, such as a client’s own bug in applying a change, and are the only protection on a feed without identifiers.

**Proposition 26.4 (What a lost message costs).**

Let a book of $K$ active levels change by absolute-quantity updates, each touching one level chosen uniformly, and let each message be lost with probability $p$, independently. A client that applies what it receives without checking is wrong after a loss until the lost level is next updated: a geometric number of messages with mean about $K$, so that it is silently wrong a fraction about $pK$ of the time for small $p$. A client that detects the gap and needs $R$ messages to resynchronise is never silently wrong and is knowingly stale a fraction about $pR$ of the time.

**Proof.** Each later message updates the lost level with probability $1/K$, so the wait is geometric with mean $K$. Losses arrive at rate $p$ per message and, for small $p$, rarely overlap, so the fractions are the rate times the durations. ∎

![A locally built book of 20 active levels under message loss (). Without checks it is wrong, without knowing it, about 20p of the time; with gap detection it is never silently wrong but spends about 50p of the time resynchronising when a snapshot takes 50 messages to arrive. Simulated feed of 100 000 messages. Data: the chapter’s tutorial.](https://one-course.com/images/onecourse/chapters/quant-3/m3-crypto-data-and-infrastructure/fig-586c549cce42.svg)

***Figure 26.3.** A locally built book of 20 active levels under message loss ([Proposition 26.4](#prop-m3-crypto-data-and-infrastructure-loss)). Without checks it is wrong, without knowing it, about $20p$ of the time; with gap detection it is never silently wrong but spends about $50p$ of the time resynchronising when a snapshot takes 50 messages to arrive. Simulated feed of 100 000 messages. Data: the chapter’s tutorial.*

The careful client is stale more of the time than the careless one is wrong, and that is the point: a stale book that knows it is stale can pull its quotes and mark its data as missing; a wrong book trades and records falsehoods. Purchased histories must be checked the same way: a vendor’s dataset of incremental book updates is only as good as its recorder’s handling of gaps, and a crossed or locked book in the data is the first sign.

## 26.4 Timestamps and what they mean

Every crypto data record should carry two times: the venue’s (exchange timestamp) and the firm’s (receive timestamp), as in One Quant Book 1, chapter 28.

**As of September 2026 — Timestamps in a vendor’s data.**

Tardis.dev’s downloadable datasets give each record a `timestamp`, provided by the exchange in microseconds since the epoch (with the local time as a fallback when the exchange provides none), and a `local_timestamp`, the message’s arrival time; its incremental book data are collected from the exchanges’ real-time websocket feeds, and files may begin with updates that precede the first snapshot, which must be skipped.

The exchange’s timestamp says when the venue says it happened, at the venue’s clock resolution and sometimes when the message was sent rather than when the book changed. The receive timestamp says when the firm learned it, including the network path. Their difference is the feed’s latency, which varies with the firm’s location and the venue’s load and jumps exactly when the market is busiest. Research must use the receive timestamp to decide what a strategy could have known, and the exchange timestamp to align venues; a backtest that uses exchange timestamps for decisions assumes the firm sat inside the matching engine. On chains there is a third time, the block’s, and a fourth, when the firm’s node saw the transaction in the [mempool](https://one-course.com/books/quant/3/en/chapter/14-blockchains-for-traders#def-m3-blockchains-for-traders-mempool).

## 26.5 Where the matching engines sit

**Definition 26.5 (Cloud region).**

A *cloud region* is a geographic area in which a cloud provider operates a cluster of data centres, within which network latency between machines is low.

Unlike the exchanges of One Quant Book 1, some crypto venues run their systems in public clouds rather than in data centres of their own (one venue’s documentation lists an API endpoint named for a cloud provider, [Box 25.3](https://one-course.com/books/quant/3/en/chapter/25-getting-access-crypto-venues#dat-m3-getting-access-crypto-venues-endpoints)), and a firm that wants to be close to such a venue rents machines in the same region and, where the provider allows, the same availability zone. Latency to a venue becomes a question of which region it runs in, which is sometimes published and sometimes inferred by measuring round-trip times from each region; and a firm trading several venues in different regions faces the same geography as the multi-venue firms of the equity and futures markets, with distances between regions instead of between buildings.

**As of September 2026 — A venue’s region.**

Glassnode’s research, reported by CoinDesk on 30 March 2026, found Hyperliquid’s 24 [validators](https://one-course.com/books/quant/3/en/chapter/14-blockchains-for-traders#def-m3-blockchains-for-traders-pow) clustered in Amazon Web Services’ Tokyo region, across several availability zones, so that orders from Tokyo reached them in 2 to 3 milliseconds and orders from Europe in more than 200.

## 26.6 Tutorial: rebuilding a book and losing a message

**Goal.** Rebuild an order book from a snapshot and a stream of update-id-numbered deltas, detect gaps and resynchronise, verify the book against a CRC32 checksum, and measure what message loss does. **End state:** [Figure 26.3](#fig-m3-crypto-data-and-infrastructure-loss) and the numbers of the weekend problem.

1. **The update rule.** Ignore old events, detect gaps, apply absolute quantities. `Status apply (std::int64_t first, std::int64_t last, const Levels& bids, const Levels& asks) { if (!synced_) return Status::gap; if (last < update_id_ + 1 ) return Status::ignored; if (first > update_id_ + 1 ) { synced_ = false ; return Status::gap; } for (auto [p, q] : bids) q == 0 ? void (bids_.erase(p)) : void (bids_[p] = q); for (auto [p, q] : asks) q == 0 ? void (asks_.erase(p)) : void (asks_[p] = q); update_id_ = last; return Status::applied; }` **Listing 26.1.** The C++20 update rule: ignore, detect a gap, or apply. code/firm/wsbook/cpp/firm_wsbook.hpp
2. **The checksum.** Ten best asks, then ten best bids, as integer strings; CRC32. `pub fn checksum (&self ) -> u32 { let mut s = String ::new(); for (p, q) in self .asks.iter().take(10 ) { s.push_str(&format!(" {p}{q} " )); } for (p, q) in self .bids.iter().rev().take(10 ) { s.push_str(&format!(" {p}{q} " )); } crc32(s.as_bytes()) }` **Listing 26.2.** The Rust checksum over the top of the book. code/firm/wsbook/rust/src/lib.rs
3. **Run** `simulate(p)` for loss rates from 0.01% to 1% and `fig_data.py` .

**What to change next.** Add a checksum policy (detect a divergence at the next message, then resync) and compare it with gap detection; let losses come in bursts rather than independently; replay a vendor’s file and count crossed books.

## 26.7 Build: the websocket book builder

**Purpose.** Every price the miniature firm uses on a crypto venue comes from a book it built itself from the venue’s feed; the builder must never present a wrong book as right, and must say when it is stale. It feeds the arbitrage scanner of [Chapter 16](https://one-course.com/books/quant/3/en/chapter/16-spot-markets#ch-m3-spot-markets).

**Interface.** `Book` with `snapshot(update_id, bids, asks)`, `apply(first, last, bids, asks)` returning ignored, applied or gap, `top(n)` and `checksum()`; `crc32`. C++20 header `firm_wsbook.hpp`, Rust crate `firm_wsbook`, Python reference.

**Rules.** Integer ticks and lots; a gap leaves the book unsynced until a snapshot; the checksum computed exactly as the venue does; every update keeps both timestamps.

**Acceptance tests.** `code/firm/wsbook/tests/`, `cpp/firm_wsbook_test.cpp` and the Rust crate’s tests: ignoring, applying and gap detection; staying unsynced; the checksum against zlib’s CRC32 and the standard check value.

**Stretch.** Venue adapters (decimal strings, depth limits of snapshots); buffered resynchronisation without dropping events; burst-loss testing; a lock-free publication of the book to strategies.

Sources and further reading

- Ethereum.org developer documentation, “Nodes and clients” (repository ethereum/ethereum-org-website). Accessed September 2026.
- Binance, spot WebSocket streams documentation (repository binance/binance-spot-api-docs); Kraken, WebSocket v2 book checksum guide. Accessed September 2026.
- Tardis.dev, downloadable CSV files: data types. Accessed September 2026.
- OKX, API v5 change log, entry of 23 June 2026; CoinDesk, “Hyperliquid traders in Tokyo get 200-millisecond edge, Glassnode research shows”, 30 March 2026.

## 26.8 Exercises

**Exercise 26.1 ★.**

A book’s id is 100. What does the client do with events $(U, u) = (95, 99)$, $(101, 103)$ and then $(105, 106)$?

**Solution of Exercise 26.1.**

Ignore $(95, 99)$ (already in the book); apply $(101, 103)$, the book’s id becoming 103; $(105, 106)$ starts after 104, so update 104 was missed: discard the book and resynchronise from a new snapshot.

**Exercise 26.2 ★.**

When does a firm need an [archive node](#def-m3-crypto-data-and-infrastructure-node) rather than a [full node](#def-m3-crypto-data-and-infrastructure-node)?

**Solution of Exercise 26.2.**

When it must query historical state at arbitrary past blocks (balances, positions, contract storage then) quickly and reliably, for research, reconciliation or tracing; a [full node](#def-m3-crypto-data-and-infrastructure-node) can regenerate some of it but slowly.

**Exercise 26.3 ★.**

Write the checksum string for asks (1001, 4), (1002, 6) and bids (999, 5), (998, 7).

**Solution of Exercise 26.3.**

“10014” “10026” “9995” “9987” concatenated, asks first from low to high, then bids from high to low; its CRC32 is 3 348 617 501.

**Exercise 26.4 ★★.**

A feed loses one message in 1 000 and the book has 20 active levels. For what share of the time is an unchecked book wrong? How long is a typical error?

**Solution of Exercise 26.4.**

About $pK = 0.1\% \times 20 = 2\%$ of the time (1.9% in the simulation); a typical error lasts about 20 messages (21 simulated), until the lost level is next updated.

**Exercise 26.5 ★★.**

Why should research use receive timestamps to decide what a strategy knew?

**Solution of Exercise 26.5.**

A strategy can act only on what it has received; exchange timestamps say when the venue says something happened, earlier than the firm could know it, so decisions timed by them look into the future by the feed’s latency.

**Exercise 26.6 ★★.**

What must an [indexer](#def-m3-crypto-data-and-infrastructure-log) do when the chain reorganises two blocks?

**Solution of Exercise 26.6.**

Detect it from the block hashes, delete or mark as orphaned the records from the two abandoned blocks, re-read the new blocks at those heights, and propagate the corrections to anything computed from them; recent records should carry their block hash and a [finality](https://one-course.com/books/quant/3/en/chapter/14-blockchains-for-traders#def-m3-blockchains-for-traders-finality) flag.

**Exercise 26.7 ★★★.**

*Coding.* With `simulate`, find the resynchronisation time at which the careful client is stale as often as the careless one is wrong, at a loss rate of 0.1%.

**Solution of Exercise 26.7.**

About 20 messages: the stale share is $pR$ and the silent-error share $pK$, equal when $R = K = 20$ (both 1.92% in the simulation).

**Exercise 26.8 ★★★.**

*Find the flaw.* “Our data vendor records every exchange’s feed, so its order books are correct.”

**Solution of Exercise 26.8.**

Recording is not rebuilding: the vendor’s recorder can lose messages, and its dataset reflects how it handled gaps; check for crossed books, sequence gaps and resynchronisations, and compare with a second source.

## 26.9 Problem: The Missing Delta

**Problem 26.1.**

Weekend problem — how wrong is a book that lost a message?

A firm builds books from a venue’s feed of 20 active levels near the top. It loses one message in 1 000. A resynchronisation from a snapshot takes about 50 messages.

**Part I — The careless client.**

1. For what share of messages is its book wrong?
2. How long does a typical error last, in messages?
3. Why is that close to the number of levels?
4. What can a wrong book do to a quoting strategy?
5. Why might a lost deletion be worse than a lost change of quantity?

**Part II — The careful client.**

6. For what share of messages is its book flagged stale?
7. Why is that larger than the careless client’s error share?
8. What does the strategy do while the book is stale?
9. How would buffering events during resynchronisation shorten the stale time?
10. What would checksums add on a venue that publishes update identifiers?

**Part III — Data for research.**

11. How would you detect lost messages in a purchased dataset?
12. What does a crossed book in the data indicate?
13. Which timestamp should a backtest use for decisions?
14. How do you align two venues’ books in time?
15. What should the dataset record about resynchronisations?

**Part IV — Judgement.**

16. Is a stale book better than a wrong one?
17. When is a provider’s node enough, and when must a firm run its own?
18. What makes crypto data harder than equity data, and what makes it easier?
19. State the *named result* : the share of time the book is wrong without checks and stale with resynchronisation, and the expected duration of an error.
20. In one sentence: what is a local order book?

**Solution of Problem 26.1.**

**1.** About 1.9% of messages. **2.** About 21 messages. **3.** Each message updates the lost level with probability $1/20$, so the wait is geometric with mean about 20. **4.** Quote around a wrong mid, think a level exists when it does not, or miss an arbitrage that is not there; in the worst case trade against a phantom crossed book. **5.** A lost deletion leaves a phantom level that stays until that exact price is updated again, which, if the price moves away, may be never. **6.** About 4.75%. **7.** Resynchronising (50 messages) takes longer than a typical silent error (20); the careful client trades knowledge of being stale for being stale longer. **8.** Pulls or widens its quotes and marks its data as missing. **9.** By applying buffered events on top of the new snapshot rather than waiting for fresh ones, the stale window shrinks to the snapshot’s latency. **10.** Protection against the client’s own bugs and against identifier mistakes, and verification of the book it computes. **11.** Check the update identifiers for gaps, look for crossed or locked books, compare the top of book with trades, and compare with a second source. **12.** A lost update, usually a missed deletion or a late snapshot, since a real venue’s book cannot cross. **13.** The receive timestamp. **14.** By their exchange timestamps corrected for each feed’s measured latency, or by receive timestamps at a single location. **15.** When each resynchronisation began and ended, so that users can drop or flag those periods. **16.** Yes: a stale book knows it cannot be trusted; a wrong one trades. **17.** A provider suffices for occasional queries and research; latency-sensitive trading, high query volumes and independence need own nodes. **18.** Harder: many venues, each with its own feed conventions, public clouds, no consolidated tape, chains that reorganise; easier: free, public, full-depth data from every venue and every chain. **19.** *Named result:* at a loss rate of 0.1% and 20 levels, 1.9% of the time silently wrong without checks (errors lasting about 21 messages), against 4.75% knowingly stale with resynchronisation taking 50 messages. **20.** The firm’s own copy of a venue’s book, as good as its handling of the messages it did not receive.

## 26.10 Interview questions

**Interview question 26.1 ★ developer.**

Describe how to maintain a local order book from a snapshot and a delta stream.

**Solution of Interview question 26.1.**

Buffer the stream, fetch a snapshot, drop buffered events already contained in it, apply the rest in order checking that each event starts right after the book’s id, set levels to absolute quantities (zero removes), and resynchronise on any gap; verify with checksums where the venue publishes them.

*What the interviewer is looking for: the sequence rule and resync.*

**Interview question 26.2 ★ researcher.**

What is the difference between an exchange timestamp and a receive timestamp, and when do you use each?

**Solution of Interview question 26.2.**

The exchange timestamp is the venue’s time of the event (or message); the receive timestamp is when the firm got it. Use receive times for what a strategy knew and exchange times to align venues and measure latency.

*What the interviewer is looking for: knowledge versus alignment.*

**Interview question 26.3 ★★ developer.**

How would you design an [indexer](#def-m3-crypto-data-and-infrastructure-log) for every swap on a [decentralised exchange](https://one-course.com/books/quant/3/en/chapter/20-automated-market-makers#def-m3-automated-market-makers-amm), robust to reorganisations?

**Solution of Interview question 26.3.**

Read blocks from own nodes; decode swap logs by pool address and signature; store with block number, hash, log index and [finality](https://one-course.com/books/quant/3/en/chapter/14-blockchains-for-traders#def-m3-blockchains-for-traders-finality); handle reorganisations by comparing parent hashes and rewriting orphaned ranges; backfill from an [archive node](#def-m3-crypto-data-and-infrastructure-node); version decoders with contract upgrades.

*What the interviewer is looking for: reorg handling and decoding discipline.*

**Interview question 26.4 ★★ developer.**

Your [book checksum](#def-m3-crypto-data-and-infrastructure-feed) fails once an hour on one venue. How do you find out why?

**Solution of Interview question 26.4.**

Log the book and the message at each failure; check the formatting rule (decimal places, trailing zeros, level count), ordering of levels, the handling of zero quantities and of levels beyond the snapshot’s depth, and message loss or reordering around the failures; reproduce offline from the recorded stream.

*What the interviewer is looking for: formatting first, then sequence.*

**Interview question 26.5 ★★ researcher.**

How would you measure the latency of your feed from a venue running in a public cloud?

**Solution of Interview question 26.5.**

Compare exchange and receive timestamps on quiet messages, measure round trips of cheap requests from machines in several regions, and watch the distribution rather than the mean, especially in busy periods; synchronise the firm’s clocks first.

*What the interviewer is looking for: clock discipline and distributions.*

**Interview question 26.6 ★★★ developer.**

Design the market-data layer for a firm trading on ten crypto venues and three chains.

**Solution of Interview question 26.6.**

Per-venue adapters that normalise messages and keep both timestamps; book builders with sequence and checksum checks; per-chain node clusters and [indexers](#def-m3-crypto-data-and-infrastructure-log); a normalised bus to strategies with staleness flags; recording of raw messages for replay; monitoring of gaps, latency and resyncs; and deployment in the regions nearest each venue.

*What the interviewer is looking for: normalise, verify, flag and record.*
