---
title: "Data and Feature Stores"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 24
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/24-data-and-feature-stores
---

# Chapter 24 — Data and Feature Stores

A one-second order-book forecast has a rank information coefficient of 0.44 in its backtest and 0.36 live. Nothing is wrong with the model. In research its features were computed on one-second bars and joined to each decision at the end of its second; in production they are computed from every message up to the decision. The research features had seen up to a second of the future, and the model had learned to use it. This chapter is about making that impossible: declaring each feature once and computing it the same way for training and for trading, storing it with the time it became known, knowing how fresh each value is, and testing continuously that the two computations agree. Its measurements use the book’s synthetic order-book market and four planted skews; the parity test catches all four, and only one of them costs the model anything, which is a reason to catch them all rather than guess which matter.

## 24.1 The same feature twice: offline and online

**Definition 24.1 (Feature store, offline store, online store, feature definition).**

A *feature store* is the system that computes, stores and serves a firm’s model inputs from one set of declarations. Its *offline store* holds feature values for history, keyed by entity and time, for training and backtests; its *online store* holds the latest values, updated as data arrive, for models in production. A *feature definition* declares a feature as data (its inputs, its aggregation, its window, its clock and its version) so that both stores compute it from the same declaration (Sculley and co-authors, 2015, on why this matters; Hermann and Del Balso, 2017, for an early production design).

The chapter declares twelve features of the order book ([Table 24.1](#tab-ml-fs-skew)): the spread, the queue and depth imbalances and the weighted mid (the last values after each message, from Book 7’s streaming `firm.lobfeat` engine), order-flow imbalance over five and thirty seconds, traded and signed volume over thirty seconds, trade and message counts over ten, and the mid’s change over five and thirty. Each is a `FeatureDef`: a name, the per-event input, an aggregation (last, sum, count, change) and a window. The [offline store](#def-ml-data-and-feature-stores-store) computes all twelve for a whole history at a list of decision times at once, from cumulative sums and as-of lookups ([Listing 24.1](#lst-ml-fs-offline)); the online engine consumes the messages one by one and keeps running windows ([Listing 24.2](#lst-ml-fs-online)). On eight thirty-minute sessions of `firm.tape`, 313 681 messages and 7 714 decisions (one at each trade), the two agree exactly at every decision: that is the store’s contract.

## 24.2 Point-in-time correctness in the store

Every value in the [offline store](#def-ml-data-and-feature-stores-store) carries the time it became known, and a query for a decision at time $t$ sees only values known at $t$: the as-of join of Book 7 (chapter 3), applied to features. For event-driven features, knowledge time is the time of the last message the feature has read; for features from slower data (fundamentals, alternative data, chapter 15) it is the delivery time, and the store is bitemporal. The two bugs this rules out are the ones planted here: a research pipeline that reads one message past the decision (an off-by-one in a join), and one that computes features on a coarser clock and attaches to each decision the value at the end of its bar.

## 24.3 Freshness and materialisation

**Definition 24.2 (Materialisation, feature freshness).**

*Materialisation* is the job that computes feature values from their definitions and writes them to a store, in batch for the [offline store](#def-ml-data-and-feature-stores-store) or continuously for the online one. *Feature freshness* is the age of a served value: the time since the newest input it reflects, which a production system monitors and bounds.

Freshness is where production differs from research even when the code is shared: an input feed that arrives late makes every online value that depends on it stale. The fourth planted skew delays trade prints by 200 milliseconds relative to the book updates, as a separate trade feed might; the three trade-count and volume features then differ from research at 92 to 94% of decisions, by 7 to 17% of their standard deviation, while the book features are untouched.

## 24.4 Testing for skew

**Definition 24.3 (Training–serving skew).**

*Training–serving skew* is any difference between the feature values a model was trained on and those it receives in production for the same moment: different code, different clocks, different inputs, different arithmetic or different data timing.

The parity test ([Listing 24.3](#lst-ml-fs-parity)) recomputes production’s values offline for a sample of recent decisions and compares them feature by feature. [Table 24.1](#tab-ml-fs-skew) shows each planted skew’s footprint, and [Table 24.2](#tab-ml-fs-parity) what the test and the model see. A check of five random decisions per session, run five times on each of the eight sessions, flags every skew every time: each changes most decisions. The tolerance is a design choice: at $10^{-9}$ the test also flags production’s float32 storage (a change of at most $10^{-7}$, harmless to the model), and at $10^{-6}$ it stops flagging it and still flags the other three at 96 to 100% of decisions.

|  | RMS difference as a share of the feature’s standard deviation |
| --- | --- |
| feature | one event ahead | one-second bars | float32 | late trades |
| spread | 0.38 | 0.92 | 0 | 0 |
| imbalance | 0.24 | 0.73 | 0.00 | 0 |
| depth imbalance | 0.11 | 0.41 | 0.00 | 0 |
| weighted mid minus mid | 0.30 | 0.79 | 0.00 | 0 |
| OFI 5 s | 0.03 | 0.24 | 0 | 0 |
| OFI 30 s | 0.01 | 0.07 | 0 | 0 |
| volume 30 s | 0.01 | 0.07 | 0 | 0.07 |
| signed volume 30 s | 0.02 | 0.10 | 0 | 0.11 |
| trades 10 s | 0.02 | 0.18 | 0 | 0.17 |
| mid change 5 s | 0.09 | 0.39 | 0 | 0 |
| mid change 30 s | 0.03 | 0.13 | 0 | 0 |
| messages 10 s | 0.00 | 0.04 | 0 | 0 |

***Table 24.1.** Size of each planted skew, feature by feature, over 7 714 decisions (0.00: a difference smaller than half a hundredth; 0: none). Data: `ml_featstore.feature_skew`.*

|  | decisions | detected | rank IC | live P&L |
| --- | --- | --- | --- | --- |
| skew | mismatched | (5-sample check) | backtest | live | (ticks) |
| none | 0 | 0 | 0.411 | 0.411 | 0.089 |
| research one event ahead | 100% | 100% | 0.395 | 0.418 | 0.091 |
| research on one-second bars | 96% | 100% | 0.438 | 0.360 | 0.078 |
| production float32 storage | 92% | 100% | 0.411 | 0.411 | 0.089 |
| production trade prints 200 ms late | 100% | 100% | 0.411 | 0.411 | 0.089 |

***Table 24.2.** Parity test (tolerance $10^{-9}$) and a ridge forecast of the mid’s change over the next second, trained on four sessions with each research pipeline and tested on four others: its backtest on research features and its live result on production features, rank IC and mean P&L per decision of trading the forecast’s sign. Data: `ml_featstore.parity_table`, `model_effects`.*

![What each skew does to the model: the rank IC its research backtest reports and the one it earns live. Data: ml_featstore.model_effects.](https://one-course.com/images/onecourse/chapters/quant-12/ml-data-and-feature-stores/fig-a52518e95839.svg)

***Figure 24.1.** What each skew does to the model: the rank IC its research backtest reports and the one it earns live. Data: `ml_featstore.model_effects`.*

The chapter’s named result is in the table’s third row. Research on one-second bars inflates the backtest’s IC from 0.411 to 0.438 and, because the model learned to lean on features that had seen the future, delivers 0.360 live, below what the clean pipeline earns; the P&L per decision falls from 0.089 ticks to 0.078. The one-event look-ahead does the reverse here: the message after a trade is usually a quote update deep in the book, which adds noise rather than information, and the backtest understates the live result (0.395 against 0.418). On the five-second forecast the same skews move the IC by less than 0.01: whether a skew costs anything depends on how its error correlates with the target, which is not knowable in advance. The store’s job is to make the stores agree, not to judge which disagreements are harmless.

**Method 24.4 (Running features in production).**

1. Declare every feature once, as data, with its clock and window, and version it.
2. Compute offline and online from the same declarations; in the [offline store](#def-ml-data-and-feature-stores-store) , stamp every value with its knowledge time and join point in time.
3. Run a parity test on a sample of recent decisions every day, with tolerances set per feature, and alert on any mismatch.
4. Monitor each online feature’s freshness against a bound, and fall back or halt when inputs go stale.

## 24.5 Tutorial: same features twice

**Goal.** Declare twelve features, compute them offline and online on eight sessions, plant four skews, and measure what the parity test and the model see. **End state:** Tables [24.1](#tab-ml-fs-skew) and [24.2](#tab-ml-fs-parity), [Figure 24.1](#fig-ml-fs-ic).

1. **The [offline store](#def-ml-data-and-feature-stores-store): every decision at once.** `def offline (events, defs, times, lag_events=0 ): """Values at each decision time from the events with time <= t (plus lag_events more, to plant a look-ahead).""" t = events[" t " ] k = np.searchsorted(t, times, side=" right " ) - 1 + lag_events # index of the last event known k = np.clip(k, -1 , len (t) - 1 ) out = np.zeros((len (times), len (defs))) for j, d in enumerate (defs): x = events[d.input] if d.agg == " last " : out[:, j] = np.where(k >= 0 , x[np.maximum(k, 0 )], 0.0 ) continue if d.agg == " change " : k0 = np.searchsorted(t, times - d.window, side=" right " ) - 1 out[:, j] = x[np.maximum(k, 0 )] - x[np.maximum(k0, 0 )] continue c = np.r_[0.0 , np.cumsum(x if d.agg == " sum " else np.ones_like(x))] k0 = np.searchsorted(t, times - d.window, side=" right " ) # first event inside (t - w, t] out[:, j] = c[k + 1 ] - c[k0] return out` **Listing 24.1.** Vectorised point-in-time feature values. code/firm/featstore/firm_featstore.py
2. **The online engine: running windows.** `def values (self , now): out = np.zeros(len (self .defs)) for j, d in enumerate (self .defs): if d.agg == " last " : out[j] = self .last.get(d.input, 0.0 ) elif d.agg == " change " : h = self .hist.get(j, deque()) while len (h) > 1 and h[1 ][0 ] <= now - d.window: h.popleft() base = h[0 ][1 ] if h and h[0 ][0 ] <= now - d.window else (h[0 ][1 ] if h else 0.0 ) out[j] = self .last.get(d.input, 0.0 ) - base else : q = self .q[j] while q and q[0 ][0 ] <= now - d.window: self .s[j] -= q.popleft()[1 ] out[j] = self .s[j] return out` **Listing 24.2.** Online values at any moment. code/firm/featstore/firm_featstore.py
3. **The parity test.** `def parity (off, on, tol=1e-9 , sample=None , seed=0 ): """Share of decision times at which any feature differs by more than tol; with `sample`, the check a production job can afford: a random sample of that many times, flagged if any of them differs.""" bad = np.any(np.abs(off - on) > tol, axis=1 ) out = {" mismatch share " : float (bad.mean())} if sample: idx = np.random.default_rng(seed).choice(len (bad), min (sample, len (bad)), replace=False ) out[" flagged " ] = bool (bad[idx].any()) return out` **Listing 24.3.** Comparing the stores on a sample of decisions. code/firm/featstore/firm_featstore.py
4. **Run** `ml_featstore.parity_table()` , `parity_table(tol=1e-6)` , `model_effects()` , `model_effects(5.0)` , `feature_skew()` and `fig_featstore.py` .

**What to change next.** Plant a skew that affects only one decision in a thousand and find the sample size a daily check needs to catch it within a week.

## 24.6 Build: the feature store

**Purpose.** Model inputs computed once by definition, point in time offline, fresh online, and continuously checked against each other.

**Interface.** `FeatureDef(name, input, agg, window, version)`, `events_from_tape(tape, levels)` (on `firm.lobfeat`), `offline(events, defs, times, lag_events)`, `OnlineEngine(defs)` with `on` and `values`, `stream(events, defs, times)`, `parity(off, on, tol, sample, seed)`, `freshness(engine, now)`.

**Rules.** No feature exists outside a definition; the [offline store](#def-ml-data-and-feature-stores-store) never returns a value known after the query time; parity runs daily; stale inputs are alarms.

**Acceptance tests.** `code/firm/featstore/tests/`: offline equals online exactly on a session; each window aggregation matches a brute-force recomputation; a one-event look-ahead is flagged; freshness grows when inputs stop.

**Stretch.** A bitemporal [offline store](#def-ml-data-and-feature-stores-store) on `firm.pit` for slow features; tolerances per feature learned from the arithmetic; C++ or Rust online engines checked against the Python reference, as Book 7 did for `firm.lobfeat`.

Sources and further reading

- D. Sculley and co-authors, “Hidden technical debt in machine learning systems”, *Advances in Neural Information Processing Systems* 28, 2015.
- J. Hermann and M. Del Balso, “Meet Michelangelo: Uber’s machine learning platform”, Uber engineering blog, 2017.
- N. Polyzotis, S. Roy, S. E. Whang and M. Zinkevich, “Data lifecycle challenges in production machine learning: a survey”, *SIGMOD Record* 47(2), 2018.

## 24.7 Exercises

**Exercise 24.1 ★.**

A thirty-second volume feature is computed at a decision at 10:00:07.300. Which trades does the [online store](#def-ml-data-and-feature-stores-store) include, and which does a research pipeline on one-second bars joined to the end of the decision’s second include?

**Solution of Exercise 24.1.**

The [online store](#def-ml-data-and-feature-stores-store): trades stamped in $(09{:}59{:}37.300, 10{:}00{:}07.300]$. The bar pipeline: trades in the thirty one-second bars ending at 10:00:08, that is $(09{:}59{:}38, 10{:}00{:}08]$: it misses 0.7 s of the past and includes 0.7 s of the future.

**Exercise 24.2 ★.**

A skew changes one decision in a thousand. How many randomly sampled decisions must a check compare to catch it with probability 95%?

**Solution of Exercise 24.2.**

$1 - 0.999^n\ge0.95$ gives $n\ge\ln0.05/\ln0.999 = 2\,994$ decisions: about 3 000, far more than a five-decision check, which is why parity checks should also run on full days in batch.

**Exercise 24.3 ★.**

Why does float32 storage change the imbalance features but not the volume and count features?

**Solution of Exercise 24.3.**

The volume and count features are sums of integers, represented exactly in float32 up to $2^{24}$; the imbalances and the weighted mid are ratios with 24 bits of precision in float32 against 53 in float64, so they change in the seventh or eighth significant digit.

**Exercise 24.4 ★★.**

Why did the one-second bars hurt the one-second forecast and not the five-second one?

**Solution of Exercise 24.4.**

A bar’s end is up to a second after the decision: for a one-second forecast the leaked interval is a large part of the target window, and the research model learned to use it; for a five-second forecast it is at most a fifth of the window and the information is diluted.

**Exercise 24.5 ★★.**

Why did the one-event look-ahead make the backtest worse rather than better here? Would it always?

**Solution of Exercise 24.5.**

The message after a trade is usually a limit order added or cancelled deep in the book: its effect on the target is small and it adds noise to features the model relies on, so the research model is slightly worse than one trained on clean features. If the next message were typically the mid-changing one (a thin book, a fast market), the look-ahead would inflate the backtest instead; the sign of a leak’s effect is not knowable in advance.

**Exercise 24.6 ★★.**

*Find the flaw.* “We checked that production and research features have the same mean and standard deviation every day, so there is no [training–serving skew](#def-ml-data-and-feature-stores-skew).”

**Solution of Exercise 24.6.**

Equal moments do not mean equal values: one-second bars shift values in time without changing their distribution, and a skew that affects some decisions can leave the daily averages intact. Compare the values themselves, decision by decision.

**Exercise 24.7 ★★★.**

*Coding.* Retrain the one-second model on the values the production engine computes when it replays the training sessions, instead of the one-second-bar research features, and score it live. What does it recover, and what does it cost?

**Solution of Exercise 24.7.**

Trained on the production engine’s replay, the model’s live IC is 0.411, the clean result, against 0.360 for the model trained on bar features: everything the skew cost is recovered. It costs a replay of the full message history through the online engine to build the [training set](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets), slower than the bar pipeline, which is exactly what a [feature store](#def-ml-data-and-feature-stores-store)’s shared definition makes routine.

**Exercise 24.8 ★★★.**

Design the tolerance of the parity test for a sum over a window of float64 inputs whose online and offline computations add in different orders.

**Solution of Exercise 24.8.**

The error of a float64 sum of $n$ terms in any order is bounded by about $n\,u\sum|x_i|$ with $u = 2^{-53}$; use a tolerance proportional to $\sum|x_i|$ times a small multiple of $n\,u$, or make both computations exact (integer inputs, compensated summation, Book 4, chapter 25) and demand equality.

## 24.8 Problem: Same Features Twice

**Problem 24.1.**

Weekend problem — one definition, two computations

The chapter’s twelve features, eight sessions and four skews.

**Part I — The store.**

1. What does a [feature definition](#def-ml-data-and-feature-stores-store) contain, and why as data?
2. How do the offline and online computations differ, and how do we know they agree?
3. What is the knowledge time of an order-book feature, and of a quarterly fundamental?
4. What is freshness, and what made it fail here?

**Part II — The skews.**

5. Describe the four skews and which features each touches.
6. Which are research bugs and which production bugs?
7. What share of decisions does each change?
8. How does the parity test’s tolerance change what it flags?

**Part III — The model.**

9. What are the backtest and live ICs under each skew?
10. Why does the one-second-bar skew inflate the backtest and hurt the live result?
11. What changes at a five-second horizon?
12. What would the IC gap look like on a desk’s dashboard?

**Part IV — The verdict.**

13. State the *named result* : the IC and P&L lost to the misaligned window, and the parity test’s detection rate for each planted skew.
14. Which skew is the most dangerous, and why?
15. How would you roll out a change to a feature’s definition?
16. What would you do on the day the parity test fires?
17. How should a vendor’s features be brought into the store?
18. What does the store owe the model’s validator (chapter 21)?
19. What is the cost of a [feature store](#def-ml-data-and-feature-stores-store) for a small team, and when is it worth it?
20. In one sentence: what does a [feature store](#def-ml-data-and-feature-stores-store) guarantee?

**Solution of Problem 24.1.**

**Part I.**

1. Name, input, aggregation, window, clock and version; as data so that both stores, the documentation and the tests read the same declaration.
2. Offline: vectorised over a history from cumulative sums and as-of lookups; online: running windows updated message by message. They agree exactly at all 7 714 decisions.
3. The time of the last message it read; for a fundamental, its delivery time to the firm.
4. The age of a value’s newest input; trade prints arriving 200 ms late made the trade features stale.

**Part II.**

1. One event ahead (all book features, slightly the windows), one-second bars (every feature, most strongly the book’s last values), float32 storage (the ratio features, invisibly), late trade prints (the three trade features).
2. The first two are research bugs; the last two are production conditions.
3. 100%, 96%, 92% and 100% of decisions (at tolerance $10^{-9}$ ).
4. At $10^{-9}$ it flags the harmless float32 storage; at $10^{-6}$ it ignores it and still flags the others.

**Part III.**

1. Backtest and live: clean 0.411 and 0.411; one event ahead 0.395 and 0.418; one-second bars 0.438 and 0.360; float32 0.411 and 0.411; late trades 0.411 and 0.411.
2. Its features saw part of the target window; the model weighted them; live they carry no such information.
3. The same skews move the IC by less than 0.01.
4. A live IC well below the backtest’s from the first day, stable thereafter: the signature of a pipeline difference, not of decay.

**Part IV.**

1. *Same features twice.* Research on one-second bars inflates the one-second forecast’s backtest IC to 0.438 and delivers 0.360 live against 0.411 for the clean pipeline, with P&L per decision falling from 0.089 to 0.078 ticks; a five-decision parity check catches each of the four planted skews every time.
2. The bar misalignment: it inflates research and hurts production, and nothing in either alone reveals it.
3. As a new version computed alongside the old, backfilled offline, compared online, and switched only when models are retrained on it.
4. Stop trading the affected models or fall back, find the feature and the cause, fix, recompute, and retrain if the training data were wrong.
5. Through the [ingestion pipeline](https://one-course.com/books/quant/12/en/chapter/15-alternative-data-pipelines#def-ml-alternative-data-pipelines-ingest) of chapter 15, with first-seen times as knowledge times, then as ordinary definitions.
6. Definitions, versions, parity history and freshness records for every feature the model uses.
7. Engineering time and discipline; worth it as soon as two models share features or one model trades live.
8. That the model sees in production what it saw in training, as of the moment it decides.

## 24.9 Interview questions

**Interview question 24.1 ★ mle.**

What is [training–serving skew](#def-ml-data-and-feature-stores-skew)? Give three causes.

**Solution of Interview question 24.1.**

A difference between the features a model trained on and those it gets live: different code paths, different clocks or bars, late or missing inputs, different arithmetic, look-ahead in the research join.

*What the interviewer is looking for: the definition and three distinct causes.*

**Interview question 24.2 ★★ mle, developer.**

Design a [feature store](#def-ml-data-and-feature-stores-store) for order-book features used by a model that trades on every message.

**Solution of Interview question 24.2.**

Declarative definitions; an online engine updating running windows per message with bounded latency; an [offline store](#def-ml-data-and-feature-stores-store) replaying the same engine or an equivalent vectorised computation over history, stamped with knowledge times; point-in-time retrieval for training; daily parity and freshness monitoring; versioning.

*What the interviewer is looking for: shared definitions, point-in-time offline, parity and freshness.*

**Interview question 24.3 ★★ researcher, mle.**

Your model’s live IC is half its backtest’s. How do you find out whether the features are to blame?

**Solution of Interview question 24.3.**

Replay the live period’s messages through the research pipeline and compare features with what production logged; retrain on production’s replayed features and see whether the backtest falls to the live IC; check freshness logs.

*What the interviewer is looking for: a replay comparison before any modelling explanation.*

**Interview question 24.4 ★★ mle.**

What is a point-in-time join, and how does a [feature store](#def-ml-data-and-feature-stores-store) implement it?

**Solution of Interview question 24.4.**

For each (entity, decision time), take each feature’s latest value with knowledge time at or before the decision; implemented by as-of joins on sorted times, never by joining on calendar keys.

*What the interviewer is looking for: knowledge time and the as-of rule.*

**Interview question 24.5 ★★ mle, developer.**

How do you monitor [feature freshness](#def-ml-data-and-feature-stores-fresh), and what do you do when a feature goes stale?

**Solution of Interview question 24.5.**

Record each input’s last arrival and each feature’s newest input; alert when their age exceeds a bound set by the feature’s window and the model’s horizon; have the model refuse or fall back on stale features rather than trade on them.

*What the interviewer is looking for: measured age, bounds and a defined behaviour when stale.*

**Interview question 24.6 ★★★ mle.**

How would you make an online and an offline implementation of a windowed sum agree bit for bit?

**Solution of Interview question 24.6.**

Make the inputs integers (ticks, shares) or fixed point, so sums are exact; or use the same summation order and compensated summation in both; and test on a replay that every value matches exactly.

*What the interviewer is looking for: exact arithmetic or controlled order, and a replay test.*
