---
title: "Machine-Learning and Data Engineer"
book: "The Industry: Firms, Roles and Careers"
subject: quant
language: en
chapter: 21
exercises: 8
source: https://one-course.com/books/quant/17/en/chapter/21-machine-learning-and-data-engineer
---

# Chapter 21 — Machine-Learning and Data Engineer

In fiscal 2021 the sourced finance employers filed 136 labour condition applications for jobs titled in machine learning, data science or data engineering; in fiscal 2025 they filed 1 051, almost eight times as many. Over the same four years their filings for software engineers doubled and those for [quantitative researchers](https://one-course.com/books/quant/17/en/chapter/17-quantitative-researcher#def-in-quant-researcher-def) did not move. The occupation that most of the 2025 filings name, [data scientist](#def-in-machine-learning-and-data-engineer-roles), did not exist in the US classification of 2010; it was created in the 2018 revision. This chapter describes the machine-learning and data jobs in trading and finance, measures how fast their filings grew and with what uncertainty, and compares their pay with the software engineers’ beside them.

| **Role cards: machine learning and data** |
| --- |
|  | [data scientist](#def-in-machine-learning-and-data-engineer-roles) | machine-learning engineer | [data engineer](#def-in-machine-learning-and-data-engineer-roles) |
| works on | models from data, for a research or business question | the systems that train, serve and monitor models | the pipelines and stores that data lives in |
| sits in | research groups or business lines | a platform team or a research group | a data or platform team |
| codes in filings | 15-2051 | 15-2051, 15-1252 | 15-1252, 15-1243 |
| taught in | Book 12, ch. 1; Book 7, ch. 12 | Book 12, ch. 23–28 | Book 15, ch. 2–4; Book 12, ch. 24 |

**Definition 21.1 (Data scientist, data engineer).**

A *data scientist* is an analyst who builds statistical and machine-learning models from data to answer a question or to drive a decision, and reports what the data show. A *data engineer* is an engineer who builds and runs the pipelines, storage and interfaces through which data is acquired, cleaned, versioned and served to the people and systems that use it.

The occupational definition of the first is broad: [data scientists](#def-in-machine-learning-and-data-engineer-roles) “Develop and implement a set of techniques or analytics applications to transform raw data into meaningful information” and “Apply data mining, data modeling, natural language processing, and machine learning”. The classification has no occupation for the second: O*NET lists “Data Engineer” among the titles reported for database architects (15-1243), and employers file most [data engineers](#def-in-machine-learning-and-data-engineer-roles) as software developers. The machine-learning engineer, who builds the systems that train, serve and monitor models, is Book 12’s (chapter 28).

## 21.1 Research-embedded machine-learning roles

In a trading firm’s research group, machine learning is a method, not a department. The researcher of chapter 17 who fits a gradient-boosted model to order-book features is doing machine learning; the title on the contract may say [quantitative researcher](https://one-course.com/books/quant/17/en/chapter/17-quantitative-researcher#def-in-quant-researcher-def) or [data scientist](#def-in-machine-learning-and-data-engineer-roles), and the filings follow the title (chapter 19). What distinguishes the research-embedded role is the problem: a signal with a low signal-to-noise ratio, few independent observations, non-stationary data and a cost of being wrong that is paid in money. Book 12 is about exactly that difference (chapter 1), and its tools (validation that respects time, labels built from future prices, models that survive noise) are the role’s daily work.

A research-embedded machine-learning specialist is judged as a researcher is: on out-of-sample results, and on whether the model adds to what simpler models already do. Book 12’s habit of setting every model against a linear and regularised baseline (chapter 4) is the role’s discipline. The [feedback speed](https://one-course.com/books/quant/17/en/chapter/17-quantitative-researcher#def-in-quant-researcher-feedback) of chapter 17 applies unchanged: a model rebalanced monthly earns evidence as slowly as any monthly signal.

## 21.2 Platform machine-learning roles

A firm that runs many models needs a machine-learning platform (Book 12, chapter 28): data and feature stores, training infrastructure, experiment tracking and a model registry, compilation and serving, monitoring and retraining. The machine-learning engineer builds and runs it, and the handoff specification of Book 12 is the contract between the people who build models and the people who run them. Most of the growth in the filings (below) is at banks, and the filings do not say which of a bank’s businesses each job serves.

The platform role is judged like the [research engineer](https://one-course.com/books/quant/17/en/chapter/19-quant-developer#def-in-quant-developer-research)’s of chapter 19: on how fast and how safely models reach production and on how rarely they fail there. It is closer to software engineering than to research, and its pay and titles follow software engineering’s more than research’s.

## 21.3 The data engineer

A trading firm is a consumer of data before it is anything else: market data by venue and by day, reference data, corporate actions, alternative data (Book 7, chapter 12), its own orders and fills. The [data engineer](#def-in-machine-learning-and-data-engineer-roles) owns the path from the vendor’s or venue’s file to the researcher’s query: capture and storage of ticks (Book 15, chapter 2), columnar formats and tick stores (chapters 3 and 4), point-in-time versions that let a backtest see only what was known at the time, feature stores that serve the same values to training and to production (Book 12, chapter 24), and the checks that catch a vendor’s silent change before a model trains on it.

The work is unglamorous and indispensable: a research result is only as good as the data under it, and a missing day, a wrong timestamp or a survivor-biased universe can produce a backtest that is false in every detail. [Data engineers](#def-in-machine-learning-and-data-engineer-roles) are judged on completeness, correctness and timeliness of what they serve, and on the cost of storing and moving it.

## 21.4 Growth and pay

**As of September 2025 — Machine-learning and data roles in the filings.**

Labour condition applications at the sourced finance employers, fiscal 2021 and 2025: machine-learning and data titles 136 and 1 051 (banks 920 in 2025); software engineers 2 618 and 5 727; [quantitative researchers](https://one-course.com/books/quant/17/en/chapter/17-quantitative-researcher#def-in-quant-researcher-def) 1 287 and 1 297; all families 4 422 and 8 808. Fiscal 2025, [data scientists](#def-in-machine-learning-and-data-engineer-roles) (15-2051) against software developers (15-1252) at the same employers: median offered base $138 000 against $155 000; by [wage level](https://one-course.com/books/quant/17/en/chapter/14-pay-levels-by-role-firm-type-and-seniority#def-in-pay-levels-by-role-firm-type-and-seniority-soc) $-\$20\,500$ (I), $-\$18\,300$ (II), $-\$19\,000$ (III), $+\$2\,220$ (IV). [Base salary](https://one-course.com/books/quant/17/en/chapter/13-how-pay-works#def-in-how-pay-works-base) only.

**Method 21.2 (Growth of counts).**

Let $n_t$ be the filings in year $t$. Model them as Poisson with mean $\lambda_0e^{g t}$; the weighted least-squares fit of $\log n_t$ on $t$ with weights $n_t$ estimates the continuous growth rate $g$ with standard error $1/\sqrt{\sum_t n_t(t-\bar t)^2}$. With two years $t_0,t_1$ it is $g=\log(n_1/n_0)/(t_1-t_0)$ with standard error $\sqrt{1/n_0+1/n_1}/(t_1-t_0)$, and the annual rate is $e^g-1$. Each count has an exact Poisson interval.

The machine-learning and data filings grew at 66.7% a year (95% interval 59.4% to 74.3%), the software engineers’ at 21.6% (20.2% to 23.0%); the [quantitative researchers](https://one-course.com/books/quant/17/en/chapter/17-quantitative-researcher#def-in-quant-researcher-def)’ did not grow (0.2%, $-1.7\%$ to 2.1%) and the traders’ fell slightly ([Figure 21.2](#fig-in-machine-learning-and-data-engineer-growth)). The family’s share of all filings rose from 3.1% to 11.9%. The Poisson interval is the smallest honest one: filings come in batches from a few employers, and nine in ten of the 2025 machine-learning filings came from banks, so the true uncertainty about the industry’s hiring is larger than the interval says.

![Labour condition applications by role family at the sourced finance employers, fiscal 2021 and 2025, with exact 95% Poisson intervals (barely visible at these counts). Data: data/industry/lca_ranges.csv, through in_mldata.counts.](https://one-course.com/images/onecourse/chapters/quant-17/in-machine-learning-and-data-engineer/fig-5eb26f06ea05.svg)

***Figure 21.1.** Labour condition applications by role family at the sourced finance employers, fiscal 2021 and 2025, with exact 95% Poisson intervals (barely visible at these counts). Data: `data/industry/lca_ranges.csv`, through `in_mldata.counts`.*

![Annual growth of filings by role family between fiscal 2021 and 2025, with 95% intervals under the Poisson model. Data: in_mldata.trends, from firm.roles.filing_trend.](https://one-course.com/images/onecourse/chapters/quant-17/in-machine-learning-and-data-engineer/fig-d7e40bcbc26f.svg)

***Figure 21.2.** Annual growth of filings by role family between fiscal 2021 and 2025, with 95% intervals under the Poisson model. Data: `in_mldata.trends`, from `firm.roles.filing_trend`.*

The pay comparison is less flattering than the growth. At the same employers, filings coded as [data scientists](#def-in-machine-learning-and-data-engineer-roles) carry a median offered base $17 000 below the software developers’ (95% interval $-\$20\,000$ to $-\$15\,000$), and about $19 000 below at each of levels I to III; only at level IV are the two level ([Figure 21.3](#fig-in-machine-learning-and-data-engineer-gap)). Most of the growth is at banks, whose filings do not say which business a job serves; the few machine-learning filings at trading firms carry higher medians ($175 000 at market makers, $193 000 at platforms and $192 912 at systematic funds, each from ten to twelve filings), in line with those firms’ researchers and engineers.

![Median offered base of data-scientist filings minus software developers’ at the sourced finance employers, fiscal 2025, overall and by wage level, with 95% bootstrap intervals. Data: data/industry/lca_ds_gap.csv, through in_mldata.gaps.](https://one-course.com/images/onecourse/chapters/quant-17/in-machine-learning-and-data-engineer/fig-f87e983ebfd5.svg)

***Figure 21.3.** Median offered base of data-scientist filings minus software developers’ at the sourced finance employers, fiscal 2025, overall and by [wage level](https://one-course.com/books/quant/17/en/chapter/14-pay-levels-by-role-firm-type-and-seniority#def-in-pay-levels-by-role-firm-type-and-seniority-soc), with 95% bootstrap intervals. Data: `data/industry/lca_ds_gap.csv`, through `in_mldata.gaps`.*

**As of May 2025 — Where data scientists work: the survey.**

Occupational survey, May 2025, [data scientists](#def-in-machine-learning-and-data-engineer-roles) (15-2051): finance and insurance 46 730 employed, median $124 770; information 32 410, median $141 440; professional, scientific and technical services 69 730, median $126 730. Within finance: credit intermediation 12 450 (median $130 300), insurance carriers 15 090 ($107 680), securities 7 210 ($134 510). Software publishers 10 950 ($156 220).

Finance and insurance employs more [data scientists](#def-in-machine-learning-and-data-engineer-roles) than the information sector, though at a lower median, and the securities industry is a small part of it ([Figure 21.4](#fig-in-machine-learning-and-data-engineer-survey)). The picture matches the filings: data science in finance is mostly a bank and insurance job, and in trading it is mostly one of the tools of research.

![Data scientists employed by sector and industry, May 2025 occupational survey; labels give the median annual wage in $ thousand. The top three bars are sectors; the four below are industries within them (banks, insurance carriers and securities within finance and insurance; software publishers within information). Data: data/industry/oews_roles.csv, through in_mldata.survey.](https://one-course.com/images/onecourse/chapters/quant-17/in-machine-learning-and-data-engineer/fig-97742ee5cbeb.svg)

***Figure 21.4.** [Data scientists](#def-in-machine-learning-and-data-engineer-roles) employed by sector and industry, May 2025 occupational survey; labels give the median annual wage in $ thousand. The top three bars are sectors; the four below are industries within them (banks, insurance carriers and securities within finance and insurance; software publishers within information). Data: `data/industry/oews_roles.csv`, through `in_mldata.survey`.*

## 21.5 Tutorial: the fastest-growing title

**Goal.** Measure the growth of the machine-learning and data filings with its uncertainty, and put their pay beside the software engineers’. **End state:** Figures [21.1](#fig-in-machine-learning-and-data-engineer-counts), [21.2](#fig-in-machine-learning-and-data-engineer-growth) and [21.3](#fig-in-machine-learning-and-data-engineer-gap).

1. **Counts.** Sum chapter 14’s `lca_ranges.csv` over employer kinds for each role family and year; `firm.roles.poisson_interval(n)` gives each count’s exact interval.
2. **Growth.** `filing_trend(years, counts)` fits the log-linear Poisson model ([Listing 21.1](#lst-in-machine-learning-and-data-engineer-trend)). `def filing_trend (years, counts, z=1.96 ): """Log-linear growth of counts over years by Poisson-weighted least squares (weights = counts): continuous rate g, its standard error, and the annual compound rate exp(g) - 1 with a z-interval. With two years it reduces to log(n1 / n0) / (t1 - t0) and se sqrt(1/n0 + 1/n1) / (t1 - t0).""" t, n = np.asarray(years, float ), np.asarray(counts, float ) w = n tb = np.sum(w * t) / np.sum(w) sxx = np.sum(w * (t - tb) ** 2 ) g = float (np.sum(w * (t - tb) * np.log(n)) / sxx) se = float (math.sqrt(1.0 / sxx)) return {" rate " : g, " se " : se, " annual " : math.exp(g) - 1.0 , " lo " : math.exp(g - z * se) - 1.0 , " hi " : math.exp(g + z * se) - 1.0 }` **Listing 21.1.** Growth of filing counts with its standard error. code/firm/roles/firm_roles.py
3. **Pay.** `in_lca_ds_derive.py` compares data-scientist and software-developer filings level by level from chapter 14’s cache with `median_gap` (chapter 19).
4. **Survey.** `oews_roles.csv` for 15-2051 by sector and industry.

For the machine-learning family: $g=\log(1\,051/136)/4=0.511$ with standard error 0.023; 66.7% a year. Its 2025 count’s Poisson interval is 988 to 1 117.

**What to change next.** Read the other fiscal years and fit the trend over all of them; separate banks from trading firms; replace the Poisson model by one that allows for employers filing in batches (a negative binomial, or a bootstrap over employers).

## 21.6 Build: machine-learning cards and filing trends

**Purpose.** Add the machine-learning and data roles to the registry, and measure the growth of a count with its uncertainty.

**Interface.** `firm.roles`: the cards `data scientist`, `data engineer`, `machine-learning engineer`; `poisson_interval(n, level) -> (lo, hi)`; `filing_trend(years, counts, z) -> dict(rate, se, annual, lo, hi)`.

**Rules.** Counts are Poisson unless the caller says otherwise; the interval is exact (Garwood); growth is continuous in the fit and reported as an annual compound rate.

**Acceptance tests.** `code/firm/roles/tests/`: two points give the closed form; counts doubling each year give $\log 2$ exactly; the interval for ten is 4.80 to 18.39; zero gives a lower bound of zero.

**Stretch.** Overdispersion; a trend per employer kind with a test of equal growth; forecasts with intervals.

Sources and further reading

- Bureau of Labor Statistics, SOC 2010 to 2018 crosswalk; occupational survey, May 2025.
- O*NET OnLine, occupations 15-2051 and 15-1243 (CC BY 4.0).
- Chapter 14’s tables from the Department of Labor’s LCA files.
- Garwood, F. (1936), Fiducial limits for the Poisson distribution, *Biometrika* 28(3/4), 437–442.

## 21.7 Exercises

**Exercise 21.1 ★.**

By what factor did the machine-learning and data filings grow between fiscal 2021 and 2025, and what is that as an annual rate?

**Solution of Exercise 21.1.**

$1\,051/136=7.73$; over four years $7.73^{1/4}-1=66.7\%$ a year.

**Exercise 21.2 ★.**

What share of all filings did the family have in each year?

**Solution of Exercise 21.2.**

$136/4\,422=3.1\%$ in fiscal 2021 and $1\,051/8\,808=11.9\%$ in fiscal 2025.

**Exercise 21.3 ★.**

From the survey box, which employs more [data scientists](#def-in-machine-learning-and-data-engineer-roles), finance and insurance or the information sector, and which pays the higher median?

**Solution of Exercise 21.3.**

Finance and insurance employs more (46 730 against 32 410); the information sector pays the higher median ($141 440 against $124 770).

**Exercise 21.4 ★★.**

Compute the standard error of the growth rate for the software engineers and explain why it is smaller than the machine-learning family’s.

**Solution of Exercise 21.4.**

$\sqrt{1/2\,618+1/5\,727}/4=0.0059$, against 0.023: the standard error of a log count is about $1/\sqrt n$, and the software engineers’ counts are twenty to forty times larger.

**Exercise 21.5 ★★.**

Why can the growth of filings coded 15-2051 not be measured from fiscal 2021, although the growth of the title family can?

**Solution of Exercise 21.5.**

Fiscal 2021 filings were coded in the 2010 classification, which had no data-scientist occupation (the work was filed as statisticians, operations research analysts, computer scientists or other codes). The title family is defined by job titles, which exist in both years.

**Exercise 21.6 ★★.**

A graduate compares a data-scientist offer at a bank with a quantitative-researcher offer at a systematic fund. What does the chapter’s evidence say, and what does it not?

**Solution of Exercise 21.6.**

The filings say that data-scientist offers at banks run below software developers’ at the same level, and that systematic funds’ researchers are offered much more (chapter 17); growth of filings says the bank job is in demand. They do not say what either job pays in bonus, which work the bank’s [data scientist](#def-in-machine-learning-and-data-engineer-roles) would do, or how the two careers diverge.

**Exercise 21.7 ★★★.**

*Coding.* Suppose both years’ machine-learning filings came in batches of ten, so that each count behaves like ten times a Poisson count. Recompute the growth rate’s standard error and interval.

**Solution of Exercise 21.7.**

Each count’s variance is ten times its mean, so the standard error grows by $\sqrt{10}$: 0.072 instead of 0.023. The annual growth stays 66.7%, but its interval widens to 44.8% to 92.0%.

**Exercise 21.8 ★★★.**

*Find the flaw.* “Machine-learning filings grew 67% a year while researchers’ did not grow, so trading firms are replacing [quantitative researchers](https://one-course.com/books/quant/17/en/chapter/17-quantitative-researcher#def-in-quant-researcher-def) with [data scientists](#def-in-machine-learning-and-data-engineer-roles).”

**Solution of Exercise 21.8.**

Nine in ten of the 2025 machine-learning filings came from banks, not trading firms; filings count visa applications for new hires, not employment; researchers’ flat filings are not a fall; and a trading firm’s researchers use machine learning under their own title, so a change of method need not change titles.

## 21.8 Problem: The Fastest-Growing Title

**Problem 21.1.**

Weekend problem — the fastest-growing title

A student deciding between a statistics master’s with a machine-learning track and a financial mathematics master’s asks which way the industry’s hiring is moving.

**Part I — The roles.**

1. Define the [data scientist](#def-in-machine-learning-and-data-engineer-roles) and the [data engineer](#def-in-machine-learning-and-data-engineer-roles) .
2. What distinguishes a research-embedded machine-learning role?
3. What does the machine-learning engineer build, and what is the handoff specification?
4. What does a [data engineer](#def-in-machine-learning-and-data-engineer-roles) own at a trading firm?
5. How does the classification treat [data engineers](#def-in-machine-learning-and-data-engineer-roles) ?

**Part II — The counts.**

6. Give the family counts for fiscal 2021 and 2025.
7. State the growth model and its standard error.
8. Give the annual growth rates with intervals for the machine-learning family, software engineers and researchers.
9. Why is the Poisson interval the smallest honest interval?
10. Which kind of employer filed most of the 2025 machine-learning applications?

**Part III — The pay.**

11. Give the data-scientist gap against software developers overall and by level.
12. What do the trading firms’ few machine-learning filings show?
13. Where do [data scientists](#def-in-machine-learning-and-data-engineer-roles) work, by the survey?
14. Why might a bank’s [data scientist](#def-in-machine-learning-and-data-engineer-roles) be paid less than its software developer at the same level?
15. What does [base salary](https://one-course.com/books/quant/17/en/chapter/13-how-pay-works#def-in-how-pay-works-base) leave out here?

**Part IV — The verdict.**

16. State the *named result* : the annual growth rate of machine-learning and data filings at finance employers, fiscal 2021 to 2025, with its interval, and the median-base gap to software developers at the same [wage level](https://one-course.com/books/quant/17/en/chapter/14-pay-levels-by-role-firm-type-and-seniority#def-in-pay-levels-by-role-firm-type-and-seniority-soc) .
17. Is the growth about trading?
18. Which master’s prepares for which jobs?
19. What should the student ask an employer about a data-science role?
20. In two sentences, answer the student.

**Solution of Problem 21.1.**

1. As in the chapter’s definition.
2. The problem: low signal-to-noise, few independent observations, non-stationarity and a monetary cost of error.
3. The systems that train, serve and monitor models; the contract of what a model package must contain between builders and runners (Book 12).
4. The path from vendors’ and venues’ files to the researchers’ queries: capture, storage, point-in-time versions, feature stores, checks.
5. It has no data-engineer occupation; O*NET lists the title under database architects, and most are filed as software developers.
6. 136 and 1 051.
7. Poisson counts with log-linear mean; $g=\log(n_1/n_0)/4$ , standard error $\sqrt{1/n_0+1/n_1}/4$ .
8. 66.7% (59.4% to 74.3%); 21.6% (20.2% to 23.0%); 0.2% ( $-1.7\%$ to 2.1%).
9. It assumes independent filings; batches from a few employers widen the true uncertainty.
10. Banks: 920 of 1 051.
11. $-\$17\,000$ overall ( $-\$20\,000$ to $-\$15\,000$ ); $-\$20\,500$ , $-\$18\,300$ , $-\$19\,000$ at levels I–III; $+\$2\,220$ at IV.
12. Medians near those firms’ researchers and engineers: $175 000 to $193 000, from ten to twelve filings each.
13. Professional and technical services, finance and insurance, and information; in finance mostly banking and insurance.
14. The filings do not say; the job may serve business lines paid less than engineering, or be filed at the analyst end of the title.
15. Bonus and equity.
16. 66.7% a year (59.4% to 74.3%); about $19 000 below software developers at levels I–III, level at IV.
17. Mostly not: most of it is at banks.
18. The statistics master’s prepares for data science and research-embedded machine learning; the financial mathematics master’s for bank quant and research roles (chapter 29).
19. Which business the role serves, what it ships, who owns the models, and how its pay compares with engineers’.
20. Hiring of machine-learning and data titles in finance grew about two-thirds a year to fiscal 2025, mostly at banks and at pay below the engineers’. Choose the master’s by the job you want: trading research still hires by research title, and uses machine learning as a method.

## 21.9 Interview questions

**Interview question 21.1 ★ mle, researcher.**

Why is shuffled k-fold cross-validation wrong for a daily return-prediction model?

**Solution of Interview question 21.1.**

Shuffling puts future observations in the training folds of past test points and lets overlapping labels leak across folds; use walk-forward or purged and embargoed splits (Book 12, chapter 3).

*What the interviewer is looking for: leakage through time and overlapping labels.*

**Interview question 21.2 ★ mle.**

A model’s live accuracy is lower than its validation accuracy. List five causes in order of how you would check them.

**Solution of Interview question 21.2.**

A feature computed differently in production; look-ahead in the training data; a change in the data vendor or the market; overfitting to the validation set; and a population drift. Check the cheap and certain first: features, then data, then the model.

*What the interviewer is looking for: training-serving skew before drift.*

**Interview question 21.3 ★★ mle, developer.**

How do you make sure a feature has the same value in training and in production?

**Solution of Interview question 21.3.**

One definition in a feature store used by both, point-in-time joins for training, and a live comparison of served values against recomputed ones.

*What the interviewer is looking for: one definition and a check.*

**Interview question 21.4 ★★ developer.**

A vendor silently changes the definition of a field. How would your pipeline notice?

**Solution of Interview question 21.4.**

Schema and distribution checks on arrival (types, ranges, null rates, quantiles against history), reconciliation with a second source where one exists, and alerts that stop downstream jobs.

*What the interviewer is looking for: data contracts and distribution monitoring.*

**Interview question 21.5 ★★ mle, researcher.**

Your gradient-boosted model beats a linear model by 0.3% of accuracy. Is it worth deploying?

**Solution of Interview question 21.5.**

Only if the gain survives costs and time: test it out of sample over several periods, measure its value in money after costs, and weigh the maintenance and risk of the more complex model; a 0.3% gain in accuracy may be noise.

*What the interviewer is looking for: economic value and stability, not accuracy.*

**Interview question 21.6 ★★★ mle.**

Design the storage for ten years of order-book data that researchers query by symbol and time range.

**Solution of Interview question 21.6.**

Partition by date and symbol in a columnar format with compression, keep an index of time ranges per file, store raw and normalised versions, and serve queries through an engine that prunes partitions (Book 15, chapters 3 and 4).

*What the interviewer is looking for: partitioning for the access pattern and columnar storage.*
