---
title: "Estimation"
book: "The Interview Book"
subject: quant
language: en
chapter: 9
exercises: 0
source: https://one-course.com/books/quant/18/en/chapter/9-estimation
---

# Chapter 9 — Estimation

“How many messages does the US options market’s consolidated feed publish in its busiest ten milliseconds of a month?” The candidate who answers “between two hundred thousand and two million, most likely around six hundred thousand” and shows the chain (a few thousand series on a busy underlying, seventeen exchanges quoting each, several moves of the underlying in the window) has answered better than one who happens to know the published figure and cannot say how it comes about. An estimation question measures two things: the ability to decompose an unknown quantity into knowable pieces, and the honesty to say how wrong the result might be.

## 9.1 The Fermi chain

**Definition 9.1 (Fermi estimate).**

A *Fermi estimate* is an order-of-magnitude estimate of an unknown quantity obtained by writing it as a product (or sum) of factors each of which can be estimated from common knowledge, and multiplying the estimates.

**Method 9.2 (Building a Fermi chain).**

1. Write the quantity as a product of three to six factors, each with a unit, so that the units cancel to the answer’s unit.
2. Estimate each factor from something you know: a population, a price, a time, a size.
3. Where only bounds come to mind, use their geometric mean: between 200 and 20 000 means about 2 000.
4. Multiply in powers of ten first, then the mantissas.
5. Give an interval (next section) and check against a second chain built from different factors.

**Example 9.3 (The hook’s chain).**

Suppose (these are assumptions, stated as such) that a heavily traded underlying has about 5 000 listed option series, that 17 exchanges quote each series, and that in a burst the underlying moves about ten times in ten milliseconds across the names that move together, each move requoting every series once: $5\,000 \times 17
\times 10 = 850\,000$ messages. The published peak for July 2026 was 0.89 million messages in ten milliseconds ([Box 9.1](#dat-iv-estimation-anchors)). The chain’s closeness is partly luck; its structure, a product of breadth, venues and event rate, is what the interviewer wants.

Errors in a chain multiply, so they add on the log scale, and independent errors partly cancel.

**Proposition 9.4 (How uncertainty grows along a chain).**

If each of $k$ factors is estimated with an independent log error of standard deviation $s_i$, the log error of the product has standard deviation $\big(\sum s_i^2\big)^{1/2}$. With four factors each at $s = 0.35$ (each within a factor of about 1.8 with 90% probability), the product has $s = 0.70$, and a 90% interval runs from a factor $e^{1.645 \times 0.7} \approx 3.2$ below to 3.2 above the estimate.

**Proof.** The log of a product is the sum of the logs; variances of independent errors add. The interval factor is $\exp(z_{0.95}\, s)$ for a lognormal error. ∎

## 9.2 Intervals

**Definition 9.5 (Calibrated interval).**

A stated $c$ interval for an unknown quantity is a *calibrated interval* if, over many such statements, the true value falls inside a fraction $c$ of them.

Calibration is a property of the estimator, not of one interval, and people are usually overconfident: in the studies surveyed by Lichtenstein, Fischhoff and Phillips (1982), intervals stated with high confidence contained the truth much less often than stated. [Figure 9.1](#fig-iv-estimation-calib) shows the arithmetic: an estimator whose log errors have standard deviation 1 but who builds intervals as if it were 0.5 hits a stated 90% interval only 59% of the time.

![Actual against stated coverage of central intervals when log errors have standard deviation 1: a calibrated estimator (intervals built with 1) lies on the diagonal; an overconfident one (intervals built with 0.5) falls far below it. Circles: 20 000 simulated estimates. Data: fig_iv_calib.py.](https://one-course.com/images/onecourse/chapters/quant-18/iv-estimation/fig-2918ed99ed93.svg)

***Figure 9.1.** Actual against stated coverage of central intervals when log errors have standard deviation 1: a calibrated estimator (intervals built with 1) lies on the diagonal; an overconfident one (intervals built with 0.5) falls far below it. Circles: 20 000 simulated estimates. Data: `fig_iv_calib.py`.*

An interval is scored by a proper scoring rule that rewards narrowness and penalises misses in proportion to their size. The interval score of Gneiting and Raftery (2007) for a central $(1-\alpha)$ interval $[l, u]$ and outcome $x$ is $(u - l) + \tfrac2\alpha(l - x)\mathbf 1_{x<l} + \tfrac2\alpha(x - u)\mathbf 1_{x>u}$: an honest interval minimises its expectation. The Brier score of One Quant Book 16, chapter 26, plays the same role for probabilities of events.

## 9.3 Market estimation questions

Market questions are Fermi questions with a trading context: volumes, notionals, messages, bytes, capacities. The factors come from the reader’s knowledge of the markets of One Quant Books 1 to 3 and the systems of Books 13 and 14, and the published anchors for a few of them are in [Box 9.1](#dat-iv-estimation-anchors). Two habits distinguish good answers. The first is to separate the count of things from the activity per thing (series times quote rate; accounts times trades per account). The second is to state the peak-to-average ratio explicitly when a capacity is asked, because systems are sized for the burst, not the day.

**As of September 2026 — Anchors for market estimates.**

Global over-the-counter FX turnover was USD 9.6 trillion a day in April 2025 (BIS Triennial Survey). The US options consolidated feed (OPRA) published peaks in July 2026 of 63.9 million messages in a second and 0.89 million in ten milliseconds, and its capacity projection for July 2026 was 4.403 gigabits per 100 ms on one stream. CME Group’s exchanges averaged 28.1 million contracts a day in 2025. The world’s population was 8.2 billion in 2024 (UN World Population Prospects 2024).

## 9.4 Checking an estimate with a second chain

A second chain that shares no factor with the first is the best check an estimate can have. If the two agree within their intervals, combine them by a weighted geometric mean (weights inversely proportional to the log variances); if they disagree by more, one of the chains has a wrong factor, and finding it is worth more than averaging ([Interview question 9.11](#iq-iv-estimation-11)).

## 9.5 Worked answers

**Example 9.6 (Multiplying in logarithms).**

A chain ends in $3.7 \times 10^3 \times 2.2 \times 10^4 \times 0.45$. In base-10 logarithms, $\log 3.7 \approx 0.57$, $\log 2.2 \approx 0.34$ and $\log 0.45 \approx -0.35$, so the product has logarithm $0.57 + 0.34 - 0.35 + 7 = 7.56$, and $10^{0.56} \approx 3.6$: about $3.6 \times 10^7$. The exact product is $3.663 \times 10^7$. Estimators who work in logarithms carry the exponent separately, never lose a zero, and can state the uncertainty of each factor in the same units.

**Example 9.7 (A chain, its interval and a second chain).**

*“How many order messages (new orders, modifications, cancels) does an options market maker send on a trading day? Give a 90% interval.”*

*First chain.* The firm quotes, say, 2 000 option series actively on one venue; each quote is updated about 5 times a second on average, as the underlying moves; the session has 6.5 hours, $23\,400$ seconds. The product is $2\,000 \times 5
\times 23\,400 \approx 2.3 \times 10^8$. The instruments and the update rate are each uncertain to a factor of 3 at 90%; in natural logarithms each has a standard deviation of $\ln 3/1.645 \approx 0.67$, together $0.67\sqrt2 \approx 0.94$, and the 90% factor is $e^{1.645 \times 0.94} \approx 4.7$. Interval: about $5 \times 10^7$ to $1.1 \times 10^9$.

*Second chain, from the pipe.* If the firm’s order-entry traffic on that venue averages 10 megabits a second and a message is about 50 bytes, it sends $10^7/(8 \times 50) = 25\,000$ messages a second, $5.85 \times 10^8$ a day. The two chains share no factor and agree within a factor of 2.5, inside the interval; if they had disagreed by a factor of 20, the right answer would have been to find out which assumption was wrong, not to average them (the previous section).

## 9.6 Question bank

**Interview question 9.1 ★ trader • market maker.**

How many litres of coffee does a trading floor of 400 people drink in a year? Show the chain.

**Solution of Interview question 9.1.**

400 people $\times$ 2 cups a day $\times$ 220 working days $\times$ 0.25 litre $= 44\,000$ litres a year; a 90% interval of a factor of about 2 either way (15 000 to 90 000), since the cups per person is the weakest factor. There is no published figure; the question scores the chain and the units.

*What the interviewer is looking for: a clean chain with units and a stated interval.*

**Interview question 9.2 ★ trader, developer • any.**

How many seconds are there in a year, and how many seconds of regular trading does a US equity market have in a year?

**Solution of Interview question 9.2.**

$365 \times 24 \times 3\,600 = 31\,536\,000$, about $3.15 \times 10^7$ (“$\pi \times 10^7$” is a good mnemonic). Regular US equity trading: $252 \times 6.5 \times 3\,600 = 5\,896\,800$, about $5.9 \times 10^6$, under a fifth of the year.

*What the interviewer is looking for: exact arithmetic with the right calendar and session length.*

**Interview question 9.3 ★ trader, researcher • any.**

You are sure a quantity lies between 200 and 20 000 and have no other information. What single number do you give, and why not 10 100?

**Solution of Interview question 9.3.**

2 000, the geometric mean. Bounds a factor of 100 apart describe uncertainty on the log scale; the midpoint of the logs is the geometric mean, and it is a factor of 10 from each bound. The arithmetic mean 10 100 is only a factor of 2 from the upper bound and 50 from the lower, which treats a factor-of-50 miss on one side like a factor-of-2 miss on the other.

*What the interviewer is looking for: thinking in logs for multiplicative uncertainty.*

**Interview question 9.4 ★ trader, bank • bank.**

Give a 90% interval for the daily turnover of the global FX market, with a chain.

**Solution of Interview question 9.4.**

A chain, with each factor an assumption to be checked: suppose world trade in goods and services is of the order of 30 trillion dollars a year; FX turnover is dominated by financial flows and rolling of swaps, perhaps 50 to 100 times the trade flows a year, over 250 days: 6 to 12 trillion a day. A 90% interval of 3 to 15 trillion (geometric centre 6.7). The BIS survey measured 9.6 trillion a day in April 2025 ([Box 9.1](#dat-iv-estimation-anchors)), inside the interval.

*What the interviewer is looking for: a chain from a known macro quantity, a multiplier for financial activity, and a wide honest interval.*

**Interview question 9.5 ★★ developer, trader • market maker.**

Give a 90% interval for the largest number of messages the US options consolidated feed publishes in any ten milliseconds of a month. What would you ask before building a system to receive it?

**Solution of Interview question 9.5.**

The chain of [Example 9.3](#ex-iv-estimation-opra) gives about 850 000; with each factor uncertain by a factor of 2, a 90% interval of 200 000 to 2 million. The July 2026 peak was 0.89 million ([Box 9.1](#dat-iv-estimation-anchors)). Before building: ask for the published peak statistics and capacity projections, the message format (bytes per message, packets per message), the number of lines and their split, and the growth rate, because systems are sized for next year’s burst.

*What the interviewer is looking for: a structural chain, a wide interval, and the engineer’s instinct to ask for the published capacity.*

**Interview question 9.6 ★★ trader • proprietary firm.**

Estimate the number of futures and options contracts traded on an average day on the exchanges of the largest US derivatives exchange group. Give an interval.

**Solution of Interview question 9.6.**

Perhaps 50 actively traded products (interest-rate futures and options, equity index, energy, metals, agriculture) at 500 000 contracts a day on average, with a long tail: about 25 million, 90% interval 10 to 60 million. The group’s exchanges averaged 28.1 million contracts a day in 2025.

*What the interviewer is looking for: a count-times-activity chain and an interval that contains the answer without being useless.*

**Interview question 9.7 ★★ trader, researcher • any.**

How many people alive today have their birthday today?

**Solution of Interview question 9.7.**

About 8.2 billion people divided by 365.25 days: about 22 million. Birthdays are not quite uniform over the year (seasonal birth patterns), which moves the answer by some percent, not a factor.

*What the interviewer is looking for: the division, and awareness of the uniformity assumption.*

**Interview question 9.8 ★★ developer • proprietary firm.**

Estimate the storage needed for one day of top-of-book updates for all US-listed stocks, uncompressed.

**Solution of Interview question 9.8.**

About 5 000 stocks $\times$ 20 000 top-of-book changes a day on average (many more for the most active, far fewer for the tail) $\times$ 32 bytes (timestamp, symbol identifier, bid and ask prices and sizes) $= 3.2
\times 10^9$ bytes, about 3 GB a day. The distribution across stocks is very skewed, so the per-stock average is the weak factor; a second chain from total message counts would check it.

*What the interviewer is looking for: a record-size estimate and awareness of skew across symbols.*

**Interview question 9.9 ★★ researcher, risk • systematic fund.**

Over the last month you gave ten 90% intervals and five contained the true value. How surprising is that if you were calibrated? What do you change?

**Solution of Interview question 9.9.**

If each interval contains the truth with probability 0.9 independently, the number of hits is binomial(10, 0.9), and $\P(X \le 5) \approx 0.0016$: very surprising. The intervals are too narrow. Widen them, roughly doubling their log width if the misses were large ([Figure 9.1](#fig-iv-estimation-calib)), and keep a record of hits to recalibrate.

*What the interviewer is looking for: a binomial test of calibration and the practical correction.*

**Interview question 9.10 ★★★ researcher, trader • any.**

A chain has four independent factors, each of which you believe is within a factor of 2 of the truth with 90% probability. Within what factor is the product, at 90%? Why is it less than $2^4$?

**Solution of Interview question 9.10.**

A factor of 2 at 90% means a log standard deviation of $\ln 2/1.645 \approx 0.42$ per factor. Four independent factors give $0.42 \times 2 = 0.84$, so the product is within a factor of $e^{1.645 \times 0.84} = 4$ of the truth at 90%, far less than $2^4 = 16$, because the four errors are unlikely to go the same way at once.

*What the interviewer is looking for: adding log variances, and why independent errors partly cancel.*

**Interview question 9.11 ★★★ researcher • systematic fund.**

Two independent chains give $10^6$ (log standard deviation 0.7) and $10^7$ (log standard deviation 1.0). What single estimate do you report, with what uncertainty? When would you refuse to combine them?

**Solution of Interview question 9.11.**

Weights $1/0.7^2 \approx 2.04$ and $1/1.0^2 = 1$ on the logs give $\exp\big((2.04 \ln 10^6 + \ln 10^7)/3.04\big)
\approx 2.1 \times 10^6$, with log standard deviation $1/\sqrt{3.04} \approx 0.57$. Refuse to combine when the chains disagree by more than their uncertainties allow (here the difference of the logs, 2.3, is about 1.9 standard deviations of the difference, $\sqrt{0.49 + 1} \approx 1.22$: borderline) or when they share a factor, so that their errors are not independent. Then look for the wrong factor.

*What the interviewer is looking for: inverse-variance weighting on the log scale, and a test of consistency before combining.*

**Interview question 9.12 ★★★ researcher, risk • bank.**

Two candidates give 90% intervals for daily FX turnover: $[5, 12]$ trillion and $[2, 5]$ trillion dollars. The answer is 9.6 trillion. Score both with the interval score, and explain why the rule makes honest intervals the best strategy.

**Solution of Interview question 9.12.**

With $\alpha = 0.1$: $[5, 12]$ contains 9.6 and scores its width, 7. $[2, 5]$ misses by 4.6 and scores $3 + 20
\times 4.6 = 95$. The rule is proper: for any belief about the quantity, the expected score is minimised by reporting the belief’s own 5% and 95% quantiles, so shading the interval narrower (or wider) than one believes raises the expected penalty (Gneiting and Raftery, 2007).

*What the interviewer is looking for: computing the score and knowing why a proper rule makes honesty optimal.*

**Interview question 9.13 ★★★ developer • market maker.**

At a peak of 63.9 million messages a second and an average of 40 bytes a message, what bandwidth does the options feed need, in gigabits a second? Compare with a published capacity of 4.403 gigabits per 100 milliseconds, and say what your estimate leaves out.

**Solution of Interview question 9.13.**

$63.9 \times 10^6 \times 40 \times 8 \approx 20.4$ gigabits a second. The capacity projection is 4.403 gigabits per 100 milliseconds, 44 gigabits a second: about twice the estimate. The estimate leaves out packet and protocol overhead, the fact that a one-second peak averages over shorter and higher bursts (the 10 ms peak of 0.89 million messages is 89 million a second), retransmission and a second, redundant stream, and growth.

*What the interviewer is looking for: a units-careful bandwidth calculation and the reasons capacity exceeds the average peak.*

Sources and further reading

- S. Lichtenstein, B. Fischhoff and L. D. Phillips, “Calibration of probabilities: the state of the art to 1980”, in D. Kahneman, P. Slovic and A. Tversky (eds), *Judgment under Uncertainty* , Cambridge University Press, 1982, 306–334.
- T. Gneiting and A. E. Raftery, “Strictly proper scoring rules, prediction, and estimation”, *Journal of the American Statistical Association* 102(477), 2007, 359–378.
- BIS, Triennial Central Bank Survey 2025; OPRA, key operating metrics (August 2026) and capacity projections (September 2025); CME Group, Form 10-K for 2025; UN, *World Population Prospects 2024* .
