---
title: "What Each Interview Tests, by Role"
book: "The Interview Book"
subject: quant
language: en
chapter: 1
exercises: 0
source: https://one-course.com/books/quant/18/en/chapter/1-what-each-interview-tests-by-role
---

# Chapter 1 — What Each Interview Tests, by Role

Two candidates leave the same building on the same afternoon with opposite verdicts on the same probability question. The trading candidate reached the right answer in four minutes and was marked down; the developer candidate gave a wrong first answer, tested it on a small case, found the error and fixed it, and was marked up. Nobody was inconsistent. The two interviewers were measuring different jobs: a trader must be right fast and say how sure she is; a developer must be right in the end and know how he knows. This chapter is about reading which job a question is measuring, because the same question has a different best answer in each room.

## 1.1 What an interview measures

An interview is a measurement. The firm wants to know how well a person will do a job it can describe; it cannot watch the person do the job for a year, so it samples a few hours of behaviour that it believes predicts the year. Everything in this book follows from taking that sentence seriously: a question is an instrument, an answer is a reading, and the reading is compared with a scale that the interviewer, not the candidate, holds.

**Definition 1.1 (Interview loop, technical interview).**

An *interview loop* is the ordered set of interviews one candidate sits for one role, with the rule that turns their separate scores into a decision. A *technical interview* is one whose questions have answers that can be checked: a number, a proof, a program, a design whose properties can be argued. Its opposite is the behavioural interview of [Chapter 28](https://one-course.com/books/quant/18/en/chapter/28-behavioural-and-fit#ch-iv-behavioural-and-fit), whose answers are accounts of past conduct.

**Definition 1.2 (Interview rubric).**

An *interview rubric* is the scoring sheet for one question or one interview: a short list of dimensions (for example correctness, speed, method, communication), each with anchored levels that say what an answer at that level contains, filled in by the interviewer before the loop’s scores are compared.

The rubric is the instrument that makes an interview a structured interview in the sense of One Quant Book 16, chapter 10: the same questions for every candidate, answers scored against anchors agreed in advance, interviewers scoring independently before they confer. The research on selection that Book 16 summarises ranks structured interviews first among selection procedures once earlier overcorrections are removed (Sackett, Zhang, Berry and Lievens, 2022), and the components that make an interview structured were catalogued a quarter of a century earlier (Campion, Palmer and Campion, 1997). Two consequences matter to a candidate. First, the rubric rewards what it lists and nothing else: an elegant digression scores nothing on a sheet that asks for a number and a check. Second, the rubric is written for the job, so the best answer to the same question differs by role, as the opening scene shows.

A loop is also a chain of filters, and its arithmetic is worth knowing before any question is asked.

**Proposition 1.3 (Pass rates through a loop).**

Let an application pass stage $k$ with probability $p_k$, independently of the other stages, for $k = 1, \dots, K$. The probability that it ends in an offer is $\prod_{k=1}^K p_k$; the expected number of stages sat per application is $\sum_{k=1}^{K} \prod_{j<k} p_j$; and with $n$ independent applications of offer probability $q$, the chance of at least one offer is $1 - (1-q)^n$.

**Proof.** Stage $k$ is reached only if stages $1, \dots, k-1$ are passed, so it is sat with probability $\prod_{j<k} p_j$; summing these indicators gives the expected count. The last statement is the complement of $n$ independent failures. ∎

**Example 1.4 (A three-stage funnel).**

An online assessment passed three times in ten, a screen passed one time in two and a final round passed two times in five give an offer probability of $0.3 \times 0.5 \times 0.4 = 0.06$ per application. Each application costs on average $1 + 0.3 + 0.15 = 1.45$ stages, so an offer costs about $1.45/0.06 \approx 24.2$ stages sat and $1/0.06 \approx 16.7$ applications. Twenty applications give at least one offer with probability $1 - 0.94^{20} \approx 0.71$. The numbers are illustrative; the lesson is structural: most of the variation between candidates is in the first stage, which is the cheapest to improve ([Chapter 4](https://one-course.com/books/quant/18/en/chapter/4-online-assessments#ch-iv-online-assessments)).

The independence assumption is optimistic in one direction and pessimistic in the other. A candidate who fails one firm’s assessment is more likely to fail the next firm’s, because the same weakness is measured twice; and a candidate who has sat ten final rounds is better at the eleventh. Both effects mean that the early applications are the ones to spend preparation on.

## 1.2 The roles and what each is probed on

The roles themselves are described in One Quant Book 17, chapters 16 to 25. The interviewer’s question is narrower: which abilities separate a good holder of this role from a poor one in the first two years, and which of them can be measured in an hour? The table below is this book’s answer, and it is also its map.

| Question family (chapter) | Trader | Researcher | Developer | ML engineer | Bank quant | Risk |
| --- | --- | --- | --- | --- | --- | --- |
| Mental arithmetic (8) | $\bullet$ | $\circ$ |  |  | $\circ$ | $\circ$ |
| Estimation (9) | $\bullet$ | $\circ$ | $\circ$ | $\circ$ | $\circ$ | $\circ$ |
| Probability (10, 11) | $\bullet$ | $\bullet$ | $\circ$ | $\circ$ | $\bullet$ | $\bullet$ |
| Brainteasers (12) | $\bullet$ | $\circ$ | $\circ$ |  |  |  |
| Betting and market making (13) | $\bullet$ | $\circ$ |  |  |  | $\circ$ |
| Statistics (14) | $\circ$ | $\bullet$ |  | $\bullet$ | $\circ$ | $\bullet$ |
| Linear algebra, calculus (15) |  | $\bullet$ |  | $\circ$ | $\bullet$ | $\circ$ |
| Stochastic calculus (16) |  | $\circ$ |  |  | $\bullet$ | $\circ$ |
| Options, fixed income (17, 18) | $\bullet$ | $\circ$ |  |  | $\bullet$ | $\bullet$ |
| Machine learning (19) |  | $\bullet$ |  | $\bullet$ |  |  |
| Research case study (20) |  | $\bullet$ |  | $\bullet$ |  |  |
| Algorithms, languages (21–24) | $\circ$ | $\circ$ | $\bullet$ | $\bullet$ | $\circ$ |  |
| Systems, design (25, 26) |  |  | $\bullet$ | $\circ$ |  |  |
| SQL and data (27) |  | $\circ$ | $\circ$ | $\bullet$ |  | $\bullet$ |
| Behavioural (28) | $\bullet$ | $\bullet$ | $\bullet$ | $\bullet$ | $\bullet$ | $\bullet$ |

*Where each role’s interviews spend their time: $\bullet$ a core family, asked in most loops; $\circ$ asked in some. A chapter number follows each family.*

Read by column, the table describes six loops.

*Trading roles* (the risk trader and execution trader of One Quant Book 17, chapter 16, and the market maker’s trader of Book 11). The interview measures speed with numbers, calibrated judgement under uncertainty, and behaviour when a counterparty knows more: mental arithmetic, probability at speed, estimation with an interval, and games in which the candidate quotes prices and is traded against. The best answer is fast, approximately right, and says how sure it is; a slow exact answer is scored lower than a quick answer with the right error bar.

*Research roles* (the quantitative researcher of Book 17, chapter 17). The interview measures whether the candidate can turn a vague question into a testable one and refuse to be fooled: probability and statistics in depth, regression pathologies, overfitting, and a research case study or take-home. The best answer states assumptions, tests them, and says what would change the conclusion.

*Developer roles* (the quant developer of Book 16, chapter 21, the research engineer and low-latency engineer of Book 17, chapters 19 and 20). The interview measures working code, reasoning about cost, and knowledge of what the machine does: algorithms, a language in depth, concurrency and systems, and a design problem. The best answer is correct, tested, and explains its complexity and its failure modes; the trading candidate’s four-minute answer in the hook would score badly here if it had not been checked.

*Machine-learning engineers* (Book 12, chapter 28) sit between research and development: statistics and validation with time, model design, and production code.

*Bank quants* (the desk strategist, model validator and risk quant of Book 17, chapter 18). The interview measures mathematics of pricing and its implementation: stochastic calculus, derivatives, numerical methods and a language. A model validator is asked more about what can go wrong with a model, a desk strategist more about what the desk needs by tomorrow morning.

*Risk and control roles* (the risk quant and the control functions of Book 17, chapter 24). The interview measures statistics of losses, derivatives at the level of their sensitivities, and judgement about what to escalate.

**Remark 1.5 (One question, several rubrics).**

“How many rolls of a fair die, on average, until the first six?” has one answer, six. A trading rubric scores the time to answer and whether the candidate volunteers the variance (thirty) when asked how long the wait could be. A research rubric scores the argument (the geometric law, or a one-line first-step equation). A developer rubric scores a check: a ten-line simulation, or the observation that the answer must exceed one and be finite. The candidate who knows the room gives the answer that rubric can record.

## 1.3 Formats and the time each question gets

Most loops are built from a small number of formats: a timed online assessment, a screen by telephone or video, and a final round of several interviews, sometimes with a trading game, a group exercise or a presentation (Chapters [2](https://one-course.com/books/quant/18/en/chapter/2-the-process-by-firm-type#ch-iv-the-process-by-firm-type) and [6](https://one-course.com/books/quant/18/en/chapter/6-the-final-round-and-trading-games#ch-iv-the-final-round-and-trading-games)). Within one interview the question types have typical lengths, and the length tells the candidate what is expected.

**Method 1.6 (Reading a question by its time budget).**

1. *Seconds* : arithmetic, a sign, a definition. The interviewer wants the answer, not the method.
2. *One to five minutes* : a probability, an estimate, a Greek, an output to predict. Say the method in one sentence, compute, check once, stop.
3. *Ten to twenty minutes* : a derivation, a coding problem, a game. Clarify, give a baseline, improve, test; think aloud ( [Chapter 5](https://one-course.com/books/quant/18/en/chapter/5-phone-and-technical-screens#ch-iv-phone-and-technical-screens) ).
4. *Thirty minutes or more* : a design, a case study, a take-home. Requirements first, numbers early, structure before detail (Chapters [26](https://one-course.com/books/quant/18/en/chapter/26-system-design#ch-iv-system-design) and [20](https://one-course.com/books/quant/18/en/chapter/20-research-case-studies-and-take-homes#ch-iv-research-case-studies-and-take-homes) ).
5. If the time budget is unclear, ask: “Do you want the number, or the argument?” costs three seconds.

A loop rarely fails a candidate on one bad answer. It fails candidates whose answers are consistently in the lower anchors on the dimensions the role weighs: the trader who never gives a number, the researcher who never checks an assumption, the developer who never tests. [Interview question 1.5](#iq-iv-what-each-interview-tests-by-role-5) shows the arithmetic of the other failure: a loop that demands a pass in every interview rejects good candidates often, which is why well-designed loops average scores rather than multiply vetoes.

## 1.4 Preparing by role

This book is organised by question family, not by role. A reader preparing for one role reads the core chapters of its column in the table of the previous section first, then the process chapters (Part I), then [Chapter 29](https://one-course.com/books/quant/18/en/chapter/29-mock-interviews#ch-iv-mock-interviews), whose six transcripts show one complete interview for each column. Each chapter’s bank is ordered from one star to three and tagged with the roles and firm types for which the question is typical; the tags describe the kind of question, never a firm that asked it. Every question in the book is new: none is reprinted from another book of this series or elsewhere, and a classic puzzle appears only as a variant whose origin is cited.

**Example 1.7 (A preparation order for a developer with a mathematics degree).**

The degree covers probability and linear algebra in principle, not at interview speed. The order is: [Chapter 21](https://one-course.com/books/quant/18/en/chapter/21-algorithms-and-data-structures#ch-iv-algorithms-and-data-structures) and the language chapter of the role (C++ or Rust), then Chapters [25](https://one-course.com/books/quant/18/en/chapter/25-concurrency-operating-systems-and-networks#ch-iv-concurrency-operating-systems-and-networks) and [26](https://one-course.com/books/quant/18/en/chapter/26-system-design#ch-iv-system-design), then [Chapter 10](https://one-course.com/books/quant/18/en/chapter/10-probability-i#ch-iv-probability-i) for speed, then Part I and [Chapter 28](https://one-course.com/books/quant/18/en/chapter/28-behavioural-and-fit#ch-iv-behavioural-and-fit). The mathematics is revised last because it decays slowest.

## 1.5 Worked answers

Each chapter of this book ends its lesson with a few questions answered in full, as a strong candidate would answer them aloud: the reasoning first, the number, then the check. They are not in the bank, and their numbers are asserted by the chapter’s tests like every other number in the book.

**Example 1.8 (Why a strong candidate needs several loops).**

*“You are well above the bar: your score in any one interview is your ability, one standard deviation above the pass mark, plus independent standard normal noise. A loop has five interviews and needs all five passed. How likely is an offer, and what does it mean for how many processes you run?”*

*Answer.* One interview is passed with probability $\Phi(1) \approx 0.841$; five in a row with $0.841^5 \approx 0.42$. A candidate well above the bar fails more loops than they pass. With three independent processes the chance of at least one offer is $1 - 0.58^3 \approx 0.81$; at one and a half standard deviations above the bar a single loop already gives $\Phi(1.5)^5 \approx 0.71$. Two consequences: run several processes in parallel ([Chapter 2](https://one-course.com/books/quant/18/en/chapter/2-the-process-by-firm-type#ch-iv-the-process-by-firm-type)), and read a rejection as a draw from a noisy test, not as a verdict ([Figure 1.1](#fig-iv-what-each-interview-tests-by-role-loops)). *Check:* with an ability of zero a single interview is a coin toss and the loop passes one time in 32, which matches $0.5^5$.

![Chance of passing a loop that needs every interview passed, by the candidate’s ability above the bar, when each interview adds independent standard normal noise. At = 1 a single interview is passed 84% of the time and a five-interview loop 42%. Data: fig_iv_loops.py.](https://one-course.com/images/onecourse/chapters/quant-18/iv-what-each-interview-tests-by-role/fig-2562bae06420.svg)

***Figure 1.1.** Chance of passing a loop that needs every interview passed, by the candidate’s ability above the bar, when each interview adds independent standard normal noise. At $\theta = 1$ a single interview is passed 84% of the time and a five-interview loop 42%. Data: `fig_iv_loops.py`.*

**Example 1.9 (Reading an advertisement into a plan).**

*“An advertisement for a quant developer on an execution platform asks for modern C++, Linux and networking, Python, ‘strong problem-solving’, and says trading knowledge is a plus. You have forty hours over four weeks. Plan them.”*

*Answer.* Map each line to the interviews it predicts: modern C++ to output-prediction and design questions ([Chapter 22](https://one-course.com/books/quant/18/en/chapter/22-c#ch-iv-cpp)); Linux and networking to the systems interview ([Chapter 25](https://one-course.com/books/quant/18/en/chapter/25-concurrency-operating-systems-and-networks#ch-iv-concurrency-operating-systems-and-networks)); problem-solving to the coding rounds ([Chapter 21](https://one-course.com/books/quant/18/en/chapter/21-algorithms-and-data-structures#ch-iv-algorithms-and-data-structures)); Python to a screen or a take-home ([Chapter 24](https://one-course.com/books/quant/18/en/chapter/24-python#ch-iv-python)); the “plus” to one conversation about the order path ([Chapter 26](https://one-course.com/books/quant/18/en/chapter/26-system-design#ch-iv-system-design)). Weight by how many interviews each feeds and how far you are from speed: 12 hours of timed coding, 10 of C++, 8 of systems, 4 of design, 4 of probability at interview speed, 2 on the behavioural stories: 40. Book the first mock interview for the end of week two, so that weeks three and four fix what it exposes.

## 1.6 Question bank

**Interview question 1.1 ★ trader, developer • any.**

You are asked: “A fair coin is tossed until it shows heads twice in a row; how many tosses on average?” Say what a trading interviewer and a developer interviewer would each be scoring, and give the answer each would want to hear.

**Solution of Interview question 1.1.**

The answer is 6. Let $a$ be the expected number of further tosses with no head pending and $b$ with one head just seen: $a = 1 + \tfrac12 b + \tfrac12 a$ and $b = 1 + \tfrac12 a$, so $a = 6$. The trading interviewer scores speed and a sense of spread: “six; it is a waiting time, so it is skewed, and a run of fifteen tosses would not surprise me” (the variance is 22). The developer interviewer scores the method and a check: the two equations, and either a ten-line simulation or the bounds (at least two tosses; at most a geometric wait of pairs, $2 \times 4 = 8$, since disjoint pairs are HH with probability one quarter).

*What the interviewer is looking for: knowing that the same number is scored for speed and calibration in one room and for method and verification in the other.*

**Interview question 1.2 ★ researcher • systematic fund.**

Name three things a research interview checks that a trading interview usually does not, with one question that would check each.

**Solution of Interview question 1.2.**

(i) Whether the candidate turns a vague claim into a test with a null hypothesis: “this signal predicts returns; how would you check?”. (ii) Whether the candidate knows the ways a backtest lies (look-ahead, survivorship, overfitting across many trials): “of 200 variants tried, the top one shows a Sharpe ratio of 1.9; what do you conclude?”. (iii) Whether the candidate can reason about a model’s assumptions and their failure: “your regression’s coefficient flips sign when you add a variable; why?”. A trading interview asks for a number now; a research interview asks how the candidate would know the number is right.

*What the interviewer is looking for: testing, overfitting and assumptions as the core of research, each tied to a concrete question.*

**Interview question 1.3 ★ trader, researcher, developer • any.**

Your applications pass the online assessment one time in four, the screen one time in two and the final round two times in five, independently. What is the chance that one application ends in an offer? How many applications do you expect to make before the first offer, and what is the chance of at least one offer from twelve? How many would you need for a 90% chance?

**Solution of Interview question 1.3.**

The offer probability per application is $\tfrac14 \times \tfrac12 \times \tfrac25 = \tfrac1{20}$. The number of applications up to and including the first offer is geometric with mean 20. From twelve applications, $1 - 0.95^{12} \approx 0.46$. A 90% chance needs $n \ge \ln 0.1 / \ln 0.95 \approx 44.9$, so 45 applications. The independence assumption flatters the last figure: the same weakness fails several firms’ assessments, so improving the first stage (one in four) is worth more than applying more widely.

*What the interviewer is looking for: the product of stage rates, the geometric mean, the complement, and a word on independence.*

**Interview question 1.4 ★★ developer • proprietary firm.**

A coding interview is scored on four rows, each from 1 to 4: correctness, complexity, testing and communication. You wrote an $O(n \log n)$ solution to a problem with an $O(n)$ answer, found your own off-by-one error with a test you wrote, and explained as you went. Where does your answer sit on each row, and which single change would have raised the total most?

**Solution of Interview question 1.4.**

Correctness 3 or 4 (the final code is right; the bug was yours but you found it), complexity 2 or 3 (a working but not optimal bound; say that you know an $O(n)$ answer may exist and what it would need), testing 4 (you wrote the test that found the bug), communication 4. The single largest gain is on complexity: before coding, ask what the lower bound is (every element must be read, so $O(n)$), and look for the structure (a hash map, two pointers, a counting array) that reaches it. Five minutes spent there is worth more than the same five minutes spent polishing.

*What the interviewer is looking for: self-assessment against the rubric’s rows, and the habit of asking for the lower bound before coding.*

**Interview question 1.5 ★★ researcher, risk • any.**

A candidate’s ability is $\theta$; each of four interviews scores $\theta + Z$ with independent standard normal $Z$ and is passed when the score is positive. For $\theta = 0.5$, what is the chance of passing all four? Of passing at least three? What ability does a candidate need to pass all four with probability one half, and what does this say about loops that require every interviewer’s approval?

**Solution of Interview question 1.5.**

Each interview is passed with probability $\Phi(0.5) \approx 0.691$. All four: $0.691^4 \approx 0.229$. At least three: $4p^3(1-p) + p^4 \approx 0.637$. Passing all four with probability one half needs $\Phi(\theta) = 0.5^{1/4} \approx 0.841$, so $\theta \approx 1.00$: a candidate one standard deviation of noise above the bar is rejected half the time. A loop of vetoes rejects good candidates often and adds noise rather than removing it; averaging scores (or requiring three of four) keeps the information of every interview.

*What the interviewer is looking for: the binomial calculation, and the insight that a veto rule multiplies noise.*

**Interview question 1.6 ★★ bank • bank.**

The same bank interviews for a desk strategist and a model validator on the same derivatives desk. Give two questions that would appear in one loop and not the other, and say why.

**Solution of Interview question 1.6.**

A desk strategist’s loop asks “the desk wants a price for this new barrier variation by tomorrow; what do you build tonight and what do you tell the trader it is wrong about?” and “hedge this book’s vega with three listed options”. A model validator’s loop asks “this local-volatility model reprices the calibration set perfectly; what tests would still make you reject it?” and “design a benchmark model for an autocallable”. The strategist is paid to deliver a usable number fast and to know its limits; the validator is paid to find where a model fails and to be independent of the desk (One Quant Book 17, chapter 18).

*What the interviewer is looking for: the difference between building for the desk and challenging the desk’s models.*

**Interview question 1.7 ★★★ trader, risk • market maker.**

Abilities in a pool of candidates have standard deviation 1; each interviewer’s score is ability plus independent noise of standard deviation 1. What is the correlation between ability and the average score of one, two and four interviewers? Two interviewers who confer before scoring end up sharing half of their noise; what is the correlation of their average then? What rule for the loop follows?

**Solution of Interview question 1.7.**

With ability spread 1 and noise 1, the average of $n$ independent scores has noise variance $1/n$, so its correlation with ability is $1/\sqrt{1 + 1/n}$: 0.707, 0.816 and 0.894 for one, two and four interviewers. If two interviewers share half their noise (a common component of variance $\tfrac12$ and private parts of variance $\tfrac12$ each), the average’s noise variance is $\tfrac12 + \tfrac14 = 0.75$ and the correlation $1/\sqrt{1.75} \approx 0.756$: conferring first throws away most of the second interviewer. Rule: score independently, record the scores, then discuss.

*What the interviewer is looking for: averaging independent noisy measurements, and why independence is the point.*

**Interview question 1.8 ★★★ mle, researcher • systematic fund.**

You are asked to design a one-hour [technical interview](#def-iv-what-each-interview-tests-by-role-loop) for a machine-learning engineer who will put research models into production at a systematic fund. What do you test, in what order, and what does your rubric record?

**Solution of Interview question 1.8.**

Test what the job does, in the order it matters. (i) Ten minutes: a validation question with time (“your model’s cross-validated score is 0.71 and live is 0.50; list the three most likely causes”), scored on leakage, overlap and non-stationarity. (ii) Twenty-five minutes: code, turning a research function into a production one: remove a look-ahead, add a test, make it streaming with bounded memory. (iii) Fifteen minutes: design, serving a model with a latency budget, a fallback and monitoring for drift. (iv) Five minutes: the candidate’s questions. The rubric records, per part, correctness, whether the candidate tested or checked, and whether the candidate named failure modes before being asked, each on anchored levels written before the first candidate is seen.

*What the interviewer is looking for: a loop derived from the job’s analysis, with anchored scores per dimension.*

Sources and further reading

- P. R. Sackett, C. Zhang, C. M. Berry and F. Lievens, “Revisiting meta-analytic estimates of validity in personnel selection”, *Journal of Applied Psychology* 107(11), 2022, 2040–2068.
- F. L. Schmidt and J. E. Hunter, “The validity and utility of selection methods in personnel psychology”, *Psychological Bulletin* 124(2), 1998, 262–274.
- M. A. Campion, D. K. Palmer and J. E. Campion, “A review of structure in the selection interview”, *Personnel Psychology* 50(3), 1997, 655–702.
- One Quant Book 16, chapter 10 (structured interviews and the hiring funnel); One Quant Book 17, chapters 16–25 (the roles).
