The Interview Book · Careers
4Online Assessments
A timer in the corner of the screen counts down from eight minutes; the counter reads question 1 of 80; a wrong answer costs a mark and a blank costs nothing. The candidate who answers everything finishes with a lower expected score than the one who skipped the fifth of the questions she was least sure of, and the test was designed so that she would. An online assessment measures a small number of abilities under a scoring rule, and the rule is part of the question.
4.1 Timed arithmetic and the scoring rule
Definition 4.1 (Online assessment, numerical reasoning test)
An online assessment is a standardised test taken remotely, under a time limit, early in a process: arithmetic, numerical or logical reasoning, probability, coding, or a game. A numerical reasoning test is an online assessment of calculation with numbers, fractions and percentages, or of reading quantities from tables and charts, scored for the number of correct answers in the time.
Two things decide the score besides ability: speed, and what the rule does with wrong answers. Eighty items in eight minutes is six seconds an item, which is enough for a well-practised method (Chapter 8) and not enough for long multiplication. When wrong answers are penalised, the rule also decides which items to answer.
Proposition 4.2 (When to answer)
If a correct answer scores , a wrong one and a blank 0, answering an item that you get right with probability is worth in expectation, which is positive exactly when .
With the threshold is one half; with the classical quarter-mark penalty on five-option items it is , which is exactly the chance of a random guess, so that blind guessing is worth nothing and any elimination of an option makes guessing worthwhile (Figure 4.1). The hook’s arithmetic: with 64 items answered at 95% and 16 at 35%, answering everything under is worth , and skipping the sixteen weak items is worth 57.6.
fig_iv_skip.py.Method 4.3 (Sitting a timed numerical test)
- Before starting, read the rules for three things: the penalty for a wrong answer, whether items can be revisited, and whether difficulty rises through the test.
- Set a time per item (the limit divided by the count) and a hard ceiling at about twice that; past the ceiling, answer if above the threshold of Proposition 4.2 and move on.
- If items can be skipped and revisited, take a first pass at every item you can do inside the time per item, then return.
- Practise with the same format, timed, until the methods of Chapter 8 are automatic. Practice raises scores: a meta-analysis of 107 samples found an average gain of a quarter of a standard deviation on retesting (Hausknecht et al., 2007), which is why many firms limit retakes (Box 2.1).
4.2 Sequences, patterns and logic items
Sequence and logic items test whether the candidate can find a rule quickly and check it. Most number sequences in such tests are built from a small set of rules, and trying them in a fixed order is faster than staring.
Method 4.4 (Finding the rule of a sequence)
- Take differences; if they are constant after one or two rounds, the sequence is a polynomial of that degree and the next term follows by adding back.
- Take ratios; a constant ratio, or a ratio plus a constant (), is the next family.
- Split the odd and even positions: two interleaved sequences are common.
- Look for squares, cubes, primes, factorials and sums of the two previous terms.
- Check the rule on every given term before answering.
Example 4.5 (Differences)
For 2, 6, 12, 20, 30 the differences are 4, 6, 8, 10 and the second differences are constant at 2: the next difference is 12 and the next term 42. (The terms are , which is the check.)
Logic items are constraint problems: a few facts about an arrangement, one question about it. The method is elimination: write the positions, place the most constrained item first, and test each remaining case against every constraint. A candidate who writes the five seats on paper answers in a minute a question that takes five in the head (Interview question 4.7).
4.3 Automated coding assessments
Definition 4.6 (Automated coding assessment)
An automated coding assessment is an online assessment in which the candidate writes programs that are run against hidden test cases, with limits on time and memory, and are scored by the share of cases passed.
Hidden tests have a predictable structure: small cases, edge cases (empty input, a single element, all elements equal, the largest and smallest values), and large cases that only an efficient solution finishes in the time limit. The input bounds in the statement say which complexity is expected.
Method 4.7 (Reading the bounds)
- At about simple operations a second: allows ; to asks for or ; allows exponential search.
- Write the edge cases down before coding, then code, then run them.
- Watch integer overflow when products or sums of large values are asked for, and floating-point equality when prices are decimals.
- Submit a correct simple version early if partial credit is given, then improve.
4.4 Probability quizzes, games and the rules of the test
Some assessments are short probability quizzes under a clock, testing the speed of the methods of Chapter 10; others are games that record behaviour (how much risk the candidate takes as a reward grows, how quickly they learn a hidden rule) and score it against what the role needs. A game rarely says what it scores; the useful assumption is that it scores what a good decision-maker would do, which can usually be computed (Interview question 4.8).
Tests must be fair to candidates with disabilities, and candidates may ask for adjustments (extra time, a different format) without that request being held against them.
As of September 2026 — Adjustments in recruitment tests
United Kingdom: section 20 of the Equality Act 2010 imposes a duty to make reasonable adjustments where a provision, criterion or practice puts a disabled person at a substantial disadvantage; Schedule 8 applies it to employers, including their arrangements for deciding to whom to offer employment. United States: under the regulations implementing the Americans with Disabilities Act (29 CFR 1630.11), an employer must select and administer employment tests so that, for an applicant with a disability that impairs sensory, manual or speaking skills, the results reflect the ability the test purports to measure rather than the impairment.
Integrity rules matter as much: assessments forbid help from other people and, increasingly, from AI assistants; a result obtained that way is found out in the next stage, which asks the same kind of question live.
4.5 Worked answers
Example 4.8 (Behind the clock at item 20)
Sixty items in thirty minutes; after item 20 the clock shows twelve minutes gone. The budget was 30 seconds an item, so item 20 should have been done at ten minutes: you are two minutes behind. The remaining 40 items now have 18 minutes, 27 seconds each. The plan that loses least is not to hurry every item by 3 seconds, which raises the error rate everywhere, but to set a cut: any item not clearly on its way after 40 seconds is left blank (or guessed, if the scoring rule makes a guess worth it, Method 4.3), and the time goes to the items that can be finished. Look at the clock at items 30, 40 and 50 again, not after every item.
Example 4.9 (Reading the bounds of a coding item)
“Given up to trade prices (integers) and a target, count the pairs whose prices sum to the target.” The bound decides the method: pair checks will not finish in a one-second limit, so the double loop is only a check. One pass with a hash map of counts of prices seen so far adds, at each price , the count of already seen, then records : time. On the prices 3, 5, 2, 5, 1 with target 7 the pass adds 0, 0, 1 (the 5 before the 2), 1 (the 2 before the second 5) and 0: two pairs, which the double loop confirms. The hidden tests to expect: all prices equal (with target twice that price, pairs, which overflows 32-bit integers at ), no pair, negative prices if the statement allows them, and .
4.6 Question bank
Interview question 4.1 ★ trader • market maker
Six seconds each, no calculator: (a) ; (b) , to four decimals; (c) .
Solution
Solution of Interview question 4.1.
(a) . (b) . (c) (multiply both by 1000). Each is one rewrite: a percentage as a fraction, a common denominator, a shift of the decimal point.
What the interviewer is looking for: a rewrite that turns each item into one easy step.
Interview question 4.2 ★ trader, researcher, developer • any
Give the next term: (a) 3, 8, 15, 24, 35; (b) 2, 3, 5, 9, 17; (c) 1, 4, 2, 8, 3, 12, 4.
Solution
Solution of Interview question 4.2.
(a) 48: differences 5, 7, 9, 11 are odd numbers, so the next is 13 (the terms are ). (b) 33: each term is twice the previous minus one. (c) 16: odd positions 1, 2, 3, 4 and even positions 4, 8, 12, so the next even-position term is 16.
What the interviewer is looking for: the method of differences, ratios and interleaving, tried in order, with a check.
Interview question 4.3 ★ trader • proprietary firm
A test gives for a right answer, for a wrong one and 0 for a blank, on four-option items. Should you guess blind? After eliminating one option? Two?
Solution
Solution of Interview question 4.3.
Answer when . Blind: , worth : skip. One option eliminated: , worth exactly 0: indifferent. Two eliminated: , worth : answer.
What the interviewer is looking for: the break-even confidence of a scoring rule.
Interview question 4.4 ★★ developer • market maker
An automated assessment asks for the length of the longest run of strictly increasing consecutive prices in a list of up to trade prices. Write it, and list the hidden test cases you expect it to face.
Solution
Solution of Interview question 4.4.
One pass keeping the current run and the best: the run grows when a price exceeds the previous one and resets to 1 otherwise; time, memory. Hidden cases: the empty list (answer 0), one price (1), all equal (1, since strictly increasing), all decreasing (1), all increasing (), equal neighbours inside a rise (they break it), and prices (any answer times out). The chapter’s code checks the function against a brute force on 300 random lists.
What the interviewer is looking for: a linear scan, the edge cases named before coding, and the strictness of the inequality.
Interview question 4.5 ★★ trader • market maker
Eighty items in eight minutes, for a wrong answer. The first sixty are easy: you take 5 seconds each at 97% accuracy. The last twenty are hard: 12 seconds each at 70%. Items are answered in order and cannot be revisited. What is your plan and your expected score? What changes if your accuracy on the hard items is only 45%?
Solution
Solution of Interview question 4.5.
The easy items take seconds and are worth each, 56.4 in all. The remaining 180 seconds allow 15 hard items, each worth : 6 more, 62.4 in expectation. At 45% a hard item is worth , so answer none of them and score 56.4; if an item looks easier than its neighbours, answer it if you would give it better than even odds.
What the interviewer is looking for: a time budget combined with the expected value of each answer.
Interview question 4.6 ★★ researcher, trader • any
Fifteen seconds each: (a) two fair dice; the chance their sum is prime? (b) three fair coins; the chance of at least two heads?
Solution
Solution of Interview question 4.6.
(a) The prime sums 2, 3, 5, 7, 11 occur 1, 2, 4, 6, 2 times out of 36: . (b) By symmetry, at least two heads out of three has probability (it is the event that heads are the majority).
What the interviewer is looking for: counting by the number of ways, and a symmetry argument where one exists.
Interview question 4.7 ★★ trader, researcher, developer • any
Five traders A, B, C, D, E sit in a row of five seats. C sits at the left end. A sits at neither end and not next to B. D sits immediately to the left of E. B sits somewhere to the right of A. Who sits in the middle?
Solution
Solution of Interview question 4.7.
Seat 1 is C. D and E occupy two adjacent seats, D on the left. A is in seat 2, 3 or 4 and B is to A’s right but not next to A. If A is in seat 2, B is in seat 4 or 5; with D and E adjacent in the remaining seats this forces A, D, E, B in seats 2 to 5 (B in 4 would leave seats 3 and 5 for D and E, which are not adjacent). If A is in seat 3, B must be in seat 5 and D, E in seats 2 and 4, not adjacent: impossible. If A is in seat 4, B has no seat. The row is C, A, D, E, B and D is in the middle; an exhaustive search over the 120 orders finds this row alone.
What the interviewer is looking for: placing the most constrained people first and eliminating cases on paper.
Interview question 4.8 ★★★ trader, risk • market maker
A game shows a balloon. Each pump adds one point to a pot; the balloon bursts at a point drawn uniformly from 1 to 64 pumps, and then the pot is lost. You choose in advance how many pumps to make, then bank. How many maximise the expected pot, and what is it? What might the game be scoring besides the pot?
Solution
Solution of Interview question 4.8.
With pumps the pot is if the burst point exceeds , which happens with probability , so the expected pot is , maximised at with value 16. Stopping at 20 gives 13.75. The game probably also scores consistency (the same choice in the same situation), learning (adjusting when bursts come earlier than expected) and the gap between the candidate’s choice and the optimum, in both directions: too little risk is as much a finding as too much.
What the interviewer is looking for: an expected-value optimisation, and a view of what a behavioural game measures.
Interview question 4.9 ★★★ developer • proprietary firm
Given up to trade prices and up to queries, each asking how many trades had a price in , answer all queries within a one-second limit. What is your approach and its complexity, and which edge cases must it handle?
Solution
Solution of Interview question 4.9.
Sort the prices once, ; answer each query with two binary searches, the first index not less than and the first index greater than , and return their difference: per query, about steps in all, well inside the limit. A linear scan per query would be and time out. Edge cases: (answer 0), bounds equal to existing prices (the inclusive ends), duplicates, a window outside all prices, and an empty list. The chapter’s code checks it against a direct count on random data.
What the interviewer is looking for: preprocessing plus binary search, derived from the bounds, with inclusive-end care.
Interview question 4.10 ★★★ researcher, trader • any
A test’s score is your ability plus independent noise with standard deviation 5; the pass mark is 60 and your ability is 57. What is your chance of passing one sitting? Two independent sittings? If practice raises your ability by a quarter of the spread of candidates’ abilities, which is 8, what is your chance at one sitting then? What does this suggest about when to take a test that cannot be retaken for months?
Solution
Solution of Interview question 4.10.
One sitting: . Two independent sittings: . With practice, ability rises by to 59: at one sitting. Practice before the first attempt is worth almost as much as a second attempt, and a second attempt is often not available for months (Box 2.1): take the test when practised, not when first invited, if the process allows a few days.
What the interviewer is looking for: normal tail probabilities and the practical conclusion about timing.
Sources and further reading
- J. P. Hausknecht, J. A. Halpert, N. T. Di Paolo and M. O. Moriarty Gerrard, “Retesting in selection: a meta-analysis of coaching and practice effects for tests of cognitive ability”, Journal of Applied Psychology 92(2), 2007, 373–385.
- Equality Act 2010, section 20 and Schedule 8; 29 CFR 1630.11.
- T. H. Cormen, C. E. Leiserson, R. L. Rivest and C. Stein, Introduction to Algorithms, 4th edition, MIT Press, 2022, for the complexity rules of thumb.