---
title: "The Research Process"
book: "Research Craft: Predictors, Backtests, Measurement, Portfolios"
subject: quant
language: en
chapter: 1
exercises: 8
source: https://one-course.com/books/quant/7/en/chapter/1-the-research-process
---

# Chapter 1 — The Research Process

A researcher shows a backtest on Friday afternoon: a reversal signal on liquid stocks, two years of daily data, an annual Sharpe ratio of 2.4 after costs. On Monday the head of research asks two questions. How many variants were tried before this one? And which came first, the reason the signal should work or the chart that says it does? Nobody wrote either down, and without them the number on the screen cannot be judged. This book is about the craft that turns market data into strategies a firm can trust: predictors and how to measure them, forecasts, backtests at four levels of fidelity, performance and risk measurement, and the construction of portfolios. It starts with the process, because every later tool can be defeated by a process that does not record what it tried.

## 1.1 Hypothesis first

**Definition 1.1 (Research hypothesis, economic rationale).**

A *research hypothesis* is a falsifiable statement, written before the test, about a mechanism that moves prices: which instruments, at which horizon, which observable predicts what, and which result of the test would refute it. Its *economic rationale* is the answer to the question “who pays you, and why?”: a risk the strategy bears and others will pay to shed, a service it provides (liquidity, immediacy, the warehousing of inventory), a constraint that forces others to trade at a bad price (an index rebalance, a margin call, a mandate), or information processed faster or better than the price already reflects.

The rationale matters for two reasons. It tells the researcher where to look, and so shrinks the search; and it predicts more than the headline result, which gives the hypothesis ways to fail. A liquidity-provision story for short-horizon reversal predicts that the reversal is stronger in less liquid stocks, stronger on days when liquidity is scarce, and weaker after the market’s cost of immediacy falls. A strategy whose backtest succeeds while each side prediction fails is probably fitting something else.

**Example 1.2 (A hypothesis written down).**

*“Stocks that fall more than their sector over one day, on no news, recover part of the move over the next five days, because the sellers are demanding immediacy and market makers are paid for supplying it. The effect should be larger in the bottom half of the universe by traded value and in the top decile of market-wide volatility, and absent on days with an earnings release. If the five-day information coefficient of the sector-relative one-day return is not negative in both halves of the sample, the hypothesis is rejected.”* Every clause of it can be tested, and the last sentence says in advance what failure looks like.

The opposite case is instructive. Arnott, Harvey and Markowitz (2019) open their backtesting protocol with a long–short strategy on NYSE stocks, developed on 1963–1988 and “validated out of sample with even stronger results” over 1989–2015: nearly 6% alpha a year, no significant correlation with the market or the well-known factors, turnover below 10% a year. It was built from the letters of the stocks’ ticker symbols, chosen among thousands of letter combinations. An out-of-sample test did not save it, because nothing but the search stood behind it.

How much a successful test says depends on how many of the hypotheses a team tests are true, an uncomfortable number that Ioannidis (2005) made the centre of an argument about published research.

**Proposition 1.3 (What a significant backtest is worth).**

Suppose a fraction $\pi$ of the hypotheses a team tests are true, and each is tested by $k$ independent stages, every one with size $\alpha$ and power $1 - \beta$ against the true ones. The probability that a hypothesis that passes all $k$ stages is true (its positive predictive value) is

$$
\mathrm{PPV}_k = \frac{\pi(1 - \beta)^k}{\pi(1 - \beta)^k + (1 - \pi)\alpha^k}.
$$

**Proof.** By Bayes’ formula, with $\P(\text{pass} \mid \text{true}) = (1 - \beta)^k$ and $\P(\text{pass} \mid \text{false}) = \alpha^k$ by independence of the stages. ∎

With one stage, a prior of 5%, the conventional size of 5% and a power of one half (a modest edge on a few years of data), $\mathrm{PPV}_1 = 0.345$: two backtests in three that pass are false. A second independent stage (a holdout, another market) lifts it to 0.840 and a third to 0.981 ([Figure 1.1](#fig-rs-the-research-process-ppv)). A better test helps less than a better prior: with a prior of 10% and a power of 80%, one stage already gives 0.64.

![Probability that an idea that passed k independent stages is real, against the share of real ideas among those tested, each stage with 5% size and 50% power. At a prior of 5% (vertical line): 0.345, 0.840 and 0.981. Data: the chapter’s module, .](https://one-course.com/images/onecourse/chapters/quant-7/rs-the-research-process/fig-16538676dc03.svg)

***Figure 1.1.** Probability that an idea that passed $k$ independent stages is real, against the share of real ideas among those tested, each stage with 5% size and 50% power. At a prior of 5% (vertical line): 0.345, 0.840 and 0.981. Data: the chapter’s module, [Proposition 1.3](#prop-rs-the-research-process-ppv).*

## 1.2 The research log and the trial count

**Definition 1.4 (Research log, trial count).**

A *research log* is an append-only record, kept as the work is done, of every hypothesis registered and every test run: its parameters, the code and data it used, its results and when it ran. The *trial count* of a result is the number of tests of the same hypothesis family recorded in the log before the result was chosen, including the variants that failed and were discarded.

The [trial count](#def-rs-the-research-process-log) is what Book 4, chapter 12, needs in order to correct a result for the search behind it: its multiple-testing corrections and its deflated Sharpe ratio both take the number of trials as an input, and a researcher who does not record it cannot supply it. The trials that are hardest to count are the ones never run, only considered: the garden of forking paths of that chapter. The log does not catch those, but it catches everything else, and it makes the difference visible.

**Proposition 1.5 (The best of many worthless variants).**

Let $n$ variants have no edge, each backtested on $Y$ years of daily returns, so that each estimated annual Sharpe ratio is approximately $\mathcal N(0, 1/Y)$, and let the estimates be equicorrelated with correlation $\rho$. Then

$$
\P\Bigl(\max_i \widehat{\mathrm{SR}}_i \ge c\Bigr) = \E_W\Bigl[1 - \Phi\Bigl(\frac{c\sqrt Y - \sqrt\rho\,W}{\sqrt{1 - \rho}}\Bigr)^n\Bigr],
\qquad W \sim \mathcal N(0, 1),
$$

which is $1 - \Phi(c\sqrt Y)^n$ for independent variants.

**Proof.** Write the standardised estimates as $Z_i = \sqrt\rho\,W + \sqrt{1 - \rho}\,\varepsilon_i$ with independent standard normals; given $W$ the $Z_i$ are independent, so $\P(\max_i Z_i < c\sqrt Y \mid W)$ is the product of the $n$ conditional probabilities. Integrate over $W$. ∎

The numbers are stark ([Figure 1.2](#fig-rs-the-research-process-best)). A single worthless strategy backtested on two years shows a Sharpe ratio of 2 or more with probability 0.23%. The best of twenty independent ones does so 4.6% of the time, and the best of a thousand 90% of the time: the expected best of a thousand is 2.30, above the ratio that impressed the room on Friday. One year of data makes a Sharpe ratio of 2 a near certainty for the best of a thousand; five years bring the chance down to 0.39%. Correlation between the variants helps, since correlated variants are fewer independent chances: twenty variants correlated at 0.6 reach 2 with probability 2.6%.

![Probability that the best of n strategies with no edge shows an annual Sharpe ratio of at least 2, by the length of the backtest; the dash-dotted curve has the variants’ estimates correlated at 0.6. With two years of data: 0.23% for one, 4.6% for twenty, 90% for a thousand independent variants. Data: the chapter’s module, .](https://one-course.com/images/onecourse/chapters/quant-7/rs-the-research-process/fig-61fd542d8c36.svg)

***Figure 1.2.** Probability that the best of $n$ strategies with no edge shows an annual Sharpe ratio of at least 2, by the length of the backtest; the dash-dotted curve has the variants’ estimates correlated at 0.6. With two years of data: 0.23% for one, 4.6% for twenty, 90% for a thousand independent variants. Data: the chapter’s module, [Proposition 1.5](#prop-rs-the-research-process-best).*

**Remark 1.6 (Combinations count).**

Trials multiply faster than researchers notice. Twenty candidate variables with all their pairwise interactions are not 22 tests but $20 + \binom{20}{2} = 210$; Arnott, Harvey and Markowitz make the same count. A change of universe, of rebalance day or of cost assumption after seeing a result is a trial too, and belongs in the log.

## 1.3 What counts as evidence

**Definition 1.7 (Pre-registration, holdout set).**

*Pre-registration* is the recording of a hypothesis and of its analysis plan (data, sample, variables, test and decision rule) before the outcome is observed. A *holdout set* is a part of the data set aside before the research begins and used once, for the final test of a pre-registered analysis.

Nosek and co-authors (2018) put the case in one distinction: [pre-registration](#def-rs-the-research-process-prereg) separates analyses that test predictions from analyses that generate them after the fact, and a result’s credibility depends on which it is. In finance the distinction is harder to keep than in a clinical trial, because the data exist before the question. A holdout is the way to recover it, provided it is really held out.

**Method 1.8 (Weighing the evidence for a strategy).**

Rank what the research produced, weakest first, and ask for at least three levels:

1. an in-sample fit, however good (it measures the search as much as the idea);
2. the same test on later data that were not used to choose anything;
3. the mechanism’s side predictions: other markets, other periods, the subsets where the rationale says the effect should be stronger or absent;
4. a pre-registered test on the holdout, run once;
5. live trading at small size, measured against the backtest of the same days.

Each level counts only if it was independent of the choices made at the levels before it.

A holdout that has been looked at is not independent. If a researcher checks it five times while tuning, each time at 5%, the chance that a worthless strategy passes at least once is $1 - 0.95^5 = 22.6\%$, not 5%. In [Proposition 1.3](#prop-rs-the-research-process-ppv) the second stage then has size 0.226 instead of 0.05, and at a 5% prior the two stages together give a PPV of 0.54 instead of 0.84. A holdout is spent the moment it influences a choice.

## 1.4 From idea to production

**Definition 1.9 (Stage gate, kill criterion, research review).**

A *stage gate* is a decision point between two stages of a strategy’s life at which it must meet criteria fixed in advance to continue. A *kill criterion* is a rule, written before a stage starts, that stops the strategy when the evidence it produces falls below a threshold. A *research review* is the examination of a result by people who did not produce it, with access to the [research log](#def-rs-the-research-process-log).

![The life of a strategy. Each stage is entered through a gate whose criteria were set before the stage began, and each can end in a stop whose reason is logged: the stops are trials, and the log counts them.](https://one-course.com/images/onecourse/chapters/quant-7/rs-the-research-process/fig-161d824a28c5.svg)

***Figure 1.3.** The life of a strategy. Each stage is entered through a gate whose criteria were set before the stage began, and each can end in a stop whose reason is logged: the stops are trials, and the log counts them.*

The pipeline in [Figure 1.3](#fig-rs-the-research-process-stages) is generic; firms name the stages differently and some merge them. Its logic does not change. Exploration may search freely, because it is logged. The holdout test runs one pre-registered analysis on the holdout. Review reads the log as well as the result. Shadow trading runs the strategy on live data without orders (chapter 19), and small live trading measures what the simulator cannot: fills, impact and the discipline of the people running it. A [kill criterion](#def-rs-the-research-process-gate) needs the same care as the entry criteria, because live records are short and noisy.

**Proposition 1.10 (How much a live record says against its backtest).**

If a strategy’s true annual Sharpe ratio were the backtested $S$, the probability that $Y$ years of daily live returns show an annual Sharpe ratio of $s$ or less is approximately $\Phi\bigl((s - S)\sqrt Y\bigr)$. Telling a true $S$ from zero with a one-sided test of size $\alpha$ and power $1 - \beta$ takes about $Y = \bigl((z_{1-\alpha} + z_{1-\beta})/S\bigr)^2$ years.

**Proof.** The estimate is approximately $\mathcal N(S, 1/Y)$; the second statement is Book 4, chapter 12, with the roles written for a live record. ∎

A strategy backtested at 1.5 that earns a Sharpe ratio of 0.2 in its first live year is disappointing but not decisive: under the backtest’s own claim, a year that bad has probability 9.7%. A live record that must separate 1.5 from zero with 5% size and 80% power needs 2.7 years. Hence the rule of practice: a [kill criterion](#def-rs-the-research-process-gate) is a threshold on a statistic fixed in advance, together with a minimum time, and often with a second trigger on something measured faster than returns, such as fills against the simulator or the realised costs.

## 1.5 Research as a portfolio of bets

A research group is itself a strategy: it spends researcher time and data on ideas with a low prior and a skewed payoff. Its expected output in a year follows from [Proposition 1.3](#prop-rs-the-research-process-ppv). A group that tests 200 ideas a year, of which 5% are real, with stages of 5% size and 50% power, passes 5 real ideas and 9.5 false ones through exploration, and 2.5 real and 0.475 false through a second independent stage. The first gate halves the real ideas and divides the false ones by twenty; that trade is almost always worth making, because a false strategy that reaches production costs its trading losses, the capital it ties up and the months before the [kill criterion](#def-rs-the-research-process-gate) can fire.

Three rules of thumb follow. Raise the prior before raising the power: a hypothesis with a rationale and side predictions is worth more than a larger search. Keep stages independent: a holdout used once, a market not looked at, a period not studied. And count: the [research log](#def-rs-the-research-process-log) is cheap, and it is the only record from which anyone can later tell a discovery from a search.

## 1.6 Tutorial: logging a search

**Goal.** Run a fifty-variant search on data where nothing works, with every trial in the [research log](#def-rs-the-research-process-log), and let the log deflate the winner. **End state:** [Figure 1.4](#fig-rs-the-research-process-search); the winner at a two-day lookback with an annual Sharpe ratio of 2.14, a probabilistic Sharpe ratio of 0.999 as a single test and a deflated Sharpe ratio of 0.77 with the logged fifty trials.

1. **The log.** Each entry’s digest covers its content and the previous digest, so editing any past entry breaks every later link; `verify()` finds the first broken one. `def _digest (kind: str , family: str , ts: str , body: dict , prev: str ) -> str : """SHA-256 of the canonical JSON of an entry and the previous digest: editing any past entry changes its digest and breaks every later link.""" blob = json.dumps([kind, family, ts, body, prev], sort_keys=True , separators=(" , " , " : " )) return hashlib.sha256(blob.encode()).hexdigest() class ResearchLog : def __init__(self , path=None ): self .path = pathlib.Path(path) if path else None self .entries: list [Entry] = [] if self .path and self .path.exists(): for line in self .path.read_text().splitlines(): if line.strip(): self .entries.append(Entry(**json.loads(line))) def _append (self , kind: str , family: str , ts: str , body: dict ) -> Entry: prev = self .entries[-1 ].digest if self .entries else GENESIS e = Entry(kind, family, ts, body, prev, _digest(kind, family, ts, body, prev)) self .entries.append(e) if self .path: with self .path.open(" a " ) as f: f.write(json.dumps(asdict(e), sort_keys=True ) + " \n " ) return e` **Listing 1.1.** Hash-chained entries of the research log. code/firm/researchlog/firm_researchlog.py
2. **The search.** Register the hypothesis first, then log fifty reversal rules (hold minus the sign of the trailing $L$-day return, $L = 1, \dots, 50$) on two years of a driftless random walk with 1% daily volatility. `def search (seed: int = SEED, years: float = 2.0 , log: ResearchLog | None = None ): """The tutorial: fifty reversal rules (hold minus the sign of the trailing L-day return) on a driftless random walk with 1% daily volatility; every trial goes into the research log.""" n = int (round (years * DAYS)) rng = np.random.default_rng(seed) r = rng.normal(0.0 , 0.01 , n + max (LOOKBACKS)) log = log or ResearchLog() log.register(" reversal " , " short-horizon returns revert: liquidity providers are paid to absorb flow " , " 2026-09-24T09:00 " ) cs = np.concatenate([[0.0 ], np.cumsum(r)]) srs = [] for k, lb in enumerate (LOOKBACKS): t = np.arange(max (LOOKBACKS), len (r)) pos = -np.sign(cs[t] - cs[t - lb]) # uses returns up to t - 1 only pnl = pos * r[t] sr = float (pnl.mean() / pnl.std(ddof=1 ) * math.sqrt(DAYS)) srs.append(sr) log.trial(" reversal " , {" lookback " : lb}, {" sr " : round (sr, 6 )}, f " 2026-09-24T10: { k: 02d } " , code_hash=" rs_process.search " , data_id=f " rw-seed { seed} " ) srs = np.array(srs) best = int (np.argmax(srs)) return {" srs " : srs, " best_lookback " : LOOKBACKS[best], " best_sr " : float (srs[best]), " log " : log, " n_obs " : n}` **Listing 1.2.** Fifty logged trials of a reversal rule on noise. code/research/01-the-research-process/python/rs_process.py
3. **Deflate.** Call `log.deflated(family, sr, n_obs)` on the family `reversal` with the winner’s daily Sharpe ratio: the benchmark is the expected best of the logged number of null trials. Then run `fig_process.py` .

**What to change next.** Add a second family (momentum rules) to the same log and check that each family is deflated by its own count; edit one logged metric in the JSON file and watch `verify()` point at it.

![Left: annual Sharpe ratios of the fifty reversal rules on one simulated two-year random walk; the winner (dot) has 2.14 at a two-day lookback, five rules exceed 1 and the mean is 0.07. Right: the winner’s deflated Sharpe ratio against the number of trials the log records: 0.999 for one, 0.77 for the fifty actually run (vertical line). Data: the chapter’s tutorial, seeded (the first history drawn).](https://one-course.com/images/onecourse/chapters/quant-7/rs-the-research-process/fig-8ffdadd1072c.svg)

![Left: annual Sharpe ratios of the fifty reversal rules on one simulated two-year random walk; the winner (dot) has 2.14 at a two-day lookback, five rules exceed 1 and the mean is 0.07. Right: the winner’s deflated Sharpe ratio against the number of trials the log records: 0.999 for one, 0.77 for the fifty actually run (vertical line). Data: the chapter’s tutorial, seeded (the first history drawn).](https://one-course.com/images/onecourse/chapters/quant-7/rs-the-research-process/fig-24bd1270f41b.svg)

***Figure 1.4.** Left: annual Sharpe ratios of the fifty reversal rules on one simulated two-year random walk; the winner (dot) has 2.14 at a two-day lookback, five rules exceed 1 and the mean is 0.07. Right: the winner’s deflated Sharpe ratio against the number of trials the log records: 0.999 for one, 0.77 for the fifty actually run (vertical line). Data: the chapter’s tutorial, seeded (the first history drawn).*

## 1.7 Build: the research log

**Purpose.** The miniature firm’s memory of what its research tried. Every later tool of this book that runs a test (the backtesters of Part IV, the overfitting tools of chapter 20, the workflow of chapter 29) writes a trial here, and the review of a result starts by reading its family.

**Interface.** `ResearchLog(path)`; `register(family, hypothesis, ts)`; `trial(family, params, metrics, ts, code_hash, data_id)` returning an `Entry`; `trial_count(family)`; `post_hoc(family)`; `metric(family, name)`; `verify()`; `deflated(family, sr, n_obs, skew, kurt)`.

**Rules.** Append-only JSON lines; each entry carries the digest of its content and of the previous entry; timestamps are supplied by the caller, never read from the clock, so that a replayed log is identical; one hypothesis per family, and a family whose first entry is a trial is flagged post hoc; the deflation uses Book 4’s `firm.multitest`.

**Acceptance tests.** `code/firm/researchlog/tests/`: an intact chain verifies; an edited metric is located; a post hoc family is flagged and a second hypothesis refused; the log survives a reload and keeps counting; the deflated Sharpe ratio falls as the count grows.

**Stretch.** Sign entries with a per-researcher key; count the effective number of trials of a family from the correlation of its logged P&L series (Book 4, chapter 12) instead of the raw count.

Sources and further reading

- J. P. A. Ioannidis, “Why most published research findings are false”, *PLoS Medicine* 2(8), e124, 2005.
- R. D. Arnott, C. R. Harvey and H. Markowitz, “A backtesting protocol in the era of machine learning”, *Journal of Financial Data Science* 1(1), 2019.
- B. A. Nosek, C. R. Ebersole, A. C. DeHaven and D. T. Mellor, “The preregistration revolution”, *Proceedings of the National Academy of Sciences* 115(11), 2018.
- D. H. Bailey, J. M. Borwein, M. López de Prado and Q. J. Zhu, “Pseudo-mathematics and financial charlatanism: the effects of backtest overfitting on out-of-sample performance”, *Notices of the American Mathematical Society* 61(5), 2014.
- M. López de Prado, “The 10 reasons most machine learning funds fail”, *Journal of Portfolio Management* 44(6), 2018.
- C. R. Harvey, Y. Liu and H. Zhu, “…and the cross-section of expected returns”, *Review of Financial Studies* 29(1), 2016.

## 1.8 Exercises

**Exercise 1.1 ★.**

A team believes 10% of its ideas are real and tests each with 80% power at 5% size. What fraction of the ideas that pass are real? What if the prior is 2%?

**Solution of Exercise 1.1.**

$\mathrm{PPV} = 0.1 \times 0.8/(0.1 \times 0.8 + 0.9 \times 0.05) = 0.08/0.125 = 0.64$. With a 2% prior, $0.016/(0.016 + 0.049) = 0.246$: three passes in four are false.

**Exercise 1.2 ★.**

Twenty independent worthless strategies are each tested one-sidedly at $t \ge 2$. What is the probability that at least one passes, and how many pass on average?

**Solution of Exercise 1.2.**

$1 - \Phi(2)^{20} = 1 - 0.97725^{20} = 0.369$; on average $20 \times 0.02275 = 0.455$ pass.

**Exercise 1.3 ★.**

A strategy backtested at an annual Sharpe ratio of 1.2 earns $-0.5$ over its first six months live. How likely is a half-year that bad if the backtest were right?

**Solution of Exercise 1.3.**

$\Phi\bigl((-0.5 - 1.2)\sqrt{0.5}\bigr) = \Phi(-1.20) = 0.115$: about one half-year in nine would be this bad. Not decisive on its own.

**Exercise 1.4 ★★.**

How many years of live daily returns separate an annual Sharpe ratio of 0.8 from zero with a one-sided test of 5% size and 90% power? Compare with the answer for 1.5 at 80% power in the text.

**Solution of Exercise 1.4.**

$Y = ((1.645 + 1.282)/0.8)^2 = 13.4$ years, against 2.7 years for 1.5 at 80% power: the time grows with the inverse square of the Sharpe ratio.

**Exercise 1.5 ★★.**

A researcher tests thirty variables and all their pairwise products. How many tests is that, and what $t$-statistic does a one-sided Bonferroni correction at 5% require?

**Solution of Exercise 1.5.**

$30 + \binom{30}{2} = 30 + 435 = 465$ tests. Bonferroni needs $p \le 0.05/465 = 1.08 \times 10^{-4}$, that is $t \ge 3.70$.

**Exercise 1.6 ★★.**

A group tests 150 ideas a year with a 4% prior through three independent stages of 5% size and 60% power. How many real and how many false strategies reach production each year, and what is the PPV at each stage?

**Solution of Exercise 1.6.**

Six real and 144 false ideas. After stage one, 3.6 real and 7.2 false (PPV 0.333); after stage two, 2.16 and 0.36 (0.857); after stage three, 1.30 and 0.018 (0.986). About 1.3 real strategies and one false one every 55 years reach production.

**Exercise 1.7 ★★★.**

*Coding.* With `p_best`, find the probability that the best of a thousand worthless variants backtested on two years reaches a Sharpe ratio of 2 when the variants are equicorrelated at 0.3, 0.6 and 0.9. Which correlation makes a thousand variants worth about as much as twenty independent ones?

**Solution of Exercise 1.7.**

`p_best(2, 1000, 2, rho)` gives 0.420, 0.167 and 0.030 for $\rho = 0.3$, 0.6 and 0.9. Twenty independent variants give 0.046; a thousand variants equicorrelated at $\rho = 0.85$ give the same. Even at 0.9 a thousand variants are worth about thirteen independent ones: correlation reduces a search but rarely to one trial.

**Exercise 1.8 ★★★.**

*Find the flaw.* “We found the signal on 2015–2020 and confirmed it on the 2021–2023 holdout with a Sharpe ratio of 1.4. We then tuned two parameters on the holdout, which raised it to 2.1. The strategy has passed two independent stages and we propose to trade it.”

**Solution of Exercise 1.8.**

The holdout stopped being a holdout when it was used to tune: the 2.1 is an in-sample number, and the only independent evidence is the 1.4 measured before tuning, itself one look at a holdout that was then reused. Two parameters tuned on three years are also trials the log must count. The honest status is one exploratory stage passed; a new, untouched period (or live shadow trading) is needed before the claim “two independent stages” is true.

## 1.9 Problem: The Friday Backtest

**Problem 1.1.**

Weekend problem — how often does a worthless search produce a Sharpe ratio of 2?

A researcher with no edge tries twenty variants of a signal each week, fifty weeks a year. Each variant is backtested on two years of daily data, the twenty variants of a week are equicorrelated at 0.6, and different weeks’ searches are independent. A backtest “impresses” at an annual Sharpe ratio of 2.

**Part I — One backtest.**

1. What is the law of one variant’s estimated annual Sharpe ratio?
2. What is the probability that one variant impresses?
3. What is the expected number of impressive variants in a year, ignoring correlation?
4. Why does that expectation not answer the question the head of research asks?

**Part II — A week, a year.**

5. What is the probability that the best of one week’s twenty variants impresses if they were independent?
6. And with their correlation of 0.6?
7. What is the probability of at least one impressive backtest in a year?
8. What would it be if all thousand variants were independent?
9. What is the expected best Sharpe ratio of a thousand independent worthless variants?

**Part III — What the log changes.**

10. What is the probabilistic Sharpe ratio of a two-year backtest at 2.0 read as a single test?
11. What is its deflated Sharpe ratio if the log shows the week’s twenty trials?
12. And if it shows the year’s thousand?
13. Why is the thousand the right count, although the variants were correlated?
14. How many years of backtest would bring the single-variant probability of impressing below 0.001%?

**Part IV — The gate.**

15. The impressive variant is sent to a one-year holdout that it passes if its Sharpe ratio exceeds 1. What is the probability that a worthless variant passes?
16. With a 5% prior and the holdout’s power against a real Sharpe ratio of 1.5, what is the PPV after the holdout?
17. The researcher checked the holdout four times while tuning. What is the holdout’s effective size?
18. State the *named result* : the probability of at least one impressive backtest in a year of this search, and the deflated Sharpe ratio of that backtest with the logged count.
19. What three things should the Friday presentation have shown?
20. In one sentence: why is the [research log](#def-rs-the-research-process-log) a risk-management tool?

**Solution of Problem 1.1.**

**1.** Approximately $\mathcal N(0, 1/2)$: standard deviation 0.707. **2.** $1 - \Phi(2\sqrt2) = 1 - \Phi(2.83) = 0.00234$. **3.** $1\,000 \times 0.00234 = 2.34$. **4.** The question is the probability that at least one impresses (and which one is shown); correlation changes that probability, not the expectation. **5.** $1 - (1 - 0.00234)^{20} = 0.0458$. **6.** By [Proposition 1.5](#prop-rs-the-research-process-best) with $\rho = 0.6$: 0.0262. **7.** $1 - (1 - 0.0262)^{50} = 0.735$. **8.** $1 - \Phi(2.83)^{1\,000} = 0.904$. **9.** 2.30, above the impressive threshold. **10.** $\Phi\bigl(2/\sqrt{252} \times \sqrt{503}\bigr) = 0.998$. **11.** With the benchmark the expected best of twenty null ratios: 0.822. **12.** With a thousand: 0.336, a coin flip that the strategy is better than the best of the search. **13.** The count must cover every trial that could have been shown; correlation reduces the effective number, and the deflated ratio with the raw count is then conservative, but the effective count needs the variants’ correlations, which only the log can provide. **14.** $2\sqrt Y \ge z_{0.99999} = 4.265$, so $Y \ge 4.55$ years. **15.** $1 - \Phi(1) = 0.159$. **16.** Power $\Phi(1.5 - 1) = 0.691$; $\mathrm{PPV} = 0.05 \times 0.691/(0.05 \times 0.691 + 0.95
\times 0.159) = 0.187$: a one-year holdout at this threshold is a weak stage. **17.** $1 - (1 - 0.159)^4 = 0.499$: a coin flip. **18.** *Named result:* a worthless search of twenty correlated variants a week produces at least one backtest with a Sharpe ratio of 2 within a year with probability 0.735, and that backtest’s deflated Sharpe ratio with the logged thousand trials is 0.336. **19.** The hypothesis and the date it was registered; the [trial count](#def-rs-the-research-process-log) of the family and the distribution of all trials’ results; the evidence beyond the in-sample fit (later data, side predictions, a holdout). **20.** It records the search, and the search is the largest unmodelled risk in a backtest.

## 1.10 Interview questions

**Interview question 1.1 ★ researcher, trader.**

A candidate strategy shows an annual Sharpe ratio of 2.5 in its backtest. What are the first three questions you ask?

**Solution of Interview question 1.1.**

How many variants were tried, and how correlated were they (the [trial count](#def-rs-the-research-process-log), and so the deflated ratio)? What is the [economic rationale](#def-rs-the-research-process-hypothesis), and was it written before the test? What evidence exists beyond the in-sample fit (later data, other markets, side predictions), and are the costs, borrow and capacity realistic?

*What the interviewer is looking for: the search before the statistics; a rationale with testable side predictions; costs.*

**Interview question 1.2 ★★ researcher.**

Most of the ideas a team tests are false. Why does that matter more than the significance level of its tests?

**Solution of Interview question 1.2.**

By Bayes, the share of true strategies among those that pass is $\pi(1 - \beta)/(\pi(1 - \beta) + (1 - \pi)\alpha)$. With a 5% prior, 5% size and 50% power it is 0.345: most passes are false whatever the size, unless the prior rises or independent stages are added. Size controls the false positives per false idea tested, not the proportion of false strategies among the winners.

*What the interviewer is looking for: the positive predictive value, and that priors and stages dominate.*

**Interview question 1.3 ★★ researcher, mle.**

What makes a [holdout set](#def-rs-the-research-process-prereg) stop being a holdout?

**Solution of Interview question 1.3.**

Any use of it that influences a choice: tuning on it, looking at it and going back to change the model, re-running it after a disappointing result, or choosing the sample period with knowledge of it. Each look raises the holdout’s effective size (five looks at 5% give 22.6%).

*What the interviewer is looking for: that a holdout is spent by influence, not by being read once.*

**Interview question 1.4 ★★ trader, researcher, risk.**

A strategy with a backtested Sharpe ratio of 1.5 is flat after one live year. Do you turn it off?

**Solution of Interview question 1.4.**

Not on that evidence alone: if the true ratio were 1.5, a year at zero or below has probability $\Phi(-1.5) = 0.067$. Look at what is measured faster than returns: are fills, costs and turnover what the simulator said? Is the signal’s live information coefficient in line with the backtest? Apply the [kill criterion](#def-rs-the-research-process-gate) fixed before going live; if none was fixed, fix one now and include a minimum time.

*What the interviewer is looking for: the noise in a one-year Sharpe ratio, and diagnostics beyond P&L.*

**Interview question 1.5 ★★ developer, researcher.**

Design a [research log](#def-rs-the-research-process-log). What does each entry hold, and how do you make it tamper-evident?

**Solution of Interview question 1.5.**

Per entry: family, kind (hypothesis or trial), caller-supplied timestamp, parameters, metrics, code version, data snapshot and the digest of the previous entry, with the entry’s own digest a hash of all of these. Append-only storage; verification walks the chain. One hypothesis per family registered before any trial; trials before it flag the family post hoc.

*What the interviewer is looking for: hash chaining, reproducible timestamps, and counting trials per hypothesis family.*

**Interview question 1.6 ★★★ researcher.**

Give an [economic rationale](#def-rs-the-research-process-hypothesis) for short-term reversal in equities, and three predictions it makes besides the reversal itself.

**Solution of Interview question 1.6.**

Sellers or buyers demanding immediacy push prices away from value, and liquidity providers who absorb the flow are paid by the subsequent reversal. Predictions: the reversal is larger in less liquid stocks; larger when market-wide liquidity is scarce (high volatility, dealer balance sheets constrained); smaller or absent when the move carries news (earnings days), and shrinking as the market’s cost of immediacy falls over the years.

*What the interviewer is looking for: a rationale that says who pays, and side predictions that could refute it.*
