Research Craft: Predictors, Backtests, Measurement, Portfolios · Research
29Build a Research Workflow
A result from eighteen months ago cannot be reproduced. The vendor has revised its history since, the code has moved on, and nobody kept the random seed of the bootstrap that produced the confidence interval in the committee’s slides. Each of the three is ordinary, and together they make the result unknowable: it may have been right, but nobody can say. This last chapter of the book puts its reference study, from a data snapshot to a tear sheet, into a pipeline whose every stage is keyed by the hash of its code, parameters, seed and inputs, so that a result is a manifest that can be re-run bit for bit, and a change anywhere says exactly what it changed. The build is firm.workflow.
29.1 Reproducibility
Definition 29.1 (Reproducible result, data snapshot, content-addressed storage, pipeline stage, environment lock)
A reproducible result is one that the same data, code and settings produce again, bit for bit (Book 4, chapter 25), a minimum standard when an independent replication is not possible (Peng). A data snapshot is an immutable copy of the data as used, identified by its content. Content-addressed storage stores each object under the hash of its content, so that the same content has the same address and different content cannot share one (Git’s object store is the familiar example). A pipeline stage is a function of its inputs, parameters and seed, with no other source of variation. An environment lock records the exact versions of the interpreter and libraries a result was produced with.
Sandve and co-authors reduced the practice to ten rules; the first, sixth and ninth carry most of it: for every result, keep track of how it was produced; for analyses that include randomness, note the seeds; connect textual statements to the results behind them. The book has applied the ninth rule to itself since chapter 1: every number printed in these chapters is asserted by a test that recomputes it. What a research desk adds is scale: many studies, shared stages, vendor data that changes under them. The answer is to make the pipeline, not the analyst, remember how each result was produced.
firm.workflow keys each stage by the SHA-256 of its name, the hash of its source code, its parameters, its seed and the content hashes of its inputs’ outputs, stores the output under that key, and hashes the output in turn (Listing 29.1). A stage whose key is already stored does not run. Because a key depends on its inputs’ outputs rather than their keys, a change upstream that leaves an intermediate result unchanged stops there. A run writes a manifest: for each stage its key, code hash, parameters, seed, input hashes, output hash and whether it ran, and the environment.
29.2 Review
Definition 29.2 (Review checklist)
A review checklist is the list of questions a research review (chapter 1) answers before a result is accepted: that it is reproduced from its manifest, that its data are point in time, that its trials are counted, that its fidelity level matches the claim, and that its costs and capacity are those of the size it will run at.
The book’s chapters are its checklist, and a reviewer can ask their questions of any manifest.
- Reproduced. Does the manifest reproduce bit for bit in an empty cache, in the locked environment? Which stages have no seed?
- Point in time. Is the snapshot identified, and are its revisions recorded (chapter 3)? Are the universe and the exposures known at each date (chapters 4 and 24)?
- Counted. How many trials are in the research log’s family, and what are the deflated Sharpe ratio and the probability of backtest overfitting (chapters 1 and 20)?
- Faithful. At which fidelity level was it backtested, and is the claim within what that level can say (chapters 16 to 19)?
- Costed and sized. Are the costs those of the fitted model at the intended size, and where is the capacity (chapters 27 and 28)?
- Stated. Is every number in the write-up traceable to a stage’s output, and every chart to its data (Sandve’s ninth and seventh rules)?
Arnott, Harvey and Markowitz set out a similar protocol for research with machine learning; the pipeline makes its first questions mechanical, so that the review can spend its time on the last ones.
29.3 Build: the pipeline from raw data to a reviewed result
The reference study has eight stages (Figure 29.1, Listing 29.3): the snapshot of firm.synthmkt with any revisions the vendor has made; the 500 most liquid names each day; two features, 12–1 month momentum and the one-day reversal; their predictor-card statistics; an equal blend; a monthly dollar-neutral book; a level-1 backtest with costs of ten basis points; and a tear sheet with a block bootstrap of the Sharpe ratio, seeded. The first run computes every stage in a few seconds; the second reads everything from the cache.
The study’s own result is modest, and the pipeline reports it with its uncertainty: momentum’s mean daily rank information coefficient is 0.0076 (), the reversal’s 0.0379 () over 2 267 days, and the blended monthly book has a Sharpe ratio after costs of 0.08 (standard error 0.33; bootstrap interval from to 0.74), turning over 3.5% of its capital a day. The reversal’s strong daily signal does not survive a book rebuilt monthly, the horizon mismatch of chapter 6. What matters here is how the result behaves under change:
| change | stages re-run | outputs changed |
|---|---|---|
| none | none | none |
| revision, name never in the universe | snapshot, universe, features, cards, level 1 | snapshot |
| revision, name always in the universe | all eight | all eight |
| blend weights 0.7 and 0.3 | blend, portfolio, level 1, tear sheet | those four |
| tear sheet code gains a hit rate | tear sheet | tear sheet |
A revision to a name that is never among the liquid 500 re-runs every stage that reads the snapshot, and each of them reproduces its previous output, so nothing downstream of them runs: the manifest diff shows the snapshot’s parameters and output changed and the others’ inputs only. A revision inside the universe changes everything, and says so. A change of code changes only its stage. And reproduction, re-running the whole manifest in an empty cache, returns the same output hash for all eight stages; with the bootstrap’s seed left out it flags exactly one, the tear sheet (Listing 29.2). Finally the run is registered as a trial in the research log of chapter 1, with the hash of its manifest as its code identity and the snapshot’s hash as its data identity: the eighteen-month-old result, had it been produced this way, would be one line of the log and one command.
29.4 Tutorial: the result that could not be reproduced
Goal. Run the reference study through firm.workflow, change the data, the parameters and the code in turn, reproduce the result bit for bit, and register it. End state: the table and Figure 29.1.
Running a pipeline: keys from code, parameters, seed and input hashes; the cache; the manifest.
def run(self, target=None): outputs, out_hash, manifest = {}, {}, {"stages": {}, "env": environment()} for n in self._order(target): s = self.stages[n] code = s.code_hash() key = content_hash([n, code, s.params, s.seed, [out_hash[d] for d in s.deps]]) path = self.cache / f"{key}.pkl" ran = not path.exists() if ran: out = s.fn({d: outputs[d] for d in s.deps}, s.params, s.seed) path.write_bytes(pickle.dumps(out, protocol=4)) else: out = pickle.loads(path.read_bytes()) outputs[n], out_hash[n], self.keys[n] = out, content_hash(out), key manifest["stages"][n] = {"key": key, "code": code, "params": s.params, "seed": s.seed, "deps": list(s.deps), "inputs": [out_hash[d] for d in s.deps], "output": out_hash[n], "ran": ran} return outputs, manifestListing 29.1. A content-addressed pipeline run. code/firm/workflow/firm_workflow.py Reproduction and diff.
def reproduce(make_pipeline, manifest: dict, cache_dir) -> list: """Re-run the pipeline built by make_pipeline() in an empty cache_dir; the stages whose outputs differ.""" _, again = make_pipeline(cache_dir).run() return [n for n, m in manifest["stages"].items() if again["stages"][n]["output"] != m["output"]] def diff(m1: dict, m2: dict) -> dict: out = {} for n in m1["stages"]: a, b = m1["stages"][n], m2["stages"].get(n) if b is None: out[n] = ["removed"] continue r = [k for k in ("code", "params", "seed", "inputs", "output") if content_hash(a[k]) != content_hash(b[k])] if r: out[n] = r return outListing 29.2. Re-running from scratch, and what changed between two manifests. code/firm/workflow/firm_workflow.py The reference study as eight stages.
def make(revisions=(), weights=None, seed=7, reps=200, tear=tearsheet): weights = weights or {"momentum": 0.5, "reversal": 0.5} def build(cache): return Pipeline([ Stage("snapshot", snapshot, (), {"market_seed": 1, "revisions": [list(r) for r in revisions]}), Stage("universe", universe, ("snapshot",), {"top": 500, "window": 21}), Stage("features", features, ("snapshot", "universe")), Stage("cards", cards, ("snapshot", "features")), Stage("blend", blend, ("features", "universe"), {"weights": weights}), Stage("portfolio", portfolio, ("blend", "universe"), {"gross": 1.0, "rebalance": MONTH}), Stage("level1", level1, ("portfolio", "snapshot"), {"lag": 1, "cost": 0.0010}), Stage("tearsheet", tear, ("level1",), {"reps": reps, "block": MONTH}, seed), ], cache) return buildListing 29.3. The study’s pipeline. code/research/29-build-a-research-workflow/python/rs_workflow.py - Run
rs_workflow.make()in a cache directory; revise a name withnever_in_universeandalways_in_universe; change the weights; swap intearsheet_v2; callreproduceandregister.
What to change next. Add a level-2 stage (chapter 17) on the book’s largest name and see which changes reach it; store the cache in a shared directory and let a colleague reproduce your manifest on another machine, comparing environments first.
29.5 Build: the workflow
Purpose. Every result the firm relies on is produced by a pipeline whose manifest reproduces it, says what it depends on, and is registered in the research log.
Interface. content_hash(obj), Stage(name, fn, deps, params, seed), Pipeline(stages, cache_dir).run(target) returning outputs and a manifest, environment(), reproduce(make_pipeline, manifest, cache_dir), diff(m1, m2), register(manifest, log, family, ts, metrics).
Rules. Stages are pure functions of inputs, parameters and seed; data enter only through snapshot stages whose revisions are parameters; every stage that draws random numbers takes a seed; manifests and caches are kept with the result; a result is reviewed from its manifest, not from its slides.
Acceptance tests. code/firm/workflow/tests/: a canonical hash (key order irrelevant, types significant); a second run entirely from the cache; early cutoff when a revision is clipped away; diffs naming the changed fields; reproduction bit for bit, and a failed reproduction flagging exactly the unseeded stage; registration in the research log.
Stretch. Code hashes that follow the functions a stage calls; remote caches shared by the desk; environment locks enforced before reproduction.
Sources and further reading
- R. D. Peng, “Reproducible research in computational science”, Science 334(6060), 2011.
- G. K. Sandve, A. Nekrutenko, J. Taylor and E. Hovig, “Ten simple rules for reproducible computational research”, PLoS Computational Biology 9(10), 2013.
- Pro Git, 2nd edition, section 10.2, “Git internals: Git objects” (git-scm.com).
- R. Arnott, C. R. Harvey and H. Markowitz, “A backtesting protocol in the era of machine learning”, Journal of Financial Data Science 1(1), 2019.
29.6 Exercises
Exercise 29.1 ★
Why must a stage’s key include the hash of its code, and what does a key that includes only its parameters miss?
Solution
Solution of Exercise 29.1.
The same parameters run through different code give different results; a key without the code’s hash would serve the old result after the code changed (the tear sheet that gained a hit rate would have been read from the cache without it).
Exercise 29.2 ★
Why are keys built from the inputs’ output hashes rather than the inputs’ keys?
Solution
Solution of Exercise 29.2.
So that a change upstream that leaves an intermediate output unchanged stops there: the downstream keys depend only on what the inputs contain, not on how they were produced (early cutoff).
Exercise 29.3 ★
The vendor revises a return of a stock that was never in the universe. Which stages run, and which outputs change?
Solution
Solution of Exercise 29.3.
Snapshot, universe, features, cards and level 1 run, since each reads the snapshot; only the snapshot’s output changes, so blend, portfolio and the tear sheet are read from the cache.
Exercise 29.4 ★★
Which of Sandve’s ten rules does the pipeline enforce, and which does it leave to people?
Solution
Solution of Exercise 29.4.
It enforces keeping track of how each result was produced (1), recording intermediate results (5), noting seeds (6, by flagging unseeded stages in reproduction), and part of archiving versions (3, the environment record). Avoiding manual steps (2), version control (4), keeping plot data (7), hierarchical output (8), linking statements to results (9) and public access (10) remain practices, though the manifest makes 9 easy.
Exercise 29.5 ★★
A stage calls a helper function in another module. Its code hash is the hash of the stage’s own source. What can go wrong, and how would you fix it?
Solution
Solution of Exercise 29.5.
A change to the helper changes the stage’s output without changing its key, so the cache serves a stale result. Hash the functions the stage calls (its module’s source, or the transitive closure of the functions it uses), or version helpers explicitly as parameters.
Exercise 29.6 ★★
Reproduction fails on a colleague’s machine for every stage that uses floating-point reductions. What do you check first?
Solution
Solution of Exercise 29.6.
The environment: library versions (NumPy’s summation and BLAS builds change the order of floating-point operations), the number of threads, the platform. The manifest’s environment record says what to compare; bitwise reproduction across environments may require pinned versions and single-threaded reductions (Book 4, chapter 25).
Exercise 29.7 ★★★
Coding. Remove the tear sheet’s seed and run reproduce. Then give the bootstrap a seed derived from the manifest’s other keys and explain why that is reproducible but still honest.
Solution
Solution of Exercise 29.7.
Without a seed, reproduce returns the tear sheet alone. A seed derived from the hash of the stage’s inputs and parameters is fixed for a given computation, so it reproduces, and it is chosen by no one, so it cannot be tuned: a result that depends on the seed is still exposed by varying the inputs.
Exercise 29.8 ★★★
Find the flaw. “The study is reproducible: the notebook is in version control.”
Solution
Solution of Exercise 29.8.
Version control keeps the code, not the data it read (the vendor revised them), not the environment, not the order in which the cells were run, not the seeds, nor which version produced the published numbers. A manifest records all of them.
29.7 Problem: The Result That Could Not Be Reproduced
Problem 29.1
Weekend problem — a result, kept
The reference study in firm.workflow.
Part I — The pipeline.
- List the eight stages and their dependencies.
- What goes into a stage’s key, and what into the manifest?
- What does the second run do, and why?
- What does the study find, with its uncertainty?
Part II — Changes.
- Which stages re-run after a revision to a name never in the universe, and which outputs change?
- And after a revision to a name always in the universe?
- After a change of blend weights?
- After a change to the tear sheet’s code?
Part III — Reproduction.
- What does reproducing the manifest in an empty cache return?
- What happens without the bootstrap’s seed?
- What does the environment record, and when does it matter?
Part IV — The verdict.
- State the named result: the stages invalidated by a vendor revision, and the manifest diff that proves the reproduction.
- How is the run registered in the research log?
- Which items of the review checklist does the pipeline answer, and which not?
- Why does the reversal’s daily signal not survive the monthly book?
- How would the eighteen-month-old result have been kept?
- What would you add before the desk relies on the pipeline?
- In one sentence: what is a reproducible result?
Solution
Solution of Problem 29.1.
- Snapshot; universe (snapshot); features (snapshot, universe); cards (snapshot, features); blend (features, universe); portfolio (blend, universe); level 1 (portfolio, snapshot); tear sheet (level 1).
- The key: name, code hash, parameters, seed, the inputs’ output hashes. The manifest: for each stage its key, code hash, parameters, seed, dependencies, input hashes, output hash and whether it ran, and the environment.
- It reads every stage from the cache: all keys are stored.
- Momentum’s rank IC 0.0076 (), the reversal’s 0.0379 () over 2 267 days; the monthly book’s Sharpe ratio after costs 0.08 (standard error 0.33, bootstrap interval to 0.74), turnover 3.5% a day.
- Snapshot, universe, features, cards, level 1 re-run; only the snapshot’s output changes.
- All eight re-run, and all eight outputs change.
- Blend, portfolio, level 1 and tear sheet.
- The tear sheet only.
- The same output hash for all eight stages: an empty list of differences.
- Reproduction flags the tear sheet, and only it.
- Python and NumPy versions and the platform; when a reproduction fails on another machine, or floating-point results differ across library versions.
- Named result. A revision outside the universe re-runs five stages and changes one output (the snapshot’s); a revision inside it re-runs and changes all eight; reproducing the manifest from scratch returns identical output hashes for all eight stages.
- As a trial of its family, with the hash of the manifest’s stage keys as its code identity and the snapshot’s output hash as its data identity, and its metrics.
- Reproduction, point-in-time data (partly: the snapshot’s identity and revisions), the trial count (through registration); not the fidelity level, the costs at size, or whether the text matches the numbers.
- The book is rebuilt monthly and the reversal decays within a day (chapter 6’s horizon mismatch).
- As a manifest with its cache and environment, registered in the log: one command to re-run it, and a diff to say what the vendor’s revision changed.
- Code hashes that follow called functions, a shared cache, enforced environment locks, and review sign-off recorded with the manifest.
- One that its manifest produces again, bit for bit, from the same data, code, parameters, seeds and environment.
29.8 Interview questions
Interview question 29.1 ★ researcher, developer
A backtest you ran last year gives a different Sharpe ratio today. List the possible causes.
Solution
Solution of Interview question 29.1.
Revised or re-snapshotted data (vendor restatements, survivorship fixes), changed code or library versions, different parameters, unseeded randomness, a different universe or calendar, look-ahead fixed since, or a different environment (threads, platform).
Interview question 29.2 ★★ researcher
What practices make computational research reproducible?
Solution
Solution of Interview question 29.2.
Keep track of how every result was produced; automate every step; version code and data (snapshots); record intermediate results; seed randomness; lock environments; link every number in a write-up to its source (Sandve and co-authors’ ten rules).
Interview question 29.3 ★★ developer
Design a cache for research pipeline stages. How do you invalidate it?
Solution
Solution of Interview question 29.3.
Content-addressed: key each stage by the hash of its code, parameters, seed and its inputs’ content hashes; store outputs under their keys; never invalidate explicitly, since a changed input produces a new key; build keys on output hashes for early cutoff; garbage-collect unreferenced entries.
Interview question 29.4 ★★ researcher, risk
What would you check when reviewing a colleague’s strategy research?
Solution
Solution of Interview question 29.4.
Reproduction from a manifest; point-in-time data and universe; the trial count, deflated Sharpe ratio and PBO; the fidelity level of the backtest; costs and capacity at the intended size; stability over sub-periods; that the text’s numbers match the outputs.
Interview question 29.5 ★★ developer
How do you make a computation that uses random numbers reproducible across machines?
Solution
Solution of Interview question 29.5.
Seed every generator explicitly (never from the clock), use a generator with a documented stream (not a library default that may change), record the seed and the library version, and make the order of draws independent of threads.
Interview question 29.6 ★★★ developer, researcher
Your vendor restates five years of history. How do you find every result that depended on it, and which of them change?
Solution
Solution of Interview question 29.6.
Find every manifest whose snapshot stage used the vendor’s data (the log’s data identity); re-run each with the restated snapshot; the diffs say which stages’ outputs changed, and early cutoff says which results are untouched.