Quantitative Finance · Book 15 · Technology

Research, Data and Risk Platforms

Research, Data and Risk Platforms · Technology

15The Research Environment

In 2013 three economists replicating an influential paper on public debt and growth obtained the authors’ working spreadsheet. A coding error in it excluded five countries — Australia, Austria, Belgium, Canada and Denmark — from the analysis; with that and two other problems corrected, the average growth of countries with public debt above 90% of output was 2.2%, not the −0.1%-0.1\% published (Herndon, Ash and Pollin). The economics was not what failed. What failed was an environment in which a calculation could not be rerun from the top by anyone but its author, and not even by its author once the state of the spreadsheet had drifted from what it showed. A research platform’s notebooks have the same property, more so. This chapter audits a notebook that lies in the same way, reruns it from the top, promotes its calculation to a library, and reproduces the corrected number from a clean directory.

15.1 Notebooks and their hidden state

Definition 15.1 (Computational notebook)

A computational notebook is a document of cells — code, text and the outputs the code produced — run one cell at a time against a live interpreter (the kernel), in any order the user chooses; Jupyter’s format stores it as JSON, with each code cell’s source, its saved outputs and the kernel’s execution counter when it last ran (an integer, or null if it never ran).

Definition 15.2 (Hidden state)

A notebook has hidden state when its saved outputs depend on something its cells, run from the top, would not reproduce: a cell run out of order, a cell run twice, a cell edited or deleted after it ran, a name defined only at the prompt.

The notebook is the fastest way to explore data and the easiest way to publish a number nobody can recompute. Pimentel, Murta, Braganholo and Freire collected Jupyter notebooks from GitHub and tried to run them: out of 863 878 attempted executions of valid notebooks, only 24.11% executed without errors and only 4.03% produced the same results. Execution counters give hidden state away — a counter skipped means a cell re-run or run and deleted, a counter lower than the cell above it means execution out of document order — but only if something reads them.

The chapter’s planted notebook computes the headline of the hook’s debate on a synthetic panel of twenty countries from 1946 to 2009: the average real growth of country-years with debt above 90% of output, averaged by country. Its history is written down in the code: the cells, in document order, and what was run, in the order it was run.

CELLS = {                       # the notebook's code cells, in document order
    "load": 'import pandas as pd\ndf = pd.read_csv("panel.csv")',
    "clean": "df = df[df.year >= 1950]",
    "high": "high = df[df.debt > 90]",
    "select": "sel = high[~high.country.isin(exclude)]",
    "result": 'result = sel.groupby("country").growth.mean().mean()\n'
              'print(f"{result:.2f}")',
}
DELETED = f"exclude = {EXCLUDED!r}"
HISTORY = ["load", "high", "DELETED", "select", "result", "clean"]    # what ran, in that order
Listing 15.1. The planted notebook: five code cells in document order, a cell that was run and then deleted, and the order in which the analyst ran them — the cleaning cell, near the top, last. code/platforms/15-the-research-environment/python/pl_envaudit.py
The planted notebook as saved: code cells in document order with their execution counters. The cleaning cell, second on the page, ran sixth, after the result; the cell that set exclude ran third and was deleted. The saved 1.96 comes from that history.
Figure 15.1. The planted notebook as saved: code cells in document order with their execution counters. The cleaning cell, second on the page, ran sixth, after the result; the cell that set exclude ran third and was deleted. The saved 1.96 comes from that history.

Saved, the notebook shows 1.96 under its result cell. Nothing in the document says that five countries were excluded, or that the years before 1950 were still in the data when the number was computed.

15.2 Auditing the document

firm.envaudit reads the notebook’s JSON and nothing else (Listing 15.2). It walks the code cells in document order with their counters, and parses each cell to see which names it reads and which it binds.

def audit(nb: dict) -> list[Finding]:
    out: list[Finding] = []
    cells = code_cells(nb)
    ran = [(i, n) for i, _s, n, _o in cells if n is not None]
    for (i, n), (j, m) in zip(ran, ran[1:], strict=False):
        if m < n:
            why = f"counter {m} below the cell above it ({n}, cell {i})"
            out.append(Finding("out-of-order", j, why))
    counters = sorted(n for _i, n in ran)
    for a, b in zip(counters, counters[1:], strict=False):
        if b - a > 1:
            gap = f"{a + 1}" if b - a == 2 else f"{a + 1} to {b - 1}"
            why = f"counter {gap} missing: a re-run or a deleted cell"
            out.append(Finding("skipped-counter", -1, why))
    everywhere = set()
    for _i, s, _n, _o in cells:
        everywhere |= _names(s)[0]
    defined = set()
    for i, s, n, outputs in cells:
        if n is None and outputs:
            why = "outputs saved but no execution counter"
            out.append(Finding("output-without-run", i, why))
        if n is None and not outputs and s.strip():
            out.append(Finding("never-run", i, "code never executed"))
        bound, read = _names(s)
        for name in sorted(read - defined - BUILTINS):
            kind = "used-before-defined" if name in everywhere else "never-defined"
            out.append(Finding(kind, i, name))
        defined |= bound
    return out
Listing 15.2. The audit: counters out of document order, skipped counters, outputs without a run, cells never run, and names read before any cell above binds them — or bound by no cell at all. code/firm/envaudit/firm_envaudit.py

On the planted notebook it finds three things (Table 15.1). The cell that selects country-years above 90% ran second, but a cell above it ran sixth: the cleaning cell was run after the result, so the displayed data are not the data the result used. Counter 3 is missing: a cell was run and is gone. And the selection cell reads exclude, which no cell in the notebook defines.

findingcelldetail
out of order3counter 2 below the cell above it (6, cell 2)
skipped counter—counter 3 missing: a re-run or a deleted cell
never defined4exclude
Table 15.1. The audit of the planted notebook, from its JSON alone (cell 0 is the title).

Rerun from the top in a fresh interpreter (Listing 15.3), the notebook fails: NameError: name ’exclude’ is not defined. The number on the page cannot be produced by the page.

def rerun(nb: dict, workdir, timeout: int = 120, env: dict | None = None) -> Rerun:
    """All code cells, top to bottom, as one script in a fresh interpreter."""
    workdir = pathlib.Path(workdir)
    workdir.mkdir(parents=True, exist_ok=True)
    script = workdir / "rerun.py"
    script.write_text("\n\n".join(s for _i, s, _n, _o in code_cells(nb)) + "\n")
    r = subprocess.run([sys.executable, str(script)], cwd=workdir, capture_output=True,
                       text=True, timeout=timeout, env=env)
    return Rerun(r.returncode == 0, r.stdout, r.stderr, r.returncode)
Listing 15.3. The rerun: the code cells, top to bottom, as one script, in a fresh interpreter in a subprocess. code/firm/envaudit/firm_envaudit.py

15.3 From notebook to library

Definition 15.3 (Research library)

A research library is the shared, versioned, tested code that research notebooks and pipelines import: calculations promoted out of notebooks once they are used more than once, with their choices as explicit arguments and their behaviour pinned by tests.

The fix is to decide, in the open, what the calculation is. With no exclusion and the cells run as they stand, the notebook prints 2.44. The calculation belongs in the library (Listing 15.4): the threshold, the first year, the excluded countries and the weighting become arguments, and each has a test. The library reproduces all three numbers the debate turns on: 1.96 with the five countries excluded and all years (the saved number, explained), 2.44 with every country from 1950 (the corrected number), and 2.90 if every country-year counts once instead of every country — the weighting choice the hook’s critique also questioned. The five excluded countries hold 204 of the 455 country-years above 90% since 1950.

The headline under the three choices the notebook hid: excluding five countries or not, starting in 1946 or 1950, and averaging countries or country-years. The saved 1.96 is the top pair’s mean of country means (blue); the corrected 2.44, the bottom pair’s. Data: fig_envaudit.py.
Figure 15.2. The headline under the three choices the notebook hid: excluding five countries or not, starting in 1946 or 1950, and averaging countries or country-years. The saved 1.96 is the top pair’s mean of country means (blue); the corrected 2.44, the bottom pair’s. Data: fig_envaudit.py.
def average_growth(df: pd.DataFrame, threshold: float = 90.0, since: int | None = None,
                   exclude=(), weight: str = "country") -> float:
    """Average growth of country-years with debt above `threshold` (percent of output).

    weight='country': the mean of each country's mean (each country counts once);
    weight='country-year': the pooled mean of the country-years. `since` drops earlier
    years and `exclude` drops countries: the caller states both, no notebook remembers them."""
    d = df[df["debt"] > threshold]
    if since is not None:
        d = d[d["year"] >= since]
    if len(exclude):
        d = d[~d["country"].isin(list(exclude))]
    if weight == "country":
        return float(d.groupby("country")["growth"].mean().mean())
    if weight == "country-year":
        return float(d["growth"].mean())
    raise ValueError(weight)
Listing 15.4. The promoted function: the choices that lived in the notebook’s memory are arguments. code/platforms/15-the-research-environment/python/research_lib.py

15.4 Review

Definition 15.4 (Code review)

Code review is the examination of a change to shared code by someone other than its author before it is merged: of what the change does, whether its tests show it, and whether it does what its description says.

Review is where the research review of Book 7 (chapter 1) meets the codebase. A reviewer cannot review a notebook’s hidden state, only a diff: the promotion of a calculation to the library is what makes it reviewable. The chapter’s promotion checklist, run by promotion_checklist, is the minimum a reviewer asks for: the audit finds no hidden state, the notebook reruns from the top, the rerun reproduces the saved result, the library functions are tested, the environment matches the lock, the data match the snapshot.

Definition 15.5 (Monorepo)

A monorepo is a single version-control repository holding the code of many projects, libraries and teams, so that a change to a library and to every caller can be made, reviewed and tested together.

Google’s account of its own (Potvin and Levenberg, 2016) describes a single repository serving as a common source of truth for tens of thousands of developers, made workable by systems and workflows built for its scale. This series’ own code (code/firm/) is a monorepo in miniature: a component of one book is used by the next, and a change to it is tested against every chapter that imports it.

15.5 Environments and containers

Definition 15.6 (Container image)

A container image is a packaged file system and configuration — an operating-system layer, the interpreter, the libraries, the code — from which identical isolated environments can be started on any machine that runs containers.

The Open Container Initiative specifies the format: an image manifest, an optional image index, a set of file-system layers and a configuration, so that tools for building, transporting and running images interoperate. A container freezes an environment; an environment lock (Book 7, chapter 29) describes one. The chapter uses the lighter of the two: capture_env lists the interpreter and every installed distribution and compare_env checks them against the repository’s requirements.txt; the data are checked against a manifest of hashes (snapshot). Where the lock is enough, a container adds nothing but weight; where system libraries or compilers matter (chapter 9’s extensions), it is the only honest record.

As of September 2026 — Reproducibility tooling

The Jupyter notebook format (nbformat) stores each code cell’s execution_count as an integer or null, and its outputs as a list; the Open Container Initiative’s image specification defines an image as a manifest, an optional index, file-system layers and a configuration. Both are open specifications; notebook linters and container build tools built on them change often, the formats less.

15.6 Reproducibility as a routine

The last step reproduces the corrected number from a clean directory: the promoted library, the data file, and a three-line script that imports the library and prints the headline. The data file is checked against the snapshot’s hash (no difference), the environment against the lock (no difference), and the script prints 2.44 in a fresh interpreter. It is the reproducible result of Book 7 (chapter 29), obtained not by care but by routine: a result is published only with the command that recomputes it from a clean checkout.

Reproduction as a routine: a clean directory with the promoted library and the data file, the data checked against the snapshot’s hash and the environment against the lock, and the headline recomputed in a fresh interpreter.
Figure 15.3. Reproduction as a routine: a clean directory with the promoted library and the data file, the data checked against the snapshot’s hash and the environment against the lock, and the headline recomputed in a fresh interpreter.

Example 15.7 (Rerun from the top)

The planted notebook shows 1.96. The audit finds a cell run out of order, a deleted cell and a name defined nowhere; the rerun from the top fails on that name. Without the exclusion, and with the cells as they stand, the result is 2.44; the library shows that 1.96 is exactly the average with five countries excluded and the years before 1950 still in the data, and that pooling country-years instead of countries would give 2.90. From a clean directory, with the data snapshot and the environment lock checked, the library reproduces 2.44.

15.7 Tutorial: audit, rerun, promote, reproduce

Goal. Find a notebook’s hidden state from its JSON, show that its number cannot be recomputed, fix and promote the calculation, and reproduce it from a clean directory. End state: Table 15.1 and the reproduction’s three checks.

  1. Plant: pl_envaudit.planted(dir) writes the data and the notebook with the outputs of its history.
  2. Audit: firm_envaudit.audit(nb).
  3. Rerun: firm_envaudit.rerun(nb, dir), then with fixed(nb).
  4. Promote: research_lib.average_growth and its tests.
  5. Reproduce: pl_envaudit.reproduce(clean_dir, dir): snapshot, lock, rerun.

What to change next. Run the audit on your own notebooks; add a check that every saved output of a notebook is reproduced by its rerun (outputs compared cell by cell).

15.8 Build: the environment audit

Purpose. The platform’s gate between exploration and anything others rely on: a notebook’s number becomes a result only when it can be recomputed from the top, from a clean checkout, in the locked environment, on the snapshot of its data.

Interface. load, code_cells, audit -> [Finding], saved_output, rerun -> Rerun, capture_env, compare_env, snapshot, check_snapshot, promotion_checklist.

Rules. The audit reads the document only; the rerun uses a fresh interpreter in a subprocess, never the author’s kernel; environments are compared with the lock file, data with hashes.

Acceptance tests. code/firm/envaudit/tests/: a clean notebook has no findings and reruns; a planted one yields each kind of finding and a failed rerun; imports, functions and built-ins are not findings; lock and snapshot comparisons; the checklist passes and fails as it should.

Stretch. Cell-by-cell comparison of saved and rerun outputs; a container build from the lock; the audit as a pre-commit hook in the monorepo.

Sources and further reading

  • T. Herndon, M. Ash and R. Pollin, “Does high public debt consistently stifle economic growth? A critique of Reinhart and Rogoff”, Cambridge Journal of Economics 38(2), 2014 (PERI working paper 322, 2013).
  • J. F. Pimentel, L. Murta, V. Braganholo and J. Freire, “A large-scale study about quality and reproducibility of Jupyter notebooks”, MSR 2019.
  • R. Potvin and J. Levenberg, “Why Google stores billions of lines of code in a single repository”, Communications of the ACM 59(7), 2016.
  • Jupyter nbformat documentation; Open Container Initiative image specification.

15.9 Exercises

Exercise 15.1 ★

A notebook’s code cells show counters 1, 2, 5, 3, 6 in document order. What does the audit report?

Solution

Solution of Exercise 15.1.

The fourth code cell, with counter 3, is out of order (the cell above it has 5); counter 4 is missing (a re-run or a deleted cell). The third cell ran after the fourth, and something else ran between them.

Exercise 15.2 ★

Why does the audit report exclude as never defined rather than used before defined?

Solution

Solution of Exercise 15.2.

No cell of the notebook binds exclude, above or below: it was bound by a cell that has been deleted. “Used before defined” is for a name that a cell further down binds — a rerun from the top fails there too, but the definition is still in the document.

Exercise 15.3 ★

Which two differences between the saved and the corrected calculation account for 1.96 against 2.44?

Solution

Solution of Exercise 15.3.

The five countries excluded by the deleted cell, and the years before 1950, which were still in the data when the result ran because the cleaning cell ran afterwards. With the exclusion and all years the library gives 1.96; with every country from 1950, 2.44.

Exercise 15.4 ★★

Can a notebook with consecutive counters in document order still have hidden state? Give an example.

Solution

Solution of Exercise 15.4.

Yes. A cell edited after it ran shows its new source with the old output; a cell that appends to a list, run once with consecutive counters and then re-run after a restart that re-used counter numbers, can hide nothing in the counters; a name set at the prompt of an attached console leaves no trace. Only the rerun is conclusive.

Exercise 15.5 ★★

Why is weighting by country (2.44) or by country-year (2.90) a choice that belongs in the function’s arguments, and which would you report?

Solution

Solution of Exercise 15.5.

Both answer a question; they are different questions — the typical country’s experience against the typical country-year’s — and the difference (2.44 against 2.90) is larger than many effects reported. A function that hides the choice publishes one answer as if it were the only one; as an argument, it is stated in the call and in the report. Report the one that matches the question, and the other beside it.

Exercise 15.6 ★★

When is an environment lock enough, and when is a container image needed?

Solution

Solution of Exercise 15.6.

A lock is enough when the result depends only on the interpreter and the Python packages it pins, on a known operating system; a container is needed when it depends on system libraries, compilers or tools outside the lock (native extensions, database clients), or must run on machines you do not control.

Exercise 15.7 ★★★

Coding. Add to firm.envaudit a check that compares each cell’s saved output with the rerun’s output for that cell. What does it report on the fixed notebook, and why is it harder than comparing the last output?

Solution

Solution of Exercise 15.7.

On the fixed notebook the check has nothing to compare: the fix cleared the saved outputs, as it should, since they came from a history that no longer exists. On the planted notebook the rerun stops at the fourth code cell, so the result cell’s saved 1.96 is never compared with anything. It is harder than comparing the last output because the rerun’s output must be attributed to cells (markers printed between them), some outputs legitimately differ (times, memory addresses, the last digits of floating-point sums), and rich outputs such as plots need their own comparison.

Exercise 15.8 ★★★

Find the flaw. “The notebook is in version control and the reviewer approved it, so its result is reproducible.”

Solution

Solution of Exercise 15.8.

Version control stores the document, with its saved outputs; it does not store the kernel’s state that produced them, and a reviewer reads the document. Reproducible means rerun: from a clean checkout, in the locked environment, on the snapshot of the data, from the top, with the same result. Neither the commit nor the approval did that.

15.10 Problem: Rerun From the Top

Problem 15.1

Weekend problem — a number the page cannot produce

The planted notebook of Listing 15.1, firm.envaudit and the promoted library.

Part I — The notebook.

  1. What does a Jupyter notebook store for each code cell?
  2. In what order were the cells run, and in what order are they shown?
  3. What number does the notebook show?
  4. What is hidden state, in three forms?
  5. What does the large-scale study of notebooks on GitHub report?

Part II — The audit.

  1. What three findings does the audit make, and from what?
  2. What does the rerun from the top do?
  3. What does each finding explain about the saved number?
  4. What does the audit not see?
  5. Why rerun in a fresh interpreter in a subprocess?

Part III — The library.

  1. What becomes an argument of the promoted function?
  2. What are the three numbers it reproduces, and under which choices?
  3. How many country-years above 90% since 1950 do the five excluded countries hold?
  4. What does a reviewer check before the promotion is merged?
  5. Why does a monorepo help?

Part IV — The verdict.

  1. State the named result: the headline number as saved and as rerun, the findings that explain the difference, and the reproduction of the corrected result from a clean environment.
  2. What would you require before a notebook’s number goes into a research report?
  3. When would you build a container image?
  4. What should a research log (Book 7, chapter 1) record about this analysis?
  5. In one sentence: what makes a number reproducible?
Solution

Solution of Problem 15.1.

  1. Its source, its saved outputs, and the kernel’s execution counter when it last ran (an integer, or null).
  2. Run: load, select above 90%, the deleted exclusion, the country selection, the result, then the cleaning; shown: load, cleaning, selection, country selection, result.
  3. 1.96.
  4. A cell run out of order, a cell re-run, a cell edited or deleted after it ran (and names set at the prompt).
  5. Of 863 878 attempted executions of valid notebooks, 24.11% ran without errors and 4.03% reproduced their results.
  6. Out of order (cell 3 ran before cell 2), a skipped counter (3), and a name defined nowhere (exclude); from the JSON alone.
  7. It fails: NameError on exclude.
  8. The skipped counter and the undefined name are the deleted exclusion; the out-of-order counter says the displayed data are not those the result used.
  9. Edits after a run, re-runs that leave consecutive counters, and anything done outside the notebook.
  10. So that nothing of the author’s session — names, imports, working directory — leaks into the check.
  11. The threshold, the first year, the excluded countries and the weighting.
  12. 1.96 (five countries excluded, all years), 2.44 (all countries from 1950, by country), 2.90 (all countries from 1950, by country-year).
  13. 204 of 455.
  14. The checklist: no hidden state, rerun succeeds and reproduces, tests, environment against the lock, data against the snapshot.
  15. The library and every caller change together and are tested together.
  16. Named result. The notebook shows 1.96; rerun from the top it fails on a name defined only by a deleted cell; the audit finds that cell (a skipped counter and a name defined nowhere) and a cleaning cell run after the result; corrected, the calculation gives 2.44, which the promoted library reproduces from a clean directory with the data snapshot and environment lock checked.
  17. The promoted calculation, a rerun from a clean checkout that prints it, the lock and the data snapshot, and the choices stated in the call.
  18. When the result depends on anything outside the lock, or must run on machines I do not control.
  19. The question, the choices (threshold, years, countries, weighting), the command that recomputes the result, the code and data hashes, and the alternatives tried.
  20. That anyone can recompute it from the top, from a clean checkout, and get the same number.

15.11 Interview questions

Interview question 15.1 ★ researcher

A colleague’s notebook shows a result you cannot reproduce. What do you check?

Solution

Solution of Interview question 15.1.

Rerun it from the top in a fresh kernel; read its execution counters for skips and out-of-order runs; look for names defined nowhere; check the data version and the package versions against what the author used.

What the interviewer is looking for: Rerun first, then hidden state, data and environment.

Interview question 15.2 ★★ researcher, developer

When should code move from a notebook to a library?

Solution

Solution of Interview question 15.2.

As soon as it is used twice, feeds anything others rely on, or its choices matter: then it needs arguments for its choices, tests, and review.

What the interviewer is looking for: Reuse and consequence, not size.

Interview question 15.3 ★★ developer

What does a container give you that a requirements file does not, and what does it cost?

Solution

Solution of Interview question 15.3.

The whole environment below the Python packages — operating system, system libraries, compilers — frozen and portable. It costs build time, image storage and maintenance, and an image can still hide what is in it unless it is built from a lock.

What the interviewer is looking for: Freezing versus describing; build from a lock.

Interview question 15.4 ★★ researcher

How would you make every number in a research report reproducible a year later?

Solution

Solution of Interview question 15.4.

Produce every number by a command in the repository that runs from a clean checkout on a data snapshot in a locked environment, and record in the report the commit, the snapshot hash and the command; test in continuous integration that the commands still reproduce the numbers.

What the interviewer is looking for: Numbers as outputs of versioned commands.

Interview question 15.5 ★★★ developer, researcher

Design the research environment for a team of twenty researchers sharing data and code.

Solution

Solution of Interview question 15.5.

Notebooks for exploration, a shared tested research library in one repository with review, environments from a lock (containers where needed), data read through the platform’s snapshots, an audit and rerun gate before any result is shared, and a research log with the commands that reproduce each result.

What the interviewer is looking for: Exploration free, results gated.

Terms defined in this chapter

See all 2333 terms in the glossary