---
title: "Experiment Tracking and Reproducibility"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 25
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/25-experiment-tracking-and-reproducibility
---

# Chapter 25 — Experiment Tracking and Reproducibility

A regulator asks a firm to show the model that traded on a given morning, the data it was trained on and every alternative it was chosen over. The firm can name the model; it takes three weeks to retrain something close to it, and it cannot say how many alternatives there were. This chapter builds the records that make the answer a query: a tracker that stores every run with everything that determined it, a registry that says which model was in production when and where it came from, and an [audit trail](#def-ml-experiment-tracking-and-reproducibility-registry) that cannot be edited without the edit showing. On the chapter’s study, 200 tuning runs of a boosted forecast, the registry reproduces the March model bit for bit from its records alone; with any one of four fields missing, nothing close to it comes back. The tracker also answers the question the regulator did not ask: counting the 200 trials, the champion’s validation Sharpe ratio of 2.22 has a deflated Sharpe ratio of 0.15.

## 25.1 What must be recorded

**Definition 25.1 (Experiment tracker, run record).**

An *experiment tracker* stores a record of every training run a team makes, including the ones that are discarded. A *run record* holds everything that determines the run’s result (parameters, seeds, the data snapshot, the code version, the environment) and everything the run produced (metrics, artefacts, logs), with the time and the author.

The research log of Book 7 (chapter 1) recorded trials to count them; the tracker records them to rebuild them. The chapter’s tracker ([Listing 25.1](#lst-ml-exp-run)) writes each run as a JSON record whose identifier is the hash of its content, stores the fitted model once under the content hash of its dump (Book 7’s `firm.workflow`), and logs the run as a trial in `firm.researchlog`. Timestamps are passed in, never read from the clock, so that the tracker’s own contents are reproducible.

The study: a gradient-boosted forecast (chapter 5) of next-day returns on a synthetic panel of 40 names and eight features over 700 days, trained on the first 400, scored by the Sharpe ratio of a daily long-short portfolio on the next 150, and tested on the last 150. Two hundred runs draw five parameters and a seed at random. Their validation Sharpe ratios range from $-3.16$ to 2.22, with a median of $-0.83$ ([Figure 25.1](#fig-ml-exp-sharpes)): the signal is weak, and most configurations fit noise.

![Validation Sharpe ratios of the 200 tracked runs; circle: the champion’s (2.22); square: its test Sharpe ratio (0.91); dashed: the expected maximum of 200 independent strategies with no skill over 150 days (3.58). Data: ml_track.study.](https://one-course.com/images/onecourse/chapters/quant-12/ml-experiment-tracking-and-reproducibility/fig-f8d8e9e75d4f.svg)

***Figure 25.1.** Validation Sharpe ratios of the 200 tracked runs; circle: the champion’s (2.22); square: its test Sharpe ratio (0.91); dashed: the expected maximum of 200 independent strategies with no skill over 150 days (3.58). Data: `ml_track.study`.*

The champion’s test Sharpe ratio is 0.91, less than half its validation figure. Its probabilistic Sharpe ratio, the probability that its true Sharpe ratio is positive as if it were the only run, is 0.96; its deflated Sharpe ratio (Book 4, chapter 12), with the trial count taken from the tracker, is 0.15. The expected best of 200 skill-less strategies over 150 days is 3.58, above the champion: the correction here is conservative, because the runs are correlated and are fewer independent trials than 200, but a result that cannot beat the best of its own noise has not earned production on validation evidence alone.

## 25.2 Versioning data, code and models

**Definition 25.2 (Data versioning, model artefact).**

*Data versioning* identifies every dataset a run reads by an immutable snapshot (a content hash or a dated, frozen version), so that later corrections create new versions instead of changing old ones. A *model artefact* is the stored result of a training run, the fitted parameters in a portable form, identified by its content hash.

Versions exist because data and code change under a model. In the study, the data vendor revises 1% of the returns in March (snapshot 2026-03 replaces 2026-01), and the feature code changes its clipping from three standard deviations to two and a half (v2 replaces v1). The champion was trained on 2026-01 with v1. The record says so; without it, a team would use what it has now.

## 25.3 The registry and the audit trail

**Definition 25.3 (Model registry, model lineage, audit trail).**

A *model registry* lists the firm’s models by name and version with their stage (candidate, production, retired) and the history of stage changes. *Model lineage* links a registered version to the run, data snapshot and code that produced it. An *audit trail* is an append-only record of actions (registrations, promotions, overrides) with who, when and why, protected against silent edits; here each entry carries the hash of the previous one ([Listing 25.2](#lst-ml-exp-audit)).

**As of September 2026 — Records of algorithmic trading changes in the EU.**

Commission Delegated Regulation (EU) 2017/589 (the regulatory technical standards on organisational requirements for algorithmic trading, known as RTS 6) requires an investment firm to “keep records of any material change made to the software used for algorithmic trading”, allowing it to determine when a change was made, who made it, who approved it and its nature (Article 5(7)); order records of high-frequency traders are kept for five years (Article 28).

The registry answers the regulator’s first question by time ([Listing 25.3](#lst-ml-exp-asof)): the version in production on the morning of 10 March is found from the stage histories, and its lineage leads to the [run record](#def-ml-experiment-tracking-and-reproducibility-tracker). The [audit trail](#def-ml-experiment-tracking-and-reproducibility-registry) answers “who promoted it, and why”: version 1 of `xs-gbdt`, promoted on 2 March as the champion of 200 runs. Changing any character of any line of the trail breaks the hash chain at that line.

## 25.4 Reproducing a model, and when bitwise is too much

From the registry entry alone the chapter re-runs the training ([Listing 25.4](#lst-ml-exp-repro)) and obtains a model whose dump has the same content hash as the stored artefact: bit for bit the model that traded. Then it removes one field of the record at a time and replaces it with what a team would assume without it ([Table 25.1](#tab-ml-exp-ablation)). Every removal breaks reproduction, and none is harmless: the reproduced models’ test Sharpe ratios differ from the original’s by $-1.28$ with default parameters, $-1.25$ with the revised data, $-0.44$ with the current feature code, and $+0.85$ with the default seed.

| field missing | assumed instead | reproduced bit for bit | change in test Sharpe ratio |
| --- | --- | --- | --- |
| none (the registry’s record) |  | yes | 0 |
| parameters | the firm’s defaults | no | $-1.28$ |
| seed | 1 | no | $+0.85$ |
| data snapshot | the latest (2026-03) | no | $-1.25$ |
| code version | the current (v2) | no | $-0.44$ |

***Table 25.1.** Reproducing the production model from its record, with each field in turn replaced by an assumption. Data: `ml_track.reproduction`.*

The seed’s row carries a second lesson: a different seed of the same configuration moves the test Sharpe ratio by 0.85, as much as the difference between configurations. A result that does not survive a change of seed was not a result, and reporting the spread over seeds (chapter 7) is part of what a tracker enables.

Reproduction bit for bit is the right standard for a model’s own artefact, rebuilt on the same software and hardware; the chapter’s fits are deterministic because LightGBM runs single-threaded with its deterministic option. Across library versions, thread counts or hardware (chapter 23’s [bfloat16](https://one-course.com/books/quant/12/en/chapter/23-training-infrastructure#def-ml-training-infrastructure-mixed)), floating-point sums change order and bits change; there the standard is statistical: the same metrics within the spread over seeds, checked by re-running. What must always be exact is the record: the data, code, parameters and seed that produced the artefact.

**Method 25.4 (Tracking research that trades).**

1. Track every run, discarded ones included, with its parameters, seed, data snapshot, code version and environment, and store artefacts by content hash.
2. Take the trial count for deflation from the tracker, not from memory.
3. Register production models with lineage and stage history; promote and retire only through the registry, with reasons in an append-only trail.
4. Test reproduction from the registry alone, regularly, and treat a failure as an incident.

## 25.5 Tutorial: reproduce the March model

**Goal.** Track 200 tuning runs, register and promote the champion, find it again by date, reproduce it bitwise, and break reproduction one field at a time. **End state:** [Figure 25.1](#fig-ml-exp-sharpes), [Table 25.1](#tab-ml-exp-ablation).

1. **A [run record](#def-ml-experiment-tracking-and-reproducibility-tracker).** `def run (self , params, seed, data, code, env, metrics, artefact, ts): a_hash = content_hash(pickle.dumps(artefact)) path = self .root / " artefacts " / a_hash if not path.exists(): path.write_bytes(pickle.dumps(artefact)) rec = {" params " : params, " seed " : seed, " data " : data, " code " : code, " env " : env, " metrics " : metrics, " artefact " : a_hash, " ts " : ts} rid = hashlib.sha256(json.dumps(rec, sort_keys=True ).encode()).hexdigest()[:16 ] rec[" id " ] = rid (self .root / " runs " / f " { rid} .json " ).write_text(json.dumps(rec, sort_keys=True )) if self .log is not None : self .log.trial(self .family, params, metrics, ts, code_hash=code, data_id=data) return rid` **Listing 25.1.** Recording a run: artefact by content hash, record by its own hash, trial in the research log. code/firm/exptrack/firm_exptrack.py
2. **An append-only trail.** `def append (self , kind, body, ts): prev = self ._lines()[-1 ][" hash " ] if self ._lines() else GENESIS rec = {" kind " : kind, " ts " : ts, " body " : body, " prev " : prev} rec[" hash " ] = hashlib.sha256(json.dumps(rec, sort_keys=True ).encode()).hexdigest() with open (self .path, " a " ) as f: f.write(json.dumps(rec, sort_keys=True ) + " \n " ) return rec[" hash " ] def verify (self ): prev = GENESIS for i, rec in enumerate (self ._lines()): h = rec.pop(" hash " ) if rec[" prev " ] != prev or hashlib.sha256(json.dumps(rec, sort_keys=True ).encode()).hexdigest() != h: return i prev = h return -1` **Listing 25.2.** Hash-chained audit entries and their verification. code/firm/exptrack/firm_exptrack.py
3. **The model in production on a given date.** `def as_of (self , name, ts, stage=" production " ): """The version that was in `stage` at time ts, from the stage histories.""" best = None for e in self .db.get(name, []): state = None for t, st, _ in e[" history " ]: if t <= ts: state = st if state == stage: best = e return best` **Listing 25.3.** The registry queried as of a time. code/firm/exptrack/firm_exptrack.py
4. **Reproduction.** `def reproduce (record, train): """Re-run training from the record's fields alone and compare the artefact's content hash bit for bit.""" art = train(record[" params " ], record[" seed " ], record[" data " ], record[" code " ]) h = content_hash(pickle.dumps(art)) return h, h == record[" artefact " ]` **Listing 25.4.** Re-running from a record and comparing hashes. code/firm/exptrack/firm_exptrack.py
5. **Run** `ml_track.summary()` , `reproduction()` , `null_max()` and `fig_track.py` (about 45 seconds on one core).

**What to change next.** Estimate the effective number of independent trials from the correlation of the runs’ validation returns, and recompute the deflated Sharpe ratio with it.

## 25.6 Build: experiment tracking

**Purpose.** Every model the firm trades traceable to its run, data and code, and rebuildable from them.

**Interface.** `Tracker(root, log, family)` with `run`, `get`, `artefact`, `runs`, `best`; `Registry(root, audit)` with `register`, `promote`, `current`, `as_of`, `lineage`; `AuditTrail(path)` with `append` and `verify`; `reproduce(record, train)`; on `firm.workflow` and `firm.researchlog`.

**Rules.** Records are written once; artefacts are addressed by content; the clock never enters a record except as a supplied timestamp; promotions carry reasons.

**Acceptance tests.** `code/firm/exptrack/tests/`: identical runs share an artefact; the registry returns the version in production at a past time; promoting a new version retires the old; a tampered audit line is found; reproduction succeeds from a complete record and fails when a field changes.

**Stretch.** Environment capture (package versions, hardware) and a check that refuses to reproduce across a different environment; data snapshots on `firm.pit`.

Sources and further reading

- J. Pineau and co-authors, “Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program)”, arXiv:2003.12206, 2020.
- O. E. Gundersen and S. Kjensmo, “State of the art: reproducibility in artificial intelligence”, *AAAI* , 2018.
- Commission Delegated Regulation (EU) 2017/589 of 19 July 2016 (RTS 6), *Official Journal of the European Union* .

## 25.7 Exercises

**Exercise 25.1 ★.**

List the five fields of a [run record](#def-ml-experiment-tracking-and-reproducibility-tracker) that determine its result, and one thing each can silently change.

**Solution of Exercise 25.1.**

Parameters (a default changes between library versions), seed (the sample of rows and columns each tree sees), data snapshot (a vendor revision), code version (a feature’s definition), environment (library versions, thread counts, hardware changing the order of floating-point sums).

**Exercise 25.2 ★.**

Why is a run’s identifier the hash of its record rather than a counter or a timestamp?

**Solution of Exercise 25.2.**

A hash of the content identifies the run by what it is: two identical records get one identifier, a record cannot be changed without changing its identifier, and nothing depends on when or where it was written, so the tracker itself is reproducible.

**Exercise 25.3 ★.**

What does the deflated Sharpe ratio of 0.15 mean, and why does it differ from the probabilistic Sharpe ratio of 0.96?

**Solution of Exercise 25.3.**

The deflated Sharpe ratio is the probability that the champion’s true Sharpe ratio exceeds the best that 200 skill-less trials would be expected to show: 0.15, weak evidence. The probabilistic Sharpe ratio compares it with zero, as if it were the only trial: 0.96. The difference is the price of the search.

**Exercise 25.4 ★★.**

Why is the expected maximum of 200 null Sharpe ratios too harsh a benchmark for these 200 runs?

**Solution of Exercise 25.4.**

The formula assumes 200 independent trials; the runs share data and differ in a few parameters, so their validation returns are strongly correlated and the effective number of independent trials is much smaller. The expected maximum of fewer trials is lower, and the deflated Sharpe ratio higher; the effective number can be estimated from the runs’ correlations.

**Exercise 25.5 ★★.**

How does a hash chain make an [audit trail](#def-ml-experiment-tracking-and-reproducibility-registry) tamper-evident, and what does it not protect against?

**Solution of Exercise 25.5.**

Each entry includes the previous entry’s hash, and its own hash covers its content: changing or deleting a line changes its hash and breaks the link of every later line. It does not stop someone from rewriting the whole chain from the edited line onwards; for that the latest hash must be anchored somewhere the editor cannot reach (a separate system, a signed daily digest).

**Exercise 25.6 ★★.**

*Find the flaw.* “We save every production model’s weights, so we can always reproduce our models.”

**Solution of Exercise 25.6.**

Weights let you run the model, not rebuild it: without the data snapshot, code, parameters and seed you cannot retrain it, audit how it was chosen, or retrain a corrected version that differs only in the correction. In the chapter every missing field changed the rebuilt model’s test Sharpe ratio by 0.44 to 1.28.

**Exercise 25.7 ★★★.**

*Coding.* Retrain the champion’s configuration with ten different seeds and report the spread of test Sharpe ratios. What does it say about the champion?

**Solution of Exercise 25.7.**

Ten seeds give test Sharpe ratios from $-0.20$ to 2.88, mean 1.03 and standard deviation 0.94. The champion’s configuration is unremarkable once its lucky seed is averaged out: its test Sharpe ratio of 0.91 is near the mean, and the spread is as large as the mean.

**Exercise 25.8 ★★★.**

Design what the registry and the [audit trail](#def-ml-experiment-tracking-and-reproducibility-registry) must contain for a firm to satisfy RTS 6’s record of material changes for a model that is retrained weekly.

**Solution of Exercise 25.8.**

For each weekly version: the [run record](#def-ml-experiment-tracking-and-reproducibility-tracker) (data snapshot, code version, parameters, seed, metrics), who triggered the retrain (the scheduler, under whose authority), who approved promotion and on what evidence, the nature of the change (new data only, or new code or parameters), the times of registration, promotion and retirement; the trail anchored and retained for the period the rules require.

## 25.8 Problem: Reproduce the March Model

**Problem 25.1.**

Weekend problem — the model that traded

The chapter’s 200 runs, registry and trail.

**Part I — The study.**

1. What does each [run record](#def-ml-experiment-tracking-and-reproducibility-tracker) contain?
2. What are the distribution of validation Sharpe ratios and the champion’s figures?
3. What are its probabilistic and deflated Sharpe ratios?
4. What is the expected best of 200 null strategies, and why is it above the champion?

**Part II — Versions and registry.**

5. What changed in the data and the code after the champion was trained?
6. How is the model in production on 10 March found?
7. What does the [audit trail](#def-ml-experiment-tracking-and-reproducibility-registry) record, and how is tampering detected?
8. What does RTS 6 require the firm to record?

**Part III — Reproduction.**

9. What happens when the model is rebuilt from its record?
10. What happens when each field is missing?
11. What does the seed’s row show?
12. When is bitwise reproduction the wrong standard?

**Part IV — The verdict.**

13. State the *named result* : the number of runs behind the champion and its deflated Sharpe ratio, and the field whose absence made reproduction fail.
14. Would you have promoted the champion? What would you have asked for?
15. How would you answer the regulator’s request with these tools, and how long would it take?
16. What else would the tracker need to satisfy a five-year retention rule?
17. How should the tracker count trials that were run but never recorded?
18. What should be in the registry for a model retrained automatically every week?
19. Who should be allowed to promote a model?
20. In one sentence: what does it take to reproduce a model?

**Solution of Problem 25.1.**

**Part I.**

1. Parameters, seed, data snapshot, code version, environment, metrics, artefact hash and timestamp; the identifier is the record’s hash.
2. From $-3.16$ to 2.22, median $-0.83$ ; the champion’s validation 2.22 and test 0.91.
3. 0.96 and 0.15.
4. 3.58, above the champion’s 2.22: the champion is not distinguishable from the luck of 200 independent tries (the benchmark is conservative for correlated runs).

**Part II.**

1. The vendor revised 1% of returns (snapshot 2026-03), and the feature clipping moved from 3 to 2.5 standard deviations (v2).
2. From the registry’s stage histories, as of 10 March 08:00: version 1, promoted on 2 March.
3. Registration and promotion with times and reasons, each entry hash-linked to the previous; editing a line breaks the chain.
4. Records of every material change to algorithmic trading software (when, by whom, approved by whom, what); order records of high-frequency traders kept five years.

**Part III.**

1. The rebuilt model’s dump has the stored artefact’s content hash: bit for bit the same model.
2. Every missing field breaks reproduction: default parameters $-1.28$ , seed $+0.85$ , latest data $-1.25$ , current code $-0.44$ in test Sharpe ratio.
3. That seed noise alone moves the result by as much as the choice of configuration.
4. Across library versions, thread counts or hardware, where floating-point order changes; there the standard is the same metrics within the seed spread.

**Part IV.**

1. *Reproduce the March model.* Two hundred runs stand behind the champion, whose validation Sharpe ratio of 2.22 has a deflated Sharpe ratio of 0.15; reproduction from the registry is exact, and it fails with any one of parameters, seed, data snapshot or code version missing.
2. Not on this evidence: ask for the seed spread, the effective number of trials, and a longer validation.
3. Query the registry as of the date, follow the lineage to the run, rebuild it, and list all 200 trials from the tracker: minutes, not weeks.
4. Retention and immutability of records and artefacts for the period, and anchoring of the [audit trail](#def-ml-experiment-tracking-and-reproducibility-registry) .
5. As a lower bound; the true count is higher, and the deflated Sharpe ratio is an upper bound.
6. Each retrain as a version with its snapshot, an automatic promotion rule recorded as the approver, and alerts when the rule is overridden.
7. A person other than the model’s author, under the firm’s model-risk [policy](https://one-course.com/books/quant/12/en/chapter/17-reinforcement-learning-foundations#def-ml-reinforcement-learning-foundations-mdp) (Book 6, chapter 26).
8. The run’s full record and the same software.

## 25.9 Interview questions

**Interview question 25.1 ★ mle.**

What does an [experiment tracker](#def-ml-experiment-tracking-and-reproducibility-tracker) record, and why record the failed runs?

**Solution of Interview question 25.1.**

Parameters, seeds, data and code versions, environment, metrics and artefacts, for every run. Failed runs count as trials: without them the multiple-testing correction is impossible and the chosen result looks better than it is.

*What the interviewer is looking for: the fields, and failed runs as trials.*

**Interview question 25.2 ★★ mle, developer.**

How do you make a training run reproducible bit for bit?

**Solution of Interview question 25.2.**

Fix and record seeds, use deterministic algorithms and single-threaded or fixed-order reductions, pin library versions, version data and code, and verify by rebuilding and comparing content hashes.

*What the interviewer is looking for: determinism settings, pinned versions, and a hash comparison.*

**Interview question 25.3 ★★ researcher.**

How does the number of trials change what a backtest’s Sharpe ratio means?

**Solution of Interview question 25.3.**

The best of many trials is biased upwards; the deflated Sharpe ratio compares it with the expected maximum of that many null trials, so the count (and the trials’ correlation) decides how impressive a Sharpe ratio is.

*What the interviewer is looking for: selection bias and the deflated Sharpe ratio.*

**Interview question 25.4 ★★ mle.**

Design a [model registry](#def-ml-experiment-tracking-and-reproducibility-registry) for a trading firm. What stages and fields does it need?

**Solution of Interview question 25.4.**

Name, version, stage (candidate, production, retired), lineage (run, data, code), metrics, owner and approver, stage history with reasons, links to the [model card](https://one-course.com/books/quant/12/en/chapter/21-interpretability-and-model-governance#def-ml-interpretability-and-model-governance-card) and validation report, and an append-only [audit trail](#def-ml-experiment-tracking-and-reproducibility-registry) of changes.

*What the interviewer is looking for: stages, lineage, approvals and history.*

**Interview question 25.5 ★★ mle, developer.**

Why version data, and how would you version a dataset that is corrected after the fact?

**Solution of Interview question 25.5.**

Because the same name comes to mean different data; corrections create new immutable snapshots (content-hashed or dated) and runs record which they used, so old results stay reproducible and new ones can be compared.

*What the interviewer is looking for: immutability and snapshot identifiers.*

**Interview question 25.6 ★★★ mle.**

Your production model cannot be reproduced. Walk through the investigation.

**Solution of Interview question 25.6.**

Start from the registry’s lineage: is the [run record](#def-ml-experiment-tracking-and-reproducibility-tracker) complete, is the data snapshot still available and unchanged (hash), the code version checked out, the environment the same; rebuild and compare hashes; diff intermediate outputs stage by stage (Book 7’s pipeline manifests) to find where they diverge.

*What the interviewer is looking for: a systematic walk from lineage to the first diverging stage.*
