Research, Data and Risk Platforms · Technology
27Testing and Continuous Delivery
In the last days of July 2012 a broker-dealer deployed new order-routing code to its eight routing servers, in stages. The new code reused a flag that had once switched on a function retired years before and never deleted. A technician did not copy the new code to one of the eight servers; no second person reviewed the deployment, and no written procedure required one. On 1 August orders carrying the reused flag reached the eighth server, woke the retired function, and in forty-five minutes the firm sent millions of orders and lost more than $460 million. Book 11 reads the episode for its missing risk limits; this chapter reads it for its deployment. Every piece of it — the untested old code, the reused flag, the partial deployment, the missing check — is something a test pipeline and a deployment procedure exist to catch.
27.1 The test pyramid for trading code
Definition 27.1 (Unit, integration and end-to-end tests)
A unit test checks one function or class in isolation, in milliseconds. An integration test checks several components working together — a pricer with its market data, a strategy with the simulator — in seconds. An end-to-end test runs a whole flow as production would — orders from a strategy through the gateway to a simulated venue and back to the position service — in minutes.
Definition 27.2 (Test pyramid)
The test pyramid is the rule that a code base should have many unit tests, fewer integration tests and few end-to-end tests, because each level up is slower, more fragile and less precise about what broke; it became widely known through Mike Cohn’s 2009 book Succeeding with Agile.
Trading code adds levels of its own. Book 13, chapter 25, describes property-based, differential and fuzz tests, performance regression gates and exchange certification for the low-latency path; chapter 12’s parity harness is an integration test between the simulator and the production adapter; Book 6’s model validation is a review, not a test, but it produces the reference values tests use. This chapter adds the test a pricing library needs most, and the controls that decide what reaches production.
27.2 Golden files for pricers
Definition 27.3 (Golden-file test)
A golden-file test compares a program’s outputs with reviewed values stored with the code — the golden file — each with a tolerance and the provenance of the value; any difference beyond the tolerance fails the test, and changing a golden value is itself a reviewed change.
A pricing library’s outputs are numbers without an independent oracle for most products: a closed form covers the simplest, and Book 5’s reference pricers cover some more. What a golden file adds is memory: the value the library gave yesterday, reviewed then, against which today’s build is compared. The JPMorgan task force report on the 2012 losses in the Chief Investment Office found that the new VaR model ran through spreadsheets filled by copying and pasting, and that one of them divided by the sum of two rates instead of their average, which likely muted volatility by a factor of two. A golden value for the VaR of a reference portfolio, reviewed once, would have moved when the model changed.
The chapter’s portfolio has sixty instruments priced by Book 5’s library: twenty European options (analytic engine), ten American puts (finite differences), twenty Asian calls (Monte Carlo, 20 000 paths) and ten down-and-out calls (closed form). The tolerance is the engine’s own error times a factor : a relative for the closed forms, times the gap between the 200-point and the 400-point grid for the PDE, times the Monte Carlo standard error — estimated from eight independent seeds — for Monte Carlo.
fig_goldtest.py.def tolerance(key: str, inst, eng, k: float) -> float:
"""The engine's own error, times k (closed forms: a relative 1e-10)."""
if isinstance(eng, FP.MonteCarloEngine):
xs = [pv(inst, dataclasses.replace(eng, seed_offset=100 + j))
for j in range(N_SE_SEEDS)]
return k * float(np.std(xs, ddof=1))
if isinstance(eng, FP.PDEEngine):
return k * abs(pv(inst, eng) - pv(inst, FP.PDEEngine(m=400, n_t=400)))
return 1e-10 * max(1.0, abs(pv(inst, eng)))
def compare(store: GoldenStore, results: dict, product_of=None) -> Report:
"""Each result against its golden value; a key missing from either side is a failure."""
product_of = product_of or (lambda k: k.split(":")[0])
fails, by = [], {}
for k in sorted(set(store.items) | set(results)):
if k not in store.items or k not in results:
fails.append((k, math.inf, 0.0))
else:
g = store.items[k]
diff = results[k] - g.value
if not abs(diff) <= g.tol:
fails.append((k, diff, g.tol))
by[product_of(k)] = by.get(product_of(k), 0) + 1
return Report(fails, by, len(results))
Five builds are then compared with the golden file. Each plants one change in the library’s behaviour.
def build(change: str, key: str, inst, eng, seed: int = 1) -> float:
"""The new build's price of one instrument, for each planted change."""
if change == "seed" and isinstance(eng, FP.MonteCarloEngine):
return pv(inst, dataclasses.replace(eng, seed_offset=seed))
if change == "day count":
kw = {"expiry": _stretch(inst.expiry)}
if isinstance(inst, FP.AsianOption):
kw["fixing_dates"] = tuple(_stretch(x) for x in inst.fixing_dates)
return pv(dataclasses.replace(inst, **kw), eng)
if change == "put sign" and getattr(inst, "right", "") == "P":
return pv(inst, eng, market(rate=-0.03))
if change == "Asian fixing" and isinstance(inst, FP.AsianOption):
return pv(dataclasses.replace(inst, fixing_dates=inst.fixing_dates[:-1]), eng)
if change == "finer grid" and isinstance(eng, FP.PDEEngine):
return pv(inst, FP.PDEEngine(m=400, n_t=400))
return pv(inst, eng)
fig_goldtest.py.The two large regressions are caught at every tolerance: the sign error changes all twenty puts by hundreds of error units and more, and the day count changes every European, American and barrier price by far more than its tolerance. The trouble is Monte Carlo (Figure 27.2). With the seed free to change, the golden test faces two estimates of the same price and must tolerate their difference. At noise alone flags 18% of the Asian calls and fails every one of forty runs; at it fails 2.5% of runs, but the Asian-fixing bug — a small change, two to ten standard errors — is caught on only 9 of the 20 Asian calls, and the day count’s effect on them, never more than two standard errors, on none. With a false alarm costing $1 000 of investigation and a missed single-product regression $500 000 — the model’s assumptions — the cheapest tolerance is , a test that fails on every run. No tolerance works.
The fix is not a tolerance but a seed. Book 5’s Monte Carlo engine draws its random numbers from a seed fixed per instrument; if the build pins it, today’s Monte Carlo price is a deterministic function of the code, the same random numbers give the same price to the last digit, and the Monte Carlo instruments can be compared at a closed form’s tolerance. Pinned, the golden test catches the Asian-fixing bug on all 20 instruments and the day count on all 60. A new seed then fails the test like any other change — which it is — and is released as an approved update of the golden values. The same holds for the finer PDE grid: it moves each American put by exactly one error unit, passes every , and should reach production as a reviewed golden update with the new grid in its provenance, not as a silent pass.
27.3 Continuous integration
Definition 27.4 (Continuous integration, deployment pipeline)
Continuous integration merges every change into the main line at least daily, and builds and tests the merged code automatically on each merge. A deployment pipeline is the automated sequence of stages — build, unit tests, integration and golden tests, end-to-end tests, deployment — that every change passes through, in order, stopping at the first failure.
Definition 27.5 (Continuous delivery)
Continuous delivery keeps the main line releasable at all times, so that any change that passes the pipeline can be deployed to production by a decision rather than a project.
bench_goldtest.py). Each stage runs only if the one before passed; the last check before the traffic switch is the consistency of every host with the release.The pipeline’s order is the pyramid’s: cheap tests first, so that most failures are found in seconds. On the chapter’s portfolio, the unit tier (put–call parity on a thousand options) takes 0.02 seconds, the golden tier (the sixty instruments against their pinned values) 0.33 and the end-to-end tier (the portfolio revalued in five scenarios) 1.62, measured on an otherwise idle machine. The golden file is small enough to run on every merge; a firm’s real one, with thousands of instruments, runs in the nightly integration build and on every change to the library.
27.4 Release trains and change control
Definition 27.6 (Release train)
A release train deploys, at fixed times, everything that has passed the pipeline since the previous departure; a change that misses a train waits for the next.
Definition 27.7 (Change control)
Change control is the procedure every production change follows — recorded, reviewed, approved by someone other than its author, deployed by a stated method, verified — so that the firm can say for any change who made it, who approved it and when.
The train’s cadence sets a trade-off (Table 27.1). Over a model year of 533 changes, deploying each change as it passes the pipeline gives no wait and one change per deployment; a daily train makes changes wait half a day on average and ships 1.8 at a time; a weekly train, 3.5 days and 10.3 changes; a fortnightly one, 7 days and 20.5. The batch matters when a deployment goes wrong: finding the bad change among ten takes about four bisection steps, among twenty about five, and each step is a rollback, a redeployment and a wait for the symptom. Trading firms usually settle between the two: small changes often for research and analytics, trains with a change window for the order path, where RTS 6 requires a clearly delineated methodology to develop and test a system before its deployment or substantial update, in an environment separated from production.
| cadence | mean wait (days) | changes per deployment | bisection steps |
|---|---|---|---|
| each change at once | 0.00 | 1.00 | 0.00 |
| daily train | 0.51 | 1.83 | 1.16 |
| weekly train | 3.46 | 10.25 | 3.93 |
| fortnightly train | 6.99 | 20.50 | 4.86 |
27.5 Deploying and rolling back
Definition 27.8 (Blue–green deployment, rollback)
A blue–green deployment installs a release on a second, idle set of hosts (green) while the current set (blue) serves, then switches traffic to green at once; switching back is the rollback, the return to the previous release, which a blue–green deployment makes a switch rather than a reinstallation.
The hook’s deployment failed in the step between installation and switch: seven hosts ran the new release, one ran the old, and nothing compared them. A deployment check is simple to write. Each host reports the hash of the build it runs and of its configuration, and the release is refused — the traffic switch does not happen — unless every host reports the release’s hashes. Run on the hook’s eight servers, it names the eighth before the market opens.
def host_state(build: dict, config: dict) -> tuple[str, str]:
"""What a host reports: the hash of the build it runs and of its configuration."""
return _h(build), _h(config)
def consistency(hosts: dict, release: tuple) -> list[str]:
"""Hosts whose build or configuration differs from the release's: the release is refused
unless the list is empty."""
return sorted(h for h, state in hosts.items() if state != release)
The order also lists what else was missing: the retired code was still present and callable, and the flag it had once used was given a new meaning. Both are testing failures in this chapter’s sense. Dead code should be deleted, not left behind a flag; a flag should never be reused, since an old host or an old message will still give it the old meaning; and a feature flag (chapter 16) switched on by a release is a deployment of its own, with its own staged rollout and its own check that every host agrees on what it means.
As of September 2026 — The rules on testing and deployment
The SEC’s order of 16 October 2013 in the matter of Knight Capital Americas LLC (Release 34-70694) found that new code repurposed a flag that had activated a retired function still present on the servers; that a technician did not copy the new code to one of eight servers; that no second technician reviewed the deployment and no written procedure required it; and that orders routed over forty-five minutes on 1 August 2012 cost the firm more than $460 million. In the European Union, Commission Delegated Regulation (EU) 2017/589 (RTS 6) requires an investment firm to establish clearly delineated methodologies to develop and test an algorithmic trading system before its deployment or substantial update, ensuring among other things that it does not behave in an unintended manner, and to test in an environment separated from production (articles 5 and 7).
27.6 Tutorial: seven of eight servers
Goal. Freeze a pricing library’s golden values, measure what each tolerance detects and falsely flags, and refuse a partial deployment. End state: Figure 27.2 and Table 27.1.
- Portfolio:
pl_goldtest.portfolio(); the error unit of each instrument witherrors. - Builds:
experiment(), the five planted builds and forty noise builds. - Tolerance:
rates(ex, k)for each , then withpinned=True;expected_cost(ex, k). - Pipeline:
bench_goldtest.pyfor the stage times;trains()for the cadences. - Deployment:
eight_servers()andfirm_goldtest.consistency.
What to change next. Store the golden file on disk with GoldenStore.save and require an approval for the grid refinement; add a check that no two releases give one flag two meanings.
27.7 Build: golden tests and deployment checks
Purpose. Know, for every build of the pricing library, which values moved and by how much, and never let a release reach a host set that does not all run it.
Interface. Golden, GoldenStore (put, get, propose, approve, save, load), compare, run_tiers, release_train, host_state, consistency.
Rules. Tolerances derived from engine error, not chosen; Monte Carlo seeds pinned in the build; golden updates proposed with a reason and approved by someone else; tiers in pyramid order, stopping at the first failure; no release switched on until every host matches it.
Acceptance tests. code/firm/goldtest/tests/: comparison with a failure, a missing value and a pass; approval by another person and persistence; tiers stopping at a failure; release-train waits, batches and bisection; consistency with a stale build and a drifted configuration.
Stretch. A change-impact report by risk factor; golden values of risk (Greeks, VaR) as well as prices; blue–green switching in the simulator.
Sources and further reading
- US Securities and Exchange Commission, In the Matter of Knight Capital Americas LLC, Release No. 34-70694, 16 October 2013.
- Report of JPMorgan Chase & Co. Management Task Force Regarding 2012 CIO Losses, 16 January 2013.
- M. Fowler, TestPyramid, 2012 (on M. Cohn, Succeeding with Agile, 2009).
- Commission Delegated Regulation (EU) 2017/589 (RTS 6), articles 5 and 7.
- One Quant Book 5, chapter 28 (reference pricers); Book 11, chapter 27 (risk controls); Book 13, chapter 25 (testing low-latency code).
27.8 Exercises
Exercise 27.1 ★
Why is the tolerance of a closed-form price a relative and not zero?
Solution
Solution of Exercise 27.1.
Because the same closed form computed on another processor, library version or compiler can differ in the last bits of a floating-point number; a relative is far below any economic difference and far above rounding.
Exercise 27.2 ★
With the seed free, what share of Asian calls does noise alone flag at , and what does theory predict?
Solution
Solution of Exercise 27.2.
4.3% of Asian calls across forty runs at , against 3.4% in theory: two independent estimates differ by more than standard errors with probability . The excess comes from the standard error itself being estimated from only eight seeds.
Exercise 27.3 ★
A weekly train carries ten changes and one is bad. About how many bisection steps does it take to find it?
Solution
Solution of Exercise 27.3.
About , rounded up: four steps, each a rollback of half the remaining changes and a wait for the symptom. The model year’s weekly train averages 3.93.
Exercise 27.4 ★★
Why does the finer PDE grid pass the golden test for every , and why should it still be approved?
Solution
Solution of Exercise 27.4.
Its tolerance is times the gap between the 200-point and 400-point grids, and the new build moves each price by exactly that gap: one error unit. It passes, yet it changes every American price the desk reports, so it should be released as an approved golden update, with the new grid in the provenance, not slip through.
Exercise 27.5 ★★
Why can the day count’s effect on the Asian calls not be seen with the seed free, and how is it seen with the seed pinned?
Solution
Solution of Exercise 27.5.
Its effect on each Asian call is at most two standard errors, smaller than the noise between two seeds, so no tolerance that tolerates the noise can see it. With the seed pinned, the old and new builds use the same random numbers, the difference is the effect alone, and a closed form’s tolerance catches it on all twenty.
Exercise 27.6 ★★
What would a blue–green deployment have changed on the hook’s morning, and what would it not have changed?
Solution
Solution of Exercise 27.6.
The switch would have been one action at one moment, and switching back a second action rather than a reinstallation; had the eighth green host been left on the old build, the switch would still have sent it traffic. It changes how fast a bad release is undone, not whether a partial installation is noticed: that is the consistency check’s job.
Exercise 27.7 ★★★
Coding. Double the Monte Carlo paths of the portfolio and rerun rates(ex, 5). What happens to the Asian-fixing detection and why?
Solution
Solution of Exercise 27.7.
The standard error falls by while the bug’s effect does not, so the bug is caught on 14 of 20 Asian calls at instead of 9. Noise, however, failed 40% of the runs: with the standard error estimated from eight seeds, some instruments’ tolerances are too tight by chance. More paths help detection; only a pinned seed removes the noise.
Exercise 27.8 ★★★
Find the flaw. "The golden test failed again on Monte Carlo noise, so we widened every tolerance to ten standard errors."
Solution
Solution of Exercise 27.8.
At ten standard errors the Monte Carlo instruments no longer see any regression smaller than ten standard errors — the chapter’s Asian-fixing bug is two to ten — and the closed forms gain nothing. Noise was the symptom of an unpinned seed; the fix is to pin it and release seed changes as approved golden updates.
27.9 Problem: Seven of Eight Servers
Problem 27.1
Weekend problem — seven of eight servers
The chapter’s portfolio, builds, pipeline, trains and eight routers.
Part I — Tests.
- What does each level of the pyramid test, and why are there fewer tests at each level up?
- What does a golden file hold for each value?
- How is each engine’s tolerance derived?
- What did the JPMorgan task force report find in the VaR spreadsheets?
- Which builds change which instruments?
Part II — Tolerance.
- Which regressions are caught at every tolerance, and why?
- What does noise alone do at and at ?
- What share of the Asian-fixing bug is caught at each ?
- Which minimises the model’s expected cost, and what is wrong with it?
- What does pinning the seed change?
Part III — Delivery.
- How long does each tier of the pipeline take?
- What do the four release cadences give?
- What does change control record?
- What does RTS 6 require before a trading algorithm is deployed?
- What is a rollback in a blue–green deployment?
Part IV — The verdict.
- State the named result: the detection rate of each planted regression and the false-alarm rate of Monte Carlo noise at each tolerance, the tolerance that minimises expected cost, and the consistency check that would have refused the partial deployment.
- What four things went wrong in the hook’s deployment, in this chapter’s terms?
- Which one would the consistency check have caught, and which would testing have caught?
- What would you require of every release of the order path?
- In one sentence: what is a deployment?
Solution
Solution of Problem 27.1.
- Units test functions in isolation, integration tests components together, end-to-end tests whole flows; each level up is slower, more fragile and less precise, so there are fewer of them.
- The value, its tolerance and its provenance (engine, parameters, library version, who approved it).
- Closed forms: a relative ; PDE: times the 200 against 400-point grid gap; Monte Carlo: times the standard error from eight seeds.
- Spreadsheets filled by copying and pasting, and a division by the sum of two rates instead of their average that likely muted volatility by a factor of two.
- The seed: the twenty Asian calls; the day count: all sixty; the put sign: the twenty puts; the Asian fixing: the twenty Asian calls; the finer grid: the ten American puts.
- The put sign error and the day count on every non-Monte-Carlo instrument: hundreds of error units or more.
- At it flags 18% of Asian calls and fails every run; at , 0.1% of calls and 2.5% of runs.
- All 20 at , 18 at 2.5 and 3, 15 at 3.5, 11 at 4, 9 at 5, 7 at 6.
- , at $980 per run: a test that fails on every run, which nobody will keep looking at.
- Monte Carlo values become deterministic: compared at a closed form’s tolerance, every planted regression is caught (20 of 20 Asian fixings, 60 of 60 day counts) and noise cannot occur; a new seed is a change to approve.
- Unit 0.02 s, golden 0.33 s, end to end 1.62 s (measured, otherwise idle machine).
- Mean waits of 0, 0.51, 3.46 and 6.99 days; 1, 1.83, 10.25 and 20.5 changes per deployment; 0, 1.16, 3.93 and 4.86 bisection steps.
- Who made each change, who approved it, when, how it was deployed and verified.
- Clearly delineated methodologies to develop and test it, ensuring among other things that it does not behave in an unintended manner, with testing in an environment separated from production.
- Switching traffic back to the previous, still-installed set of hosts.
- Named result. The sign error (20 of 20 puts) and the day count (every non-Monte-Carlo instrument) are caught at every tolerance; the Asian-fixing bug is caught on 20, 18, 11, 9 and 7 of 20 Asian calls at = 2, 3, 4, 5, 6 while noise fails 100%, 60%, 20%, 2.5% and 0% of runs; the model’s expected cost is lowest at ($980 a run) only by failing every run — pinned seeds catch everything with no noise. The consistency check refuses the release because router8 reports another build’s hash.
- Dead code left callable, a flag reused, an incomplete deployment, and no review or check of the deployment.
- The check catches the incomplete deployment; tests of the old code path with the reused flag would have shown what the flag did on an old host.
- Pinned golden and integration tests, a staged rollout with a consistency check before each switch, a tested rollback, and change control with a second person.
- The moment every host starts running the same, tested release — verified, not assumed.
27.10 Interview questions
Interview question 27.1 ★ developer
What is the test pyramid, and why is it a pyramid?
Solution
Solution of Interview question 27.1.
Many fast unit tests, fewer integration tests, few end-to-end tests: each level up is slower, more fragile and says less precisely what broke, so the bulk of the checking belongs at the bottom.
What the interviewer is looking for: Cost and precision per level.
Interview question 27.2 ★★ developer
How do you regression-test a Monte Carlo pricer?
Solution
Solution of Interview question 27.2.
Pin the seed so that the same code gives the same numbers, keep golden values with tight tolerances, release seed or path-count changes as approved golden updates, and separately test convergence and unbiasedness against closed forms with statistical tolerances.
What the interviewer is looking for: Pinned seeds; statistics only where they belong.
Interview question 27.3 ★★ developer
A release reached seven of eight servers. How would your deployment process have stopped it?
Solution
Solution of Interview question 27.3.
Every host reports its build and configuration hashes; the release is switched on only if all match; deployment is automated, reviewed by a second person, staged, and reversible.
What the interviewer is looking for: Verify every host before switching.
Interview question 27.4 ★★ developer
Would you deploy trading code continuously or on a release train? Why?
Solution
Solution of Interview question 27.4.
Small, frequent, automated deployments for research and analytics; trains with a change window, staged rollout and consistency checks for the order path, where a bad release costs money in minutes and regulation requires documented testing.
What the interviewer is looking for: Batch size against blast radius.
Interview question 27.5 ★★★ developer
Design the test and release process of a firm’s pricing library and order path.
Solution
Solution of Interview question 27.5.
Unit and property tests on every merge; pinned golden values of prices and risk with engine-derived tolerances and approved updates; simulator parity and end-to-end tests nightly; a pipeline that stops at the first failure; trains or continuous delivery by component; blue–green deployment with a consistency check and a tested rollback; change control recording author, approver and time.
What the interviewer is looking for: Tests, pinning, controlled deployment.