Quantitative Finance · Book 14 · Technology

Networks, Hardware and Trading Infrastructure

Networks, Hardware and Trading Infrastructure · Technology

28Disaster Recovery and Regulatory Requirements

When Superstorm Sandy reached New York in October 2012, the U.S. national securities exchanges “jointly decided not to open for trading on October 29 and October 30”. The securities industry’s annual test of trading from backup sites had run on October 27, two days before, and found nothing that would have stopped the markets from opening on their backup systems; many firms, the SEC later wrote, “remained uncomfortable about switching over”. Two years later Regulation SCI required the exchanges’ recovery plans to be reasonably designed to resume their critical systems within two hours of a wide-scale disruption, and required the members each venue designates to test its backup site with it at least once a year.

Book 13 described failover, primary–backup replication and the split brain; Book 6 operational risk; Book 11 the kill switch; Book 15 incident response. This chapter is about sites and rules: where a firm puts its recovery site, how exchanges test theirs with their members, which resilience rules a trading firm is examined on, and how a recovery design turns into a recovery time and a recovery point. The recovery simulation is labelled as such; the rules are quoted.

28.1 Secondary sites and how they are placed

Definition 28.1 (Primary data centre, disaster-recovery site)

A firm’s primary data centre is the site from which it normally trades; its disaster-recovery site is a second site, far enough away not to share the primary’s likely disasters, from which the firm can resume trading when the primary is lost.

Two distances pull against each other. The recovery site must be outside the hazards that could take the primary, a storm, a flood, a regional power failure; and it must be close enough for its replication to be cheap and its connections to the venues usable. The exchanges’ own choices show the scale: CME’s U.S. BrokerTec markets run in Secaucus and recover in Aurora, its European ones run in Slough and recover in Secaucus; LSEG’s spot matching runs in Slough and recovers in Secaucus (chapters 23 to 25). Recovery across a continent or an ocean, not across the river.

Candidate recovery siteFrom Secaucus NY5 (km)Outside a 100 km hazard radius
Secaucus NY4 (same campus)0.3no
Carteret26.0no
Mahwah34.2no
Aurora1 191.1yes
Table 28.1. Candidate recovery sites for a firm trading from Secaucus NY5: geodesic separation (chapter 10’s coordinates) against a hazard radius of 100 km100\,\mathrm{k}\mathrm{m} (an assumption). Data: nw_dr.separations().

The New Jersey sites are all inside any storm that closes Secaucus. Aurora is not, and costs what distance costs: synchronous replication to it would add 17.4 ms17.4\,\mathrm{m}\mathrm{s} of round trip to every write at a route factor of 1.5, against 0.5 ms0.5\,\mathrm{m}\mathrm{s} to Mahwah, so the firm replicates asynchronously and accepts a recovery point.

28.2 Exchange failover tests

Definition 28.2 (Business continuity plan, exchange failover test)

A business continuity plan is a firm’s documented set of arrangements, people and procedures for continuing or resuming its critical activities after a disruption. An exchange failover test is a scheduled exercise in which an exchange runs its markets from its recovery site or backup components and its members connect, trade and verify from theirs.

Tests exist because Sandy showed that a tested backup is not a trusted one. Regulation SCI makes them mandatory for the members an exchange designates, at least once every twelve months, coordinated across the industry. The futures industry runs its own: FIA’s 2025 test, on Saturday 25 October with a pre-test in September, had more than a hundred organisations verify round-trip connectivity and process recovery between primary and backup sites (Box 28.1). Exchanges also fail over routinely: CME’s market segment gateways switch to their backup components and back in weekly windows.

As of September 2026 — Failover tests and exchange recovery sites

FIA disaster recovery test: 25 October 2025 (pre-test 27 September 2025), more than 100 participants. CME market segment gateways: weekly failover windows to backup components. BrokerTec U.S.: NY5, recovering in Aurora, 4-hour recovery time objective; BrokerTec EU: LD4.2, recovering in NY5, 2 hours. LSEG spot matching: LD4, recovering in NY6.

28.3 The resilience rules firms are examined on

Definition 28.3 (Operational resilience, important business service, impact tolerance)

Operational resilience is a firm’s ability to prevent, adapt to, respond to, recover from and learn from operational disruptions. An important business service is a service the firm provides to clients or the market whose disruption could cause them intolerable harm or threaten the firm or market integrity; its impact tolerance is the maximum tolerable disruption to it, set by the firm, usually as a time.

As of September 2026 — Resilience rules, from the regulators’ texts

  • Regulation SCI (U.S. exchanges and other SCI entities): backup capabilities “sufficiently resilient and geographically diverse” and “reasonably designed to achieve next business day resumption of trading and two-hour resumption of critical SCI systems following a wide-scale disruption”; designated members test with the entity at least once every 12 months.
  • RTS 6, Article 14 (algorithmic trading firms in the EU): documented business continuity arrangements per venue, scenarios including the loss of data centres and suppliers, relocation to a back-up site where appropriate, orderly shutdown; reviewed and tested annually.
  • FCA PS21/3 (UK): important business services and impact tolerances set by 31 March 2022; able to remain within them by 31 March 2025. DORA (EU, Regulation 2022/2554): applies from 17 January 2025.

The rules differ by who they bind but share a shape. The exchange must resume its critical systems within two hours; the trading firm must have arrangements to resume or shut down in order, tested every year; and in the United Kingdom the firm itself names the services that matter and the disruption it will tolerate, then shows it can stay within it. A trading firm’s recovery design is therefore judged on two numbers it sets itself, a recovery time and a recovery point, against the tolerance it declared.

28.4 Recovery objectives and runbooks

Definition 28.4 (Recovery time objective, recovery point objective)

A recovery time objective is the longest a system may take, from its loss, to be back in service; a recovery point objective is the most recent data the firm may lose, measured as the time between the last change the recovery site holds and the moment of loss.

A recovery is a runbook: detect the loss and decide to fail over, switch the sessions to the recovery addresses, restore or start what is not running, load positions and reference data, reconnect and verify sessions, reconcile orders and positions against the exchanges’ drop copies, resume. Each step has a duration that varies; the recovery time is their sum.

def rto_samples(plan, n=100000, seed=0):
    rng = np.random.default_rng(seed)
    total = np.zeros(n)
    for s in plan.steps:
        total += s.median_min * np.exp(s.sigma * rng.standard_normal(n))
    return total


def p_within(plan, target_min, n=100000, seed=0):
    return float(np.mean(rto_samples(plan, n, seed) <= target_min))


def rpo(plan, table):
    if plan.replication == "sync":
        return 0.0
    if plan.replication == "async":
        d = separation_km(plan.primary, plan.recovery, table) * 1000.0
        return (plan.lag_ms + gm.floor_us(d, "fibre") / 1000.0) / 1000.0
    return plan.snapshot_min * 60.0


def sync_penalty_ms(plan, table, factor=1.5):
    d = separation_km(plan.primary, plan.recovery, table) * 1000.0
    return 2 * gm.floor_us(d, "fibre") * factor / 1000.0
Listing 28.1. Recovery time as the sum of a runbook’s steps, the recovery point by replication mode, and the price of synchronous replication. code/firm/drplan/firm_drplan.py

Proposition 28.5 (What sets each objective)

The recovery point depends only on replication: zero if synchronous, the replication lag plus the one-way delay if asynchronous, up to the snapshot interval if restored from snapshots. The recovery time depends only on the runbook: it is at least the sum of the steps’ minimum durations, whatever the replication.

Proof. Data the recovery site does not hold when the primary is lost is lost; how much is set by when it was sent, which is the replication. The runbook’s steps run in sequence after the loss, and none can take less than its minimum, so neither objective can be bought with the other’s tool. ∎

The chapter’s firm, trading from Secaucus with a recovery site in Aurora, compares two designs (Figure 28.1). The hot design keeps servers running in Aurora with asynchronous replication (a lag of 50 ms50\,\mathrm{m}\mathrm{s}) and sessions ready: its runbook is five steps, with medians of 15, 10, 20, 10 and 5 minutes. The warm design keeps the site and a few servers, restores from snapshots taken every fifteen minutes, and has six steps, including 45 minutes to start and restore and 30 to load data. All durations are assumptions, lognormal around their medians.

HOT = dr.Plan("hot", PRIMARY, "aurora",
              (S("detect and decide", 15, 0.5), S("switch sessions to recovery addresses", 10, 0.4),
               S("reconcile orders and positions", 20, 0.5), S("check venue connectivity", 10, 0.4),
               S("resume trading", 5, 0.3)),
              "async", lag_ms=50.0, annual_cost=0.70 * PRIMARY_COST)
WARM = dr.Plan("warm", PRIMARY, "aurora",
               (S("detect and decide", 15, 0.5), S("start servers and restore", 45, 0.6),
                S("load reference data and positions", 30, 0.5),
                S("reconnect and verify sessions", 20, 0.5),
                S("reconcile orders and positions", 20, 0.5), S("resume trading", 5, 0.3)),
               "snapshot", snapshot_min=15.0, annual_cost=0.25 * PRIMARY_COST)
Listing 28.2. The two designs: runbooks, replication and yearly cost. code/networks/28-disaster-recovery-and-regulatory-requirements/python/nw_dr.py
Simulation: the recovery-time distributions of the hot and warm designs, against the two-hour line. The hot design resumes within two hours in 99.2% of simulated recoveries (median 64 minutes); the warm one in 20.3% (median 148). Data: fig_dr.py, nw_dr.rto_hist().
Figure 28.1. Simulation: the recovery-time distributions of the hot and warm designs, against the two-hour line. The hot design resumes within two hours in 99.2% of simulated recoveries (median 64 minutes); the warm one in 20.3% (median 148). Data: fig_dr.py, nw_dr.rto_hist().

The hot design resumes within two hours in 99.2% of simulated recoveries, with a median of 64 minutes and a 95th percentile of 96; its recovery point is 56 ms56\,\mathrm{m}\mathrm{s}, the lag plus the one-way delay to Aurora. The warm design meets two hours in 20.3%, with a median of 148 minutes and a 95th percentile of 237, and can lose fifteen minutes of data. If the hot site costs 70% of the primary’s yearly infrastructure and the warm one 25% (chapter 27’s example stack, USD 3.19 million a year; assumptions), the difference is USD 1.44 million a year: the price of the two-hour line for this firm.

Method 28.6 (Designing a recovery for a trading firm)

  1. Name the important business services and set an impact tolerance for each; derive the recovery time and point objectives.
  2. Choose a recovery site outside the primary’s hazards and near the venues’ own recovery sites; compute what synchronous replication would cost in latency, and replicate asynchronously if it is too much.
  3. Write the runbook as timed steps; measure each step in tests, not in estimates.
  4. Simulate the recovery time against the objective; buy down the steps that dominate, or change design.
  5. Take part in the exchanges’ and the industry’s failover tests, and run your own at least once a year; record the results as evidence for the regulator.

28.5 Tutorial: the two-hour rule

Goal. Place a recovery site, simulate the hot and warm designs, and compare them with the two-hour line and their cost. End state: Figure 28.1 and Table 28.1.

  1. Sites. firm_drplan.separation_km and outside_hazard on chapter 10’s coordinates.
  2. Plans. Plan and Step (Listing 28.2); rto_samples, p_within and rpo (Listing 28.1).
  3. Calendar. load_calendar and tested_within for the annual test.

What to change next. Let some steps run in parallel (reconciliation while sessions reconnect) and see which design’s median moves; add the chance that the decision to fail over is delayed, as it was before Sandy.

28.6 Build: the disaster-recovery plan

Purpose. A recovery plan as data with its recovery time and point, checks against targets and a test calendar: the resilience rows of chapter 29’s plan.

Interface. firm_drplan: Step, Plan, separation_km, outside_hazard, rto_samples, p_within, rpo, sync_penalty_ms, tested_within, load_calendar.

Rules. Sites come from the site table with its sources; test dates are dated rows; runbook durations are stated assumptions, replaced by measured ones after each test.

Acceptance tests. code/firm/drplan/tests/: a deterministic runbook, the recovery point by mode, separations, and the annual-test check.

Stretch. Parallel steps; correlated failures of the primary and a provider; the exchange’s own recovery time added to the firm’s.

Sources and further reading

  • 17 CFR 242.1001 and 242.1004 (Regulation SCI); SEC Releases 34-69077 (2013) and 34-73639 (2014).
  • Commission Delegated Regulation (EU) 2017/589, Article 14; FCA PS21/3; ESMA on DORA.
  • FIA, 2025 disaster recovery test.

28.7 Exercises

Exercise 28.1 ★

What must an SCI entity’s backup capabilities be designed to achieve after a wide-scale disruption?

Solution

Solution of Exercise 28.1.

Next business day resumption of trading and two-hour resumption of critical SCI systems following a wide-scale disruption, with backup capabilities sufficiently resilient and geographically diverse.

Exercise 28.2 ★

What is the difference between a recovery time objective and a recovery point objective?

Solution

Solution of Exercise 28.2.

The time objective bounds how long a system may be down; the point objective bounds how much recent data may be lost. The first is set by the runbook, the second by replication.

Exercise 28.3 ★

Why is Mahwah a poor recovery site for a firm trading from Secaucus?

Solution

Solution of Exercise 28.3.

It is 34.2 km34.2\,\mathrm{k}\mathrm{m} from Secaucus, inside any regional disaster that takes the primary: a storm, a flood or a power failure would likely take both.

Exercise 28.4 ★★

Using Proposition 28.5, why does faster replication not shorten the warm design’s recovery time?

Solution

Solution of Exercise 28.4.

Replication sets only the recovery point; the warm design’s time is its runbook, dominated by starting and restoring servers and loading data, which faster replication does not change.

Exercise 28.5 ★★

What does RTS 6 require a trading firm’s business continuity arrangements to include about outstanding orders?

Solution

Solution of Exercise 28.5.

Alternative arrangements to manage outstanding orders and positions, and arrangements to shut down algorithms or systems without creating disorderly trading conditions.

Exercise 28.6 ★★

The firm’s last failover test was on 25 October 2025. Is it within twelve months on 28 September 2026, and on 1 November 2026?

Solution

Solution of Exercise 28.6.

Yes on 28 September 2026; no on 1 November 2026, more than twelve months after the test.

Exercise 28.7 ★★★

Coding. With firm_drplan, how much would halving the warm design’s restore step (45 to 22.5 minutes) raise its chance of resuming within two hours?

Solution

Solution of Exercise 28.7.

From 20.3% to 44.1%: the restore step is the largest, but the other five steps still take about two hours at their medians together.

Exercise 28.8 ★★★

Find the flaw. “We tested our backup site last month and it worked, so we will switch to it if a storm comes.”

Solution

Solution of Exercise 28.8.

A test is not a decision: before Sandy the industry’s test had succeeded two days earlier and the markets still did not switch. The firm needs a runbook with a decision owner and criteria, rehearsed switching, not only connectivity, and the exchanges’ own plans.

28.8 Problem: The Two-Hour Rule

Problem 28.1

Weekend problem — hot against warm recovery

A firm trades from Secaucus NY5 and must choose a recovery design. It has set itself the exchanges’ two-hour line as its recovery time objective. Use the chapter’s model.

Part I — The site.

  1. How far is each candidate site from NY5, and which lie outside a 100-kilometre hazard radius?
  2. What would synchronous replication to Aurora and to Mahwah add to each write?
  3. Why do the exchanges recover so far away?
  4. What recovery point does asynchronous replication to Aurora give?

Part II — The designs.

  1. What is the hot design’s chance of resuming within two hours, its median and 95th percentile?
  2. And the warm design’s?
  3. What can the warm design lose in data?
  4. Which steps dominate the warm design’s time?

Part III — The rules.

  1. What does Regulation SCI require of exchanges, and of their designated members?
  2. What does RTS 6 require of trading firms?
  3. What did the UK rules require by 31 March 2025?
  4. What did Sandy show that tests alone did not?

Part IV — The verdict.

  1. State the named result: the probability that warm and hot designs resume within two hours, and the annual cost of the difference.
  2. Which inputs are published and which assumed?
  3. Is the hot design worth USD 1.44 million a year?
  4. What would make the warm design meet the line more often?
  5. What evidence would a regulator ask for?
  6. How should the firm take part in exchange tests?
  7. Where does this go in chapter 29’s plan?
  8. In one sentence: what does a recovery design buy?
Solution

Solution of Problem 28.1.

Part I.

  1. NY4 0.3 km0.3\,\mathrm{k}\mathrm{m}, Carteret 26.0, Mahwah 34.2, Aurora 1 191.1; only Aurora is outside.
  2. 17.4 ms17.4\,\mathrm{m}\mathrm{s} of round trip to Aurora at a route factor of 1.5, 0.5 ms0.5\,\mathrm{m}\mathrm{s} to Mahwah.
  3. To be outside the disasters that could take the primary, as the rules require geographic diversity.
  4. 56 ms56\,\mathrm{m}\mathrm{s}: 50 ms50\,\mathrm{m}\mathrm{s} of lag plus the one-way delay.

Part II.

  1. 99.2%, 64 and 96 minutes.
  2. 20.3%, 148 and 237 minutes.
  3. Up to fifteen minutes, the snapshot interval.
  4. Starting and restoring servers (45 minutes at the median) and loading data (30).

Part III.

  1. Plans reasonably designed to resume critical systems within two hours and trading by the next business day, geographically diverse; designated members testing with the exchange at least every twelve months.
  2. Documented business continuity arrangements per venue, relocation to a back-up site where appropriate, orderly shutdown, management of outstanding orders and positions, annual review and test.
  3. Firms had to be able to remain within the impact tolerances of their important business services.
  4. That firms and exchanges might not switch even when the backups worked: the decision and the confidence matter as much as the equipment.

Part IV.

  1. Named result: with the chapter’s runbooks, the hot design resumes within two hours in 99.2% of recoveries and the warm design in 20.3%; at 70% and 25% of chapter 27’s USD 3.19 million a year, the difference costs USD 1.44 million a year.
  2. Published: the rules, the exchanges’ recovery sites, test dates and site coordinates. Assumed: every step’s duration, the costs, the hazard radius, the replication lag.
  3. If the firm’s impact tolerance for trading is two hours, and a lost day costs more than USD 1.44 million often enough; otherwise a warm site with an orderly shutdown may be the proportionate choice.
  4. Pre-staged servers, automated restores, parallel steps, and measured, rehearsed runbooks.
  5. Named services and tolerances, the runbook, test results with measured times, and records of participation in exchange tests.
  6. As a designated member where asked, and otherwise voluntarily, with the same runbook it would use for real.
  7. In the resilience rows: the recovery site, its lines and cross-connects, the replication, the runbook and the test calendar.
  8. Hours of trading back after a disaster, at a yearly price.

28.9 Interview questions

Interview question 28.1 ★ developer

What are RTO and RPO, and what sets each?

Solution

Solution of Interview question 28.1.

The recovery time objective is how long a system may be down; the recovery point objective how much data may be lost. The runbook sets the first; replication sets the second.

What the interviewer is looking for: Definitions; runbook against replication.

Interview question 28.2 ★★ developer

Where would you put the recovery site of a firm trading in Secaucus, and why?

Solution

Solution of Interview question 28.2.

Outside the New York metropolitan area’s shared hazards, near the venues’ own recovery sites and with good lines to them: Aurora is the natural choice for U.S. futures and BrokerTec; the costs are asynchronous replication and a recovery point.

What the interviewer is looking for: Hazards; the venues’ recovery sites; replication trade-off.

Interview question 28.3 ★★ developer

How do you avoid a split brain when failing over a trading system?

Solution

Solution of Interview question 28.3.

One authority decides which site is active (a person with criteria, or a quorum), the old primary is fenced before the new one trades, and sessions use the exchange’s separate recovery addresses so both sites cannot trade at once.

What the interviewer is looking for: Single decision; fencing; separate addresses.

Interview question 28.4 ★★ trader

The primary site is lost with open positions. What happens in the first thirty minutes?

Solution

Solution of Interview question 28.4.

Cancel-on-disconnect and the exchanges’ kill switches stop resting orders; positions are known from drop copies and the clearing firm; hedges go through the recovery site or a broker by phone; then the runbook brings trading back.

What the interviewer is looking for: Exposure first; drop copies; fallbacks; runbook.

Interview question 28.5 ★★ developer, researcher

How would you decide between a hot and a warm recovery site?

Solution

Solution of Interview question 28.5.

From the impact tolerance: simulate or measure each design’s recovery time against it, price each, and compare the difference with the expected cost of the extra downtime.

What the interviewer is looking for: Tolerance first; measured times; cost against expected loss.

Interview question 28.6 ★★★ developer, researcher

Design the business continuity arrangements of an algorithmic trading firm active in the U.S. and the EU, so that they satisfy the rules it is examined on.

Solution

Solution of Interview question 28.6.

Important business services and tolerances; a recovery site outside the primary’s hazards near the venues’ recovery sites; runbooks with measured times; orderly shutdown and management of outstanding orders per venue; annual tests with the exchanges and the industry; records kept as evidence.

What the interviewer is looking for: Rules mapped to arrangements; testing; evidence.

Terms defined in this chapter

See all 2333 terms in the glossary