---
title: "The Risk Grid"
book: "Research, Data and Risk Platforms"
subject: quant
language: en
chapter: 20
exercises: 8
source: https://one-course.com/books/quant/15/en/chapter/20-the-risk-grid
---

# Chapter 20 — The Risk Grid

The nightly risk batch — every trade, every historical scenario — starts at 04:30 and must be done by 06:00, when the morning’s risk reports are built. It finished at 06:51. Doubling the machines moved the finish to 06:25 and no earlier, because the last hour belonged to one task waiting for one trade that priced two hundred times slower than its estimate, and the scheduler, which believed the estimate, had started that task late. Nothing was broken; the grid was planned by trade counts, not by cost. This chapter plans, runs and repairs such a grid on the simulated cluster of chapter 13, with the costs of the real pricing library measured on this laptop.

## 20.1 Nightly revaluation at scale

**Definition 20.1 (Risk grid, batch window).**

A *risk grid* is the infrastructure that revalues a firm’s trades under many market scenarios on many machines: the plan that cuts the work into tasks, the scheduler and cluster that run them, and the store that collects the results for the risk engine to aggregate. Its *batch window* is the interval between the moment its inputs are ready (the end-of-day positions and market snapshot) and the moment its outputs are needed.

Book 6’s risk engine (chapter 29) computes a book’s P&L vectors under historical scenarios by full revaluation, a revaluation grid or a delta-gamma approximation, and aggregates them into VaR and expected shortfall along the risk hierarchy. At the size of a firm the computation is a grid of trades by scenarios — here 105 000 trades and 1 000 scenarios, 105 million pricing calls — and the engine becomes a scheduling problem.

The chapter’s book: 60 000 swaps and 20 000 swaptions (Book 6’s rates instruments), 20 000 equity options and 5 000 autocallables priced by Monte Carlo with the library’s 100 000 paths (Book 5). One pricing call of each, measured on this laptop with the real library, costs 14.5, 10.2 and 5.1 microseconds for the first three and 22.2 milliseconds for an autocallable: the 5 000 autocallables are 99% of the book’s 112.2 seconds of pricing per scenario. A task costs a fixed five seconds (loading the snapshot and the trades, writing the results) plus its pricing. The grid has 32 cores (eight nodes of four), and the window is 90 minutes.

## 20.2 Fanning out scenarios

**Definition 20.2 (Scenario fan-out, task granularity).**

*Scenario fan-out* is the splitting of the scenarios into batches so that one trade batch’s revaluation runs as several tasks in parallel. *Task granularity* is the size of the resulting tasks — here the number of scenarios in a batch, the trades being grouped into batches of about half a second of pricing per scenario — which trades scheduling freedom against the fixed cost of each task.

![The grid as tasks: rows are trade batches grouped by estimated cost, columns scenario batches; every cell is a task. A cost-aware plan gives a trade measured slow its own row, with finer columns, so that no single task holds the window.](https://one-course.com/images/onecourse/chapters/quant-15/pl-the-risk-grid/fig-c27ec05a02d8.svg)

***Figure 20.1.** The grid as tasks: rows are trade batches grouped by estimated cost, columns scenario batches; every cell is a task. A cost-aware plan gives a trade measured slow its own row, with finer columns, so that no single task holds the window.*

```python
def plan(trades, n_scen: int, trade_target: float, scen_batch: int, overhead: float,
         isolate_above: float | None = None) -> list[TaskSpec]:
    """Trade batches of about `trade_target` seconds per scenario times scenario batches of
    `scen_batch`; a trade costing more than `isolate_above` a task gets finer tasks."""
    big = isolate_above is not None
    heavy = [t for t in trades if big and t.est * scen_batch > isolate_above]
    light = [t for t in trades if t not in heavy]
    out = []
    for group in _batches(light, trade_target):
        for s0 in range(0, n_scen, scen_batch):
            out.append(TaskSpec(len(out), group, min(scen_batch, n_scen - s0), overhead))
    for t in heavy:
        fine = max(1, int(isolate_above // t.est))
        for s0 in range(0, n_scen, fine):
            out.append(TaskSpec(len(out), [t], min(fine, n_scen - s0), overhead))
    return out
```

***Listing 20.1.** Planning the grid: trade batches by estimated cost, scenario batches of a chosen size, and trades too costly for one task isolated and fanned out more finely. code/firm/riskgrid/firm_riskgrid.py*

With 1 000 scenarios in a batch there are 220 tasks, each a trade batch across all scenarios: few, long, and impossible to balance. With five scenarios there are 44 000, and their fixed costs — 44 000 times five seconds — are larger than the pricing itself. [Figure 20.2](#fig-pl-the-risk-grid-granularity) is the whole curve: on 32 cores the [makespan](https://one-course.com/books/quant/15/en/chapter/13-distributed-compute-and-schedulers#def-pl-distributed-compute-and-schedulers-makespan) falls from 2.35 hours at 1 000 scenarios a task to 1.19 at 100, and rises again to 2.93 at five; only batches of 20 to 250 scenarios meet the window, 20 by sixteen seconds.

![Makespan of the nightly grid on 32 cores (solid) and 64 (dashed) against the number of scenarios per task, for a plan that trusts the cost estimates (naive) and one that isolates the trade measured slow last night (cost-aware); the dotted line is the 90-minute window. Coarse tasks cannot be balanced; fine tasks drown in their fixed costs. Data: fig_riskgrid.py, with costs from bench_riskgrid.py.](https://one-course.com/images/onecourse/chapters/quant-15/pl-the-risk-grid/fig-90705f09c6f6.svg)

***Figure 20.2.** [Makespan](https://one-course.com/books/quant/15/en/chapter/13-distributed-compute-and-schedulers#def-pl-distributed-compute-and-schedulers-makespan) of the nightly grid on 32 cores (solid) and 64 (dashed) against the number of scenarios per task, for a plan that trusts the cost estimates (naive) and one that isolates the trade measured slow last night (cost-aware); the dotted line is the 90-minute window. Coarse tasks cannot be balanced; fine tasks drown in their fixed costs. Data: `fig_riskgrid.py`, with costs from `bench_riskgrid.py`.*

## 20.3 Task granularity, failure and retry

The slow trade makes coarse tasks worse still. One autocallable’s engine falls back to daily steps and prices at 4.88 seconds, 220 times its estimate. In the naive plan its trade batch across all 1 000 scenarios is a task of more than 5 300 seconds that the scheduler, sorting by estimate, starts like any other: the grid finishes 8 468 seconds after 04:30, at 06:51; on 64 cores, at 06:25 (6 922 seconds). A cost-aware plan uses last night’s measured cost: the slow trade gets tasks of its own, 122 scenarios each, and the 1 000-scenario plan finishes in 4 123 seconds on 32 cores (05:38); the best granularity, 250 scenarios, in 3 846 (05:34).

```python
def run(tasks: list[TaskSpec], nodes: int, cores: int, failures: str = "trade",
        retries: int = 2, seed: int = 0) -> GridRun:
    """Longest estimate first on the simulated cluster. failures='task': a task with a failing
    trade runs 1 + retries times and loses its cells; 'trade': the trade is quarantined."""
    jobs, wasted, lost, quarantined = [], 0.0, 0, []
    for spec in tasks:
        bad = [t for t in spec.trades if t.fails]
        d = spec.duration
        if bad and failures == "task":
            wasted += retries * d
            d *= 1 + retries
            lost += len(spec.trades) * spec.scenarios
        elif bad:
            quarantined += [t.trade_id for t in bad]
            lost += len(bad) * spec.scenarios
        jobs.append(J.Task(spec.tid, d, spec.estimate))
    s = J.simulate(jobs, J.Cluster(nodes, cores, seed=seed), J.LPT())
    return GridRun(s.makespan, s.busy, wasted, lost, sorted(set(quarantined)), len(tasks))
```

***Listing 20.2.** Running a plan: failures handled per task (retried, then lost) or per trade (quarantined), then the tasks scheduled longest estimate first on the simulated cluster. code/firm/riskgrid/firm_riskgrid.py*

Failures need the same thinking. One swaption’s volatility is missing, and it fails every time. If a failing trade fails its task, the scheduler retries the task twice and then gives up: the grid loses every result of every task that held the trade — 38.3 million of the 105 million cells, more than a third of the book — and burns 1 100 core-seconds on retries. If the pricing loop catches the failure per trade, as Book 6’s engine does (a failing trade gives a missing value and an entry in an error report, never a stopped run), the grid loses the failing trade’s 1 000 cells, quarantines it with a report, and finishes 60 seconds sooner.

| failure handling | [makespan](https://one-course.com/books/quant/15/en/chapter/13-distributed-compute-and-schedulers#def-pl-distributed-compute-and-schedulers-makespan) (s) | core-seconds wasted | results lost |
| --- | --- | --- | --- |
| per task (two retries, then lost) | 4 095 | 1 100 | 38 263 000 |
| per trade (quarantined with a report) | 4 035 | 0 | 1 000 |

***Table 20.1.** One trade that always fails, in the cost-aware plan with 100 scenarios per task on 32 cores.*

## 20.4 Adjoint sensitivities in production

The nightly batch also computes sensitivities. Bumping a swap’s eight curve pillars up and down costs sixteen repricings; the adjoint of Book 4’s tape (`firm.aad`, chapter 28 there) records one valuation and sweeps back once, and gives all eight derivatives — equal to the bumped ones to $10^{-9}$ in the tests ([Listing 20.3](#lst-pl-the-risk-grid-adjoint)). On this laptop the bumped deltas of a swap take 388 microseconds and the adjoint 66, 5.9 times faster, with a Python tape; a compiled tape keeps the adjoint’s cost at a few valuations whatever the number of pillars. Over the book’s 60 000 swaps, bumping needs 1.02 million pricing calls, the adjoint 60 000 recorded valuations.

```python
def swap_pv_adjoint(zeros, tenors, years: int, fixed: float, payer: bool = True):
    """PV of a spot-starting annual swap (unit notional) on pillar zero rates, linear between
    pillars, flat outside, and its derivative to every pillar, in one reverse sweep."""
    def pv(z):
        def zero(t):
            if t <= tenors[0]:
                return z[0]
            if t >= tenors[-1]:
                return z[-1]
            i = next(k for k in range(1, len(tenors)) if t < tenors[k])
            w = (t - tenors[i - 1]) / (tenors[i] - tenors[i - 1])
            return z[i - 1] * (1 - w) + z[i] * w
        dfs = [AD.exp(-zero(t) * t) for t in [_act365(k) for k in range(1, years + 1)]]
        value = (1.0 - dfs[-1]) - fixed * sum(dfs[1:], dfs[0])
        return value if payer else -value
    return AD.gradient(pv, list(zeros))
```

***Listing 20.3.** A swap’s value on a pillar curve written once on the tape; one reverse sweep gives its derivative to every pillar. code/firm/riskgrid/firm_riskgrid.py*

## 20.5 Intraday risk and incremental revaluation

**Definition 20.3 (Incremental revaluation).**

*Incremental revaluation* recomputes only the results whose inputs changed: each result is stored with the version of the trade and of the market data it was computed from, and a rerun reprices only the trades whose key is not in the store.

Intraday, the scenarios and the base snapshot are the morning’s, and most trades have not changed. With results keyed by (trade, trade version, market-data version), an intraday rerun after 1 500 trades are amended or booked reprices those 1 500 under the 1 000 scenarios: 1.5 million pricing calls instead of 105 million. A new market snapshot changes every key; for that the real-time system of chapter 18 carries sensitivities forward, and the grid waits for the night.

![Pricing calls of the nightly grid, of the swaps’ pillar deltas by bumping and by the adjoint, and of an intraday rerun done in full and incrementally after 1 500 trades changed.](https://one-course.com/images/onecourse/chapters/quant-15/pl-the-risk-grid/fig-7d6cbc5223ef.svg)

***Figure 20.3.** Pricing calls of the nightly grid, of the swaps’ pillar deltas by bumping and by the adjoint, and of an intraday rerun done in full and incrementally after 1 500 trades changed.*

**Example 20.4 (06:51 against 06:00).**

Planned by trade batches across all scenarios and trusting its estimates, the grid finishes at 06:51 on 32 cores and at 06:25 on 64: more cores shorten the queue, not the one long task. Cutting the scenarios into batches of 100 brings the naive plan to 05:41; isolating the slow trade by its measured cost brings every granularity from 20 to 1 000 scenarios inside the window, the best at 05:34 (250 scenarios a task). Quarantining the failing trade instead of failing its tasks keeps more than a third of the book’s results. The pricing calls saved elsewhere are larger still: the adjoint replaces 1.02 million bumped repricings of swaps by 60 000 recorded valuations, and an incremental intraday rerun makes 1.5 million calls where a full one makes 105 million.

## 20.6 Tutorial: planning a nightly grid

**Goal.** Measure pricing costs, plan the grid at several granularities, repair it for a slow and a failing trade, and count the calls adjoints and incremental runs save. **End state:** [Figure 20.2](#fig-pl-the-risk-grid-granularity) and [Table 20.1](#tab-pl-the-risk-grid-failures).

1. **Costs** : `bench_riskgrid.py` : one pricing call per kind, bumped and adjoint swap deltas.
2. **Plan and run** : `pl_riskgrid.grid(scen_batch, cost_aware)` ; `sweep()` .
3. **Failures** : `failure_policies()` .
4. **Adjoints** : `firm_riskgrid.swap_pv_adjoint` against the library’s bumped sensitivities.
5. **Intraday** : `intraday()` .

**What to change next.** Give the scheduler the measured costs of every trade and see how close to the lower bound the grid comes; add the sensitivities’ repricings to the grid’s work and find the new best granularity.

## 20.7 Build: the risk grid

**Purpose.** The nightly revaluation planned by cost, run to its window, robust to slow and failing trades, and cheaper by adjoints and incremental reruns.

**Interface.** `GridTrade`, `plan`, `run -> GridRun`, `swap_pv_adjoint`, `ResultCache`, `cells`, `lower_bound`, `hours`.

**Rules.** Plans use costs measured on the previous run where they exist; a trade’s failure never fails its task; results carry their trade and market-data versions; adjoints are checked against bumps.

**Acceptance tests.** `code/firm/riskgrid/tests/`: batching and isolation of a heavy trade; the two failure policies and the lower bound; adjoint pillar deltas equal to the library’s bumped sensitivities; the result cache’s keys.

**Stretch.** Results written to chapter 11’s [results store](https://one-course.com/books/quant/15/en/chapter/11-backtest-engine-architecture#def-pl-backtest-engine-architecture-results); the plan fed by a cost model learned from past runs; checkpointed adjoints through Monte Carlo paths (Book 4, chapter 28).

Sources and further reading

- One Quant Book 4, chapter 28 (adjoints); One Quant Book 5, chapter 28, and Book 6, chapter 29 (the library, the risk engine); chapter 13 of this book (scheduling).

## 20.8 Exercises

**Exercise 20.1 ★.**

How many tasks does the plan make with 100 scenarios per task, and how much fixed cost do they add?

**Solution of Exercise 20.1.**

2 200 tasks (220 trade batches times ten scenario batches), adding $2\,200 \times 5 = 11\,000$ core-seconds of fixed cost, about 10% of the pricing.

**Exercise 20.2 ★.**

Why does doubling the cores move the naive plan’s finish from 06:51 only to 06:25?

**Solution of Exercise 20.2.**

The finish is set by one task — the slow trade’s batch across all scenarios, about 5 380 seconds — and by when it starts. More cores start it earlier, since the queue in front of it drains faster, but cannot shorten it: on 64 cores it still ends 6 922 seconds after the start.

**Exercise 20.3 ★.**

How many pricing calls do bumped pillar deltas of a swap need, and how many valuations does the adjoint need?

**Solution of Exercise 20.3.**

Sixteen repricings (each of eight pillars up and down) plus the base: 17 calls; the adjoint records one valuation and runs one reverse sweep, at a cost of a few valuations whatever the number of pillars.

**Exercise 20.4 ★★.**

Why does the [makespan](https://one-course.com/books/quant/15/en/chapter/13-distributed-compute-and-schedulers#def-pl-distributed-compute-and-schedulers-makespan) rise again below 50 scenarios per task?

**Solution of Exercise 20.4.**

The number of tasks grows as the batches shrink — 44 000 at five scenarios — and each carries five seconds of fixed cost; below about 50 scenarios the fixed costs grow faster than the balance improves.

**Exercise 20.5 ★★.**

Why does failing a task for one trade lose 38 million results?

**Solution of Exercise 20.5.**

The failing swaption sits in a trade batch of several thousand trades; every task of that batch — one per scenario batch — fails, so the batch’s results under every scenario are lost: 38.3 million cells.

**Exercise 20.6 ★★.**

When does [incremental revaluation](#def-pl-the-risk-grid-incremental) save nothing, and what covers that case?

**Solution of Exercise 20.6.**

When the market-data version changes, since every result’s key changes. Intraday risk on a new market is covered by sensitivities carried forward (chapter 18) and, if needed, a partial revaluation of the trades most exposed to the move.

**Exercise 20.7 ★★★.**

*Coding.* Run the sweep on 64 cores. Which granularities meet the window with each plan, and which is best?

**Solution of Exercise 20.7.**

On 64 cores the naive plan meets the window from 5 to 500 scenarios a task (best 50: 2 319 seconds) and fails only at 1 000; the cost-aware plan meets it at every granularity, 5 to 1 000 (best 250: 1 989 seconds). More cores widen the good range; they do not remove the need to plan.

**Exercise 20.8 ★★★.**

*Find the flaw.* “The batch was late, so we bought twice the machines.”

**Solution of Exercise 20.8.**

Doubling the cores moved the finish from 06:51 to 06:25, still late, because the lateness came from one task and the plan, not from capacity. Measure the tasks, isolate what is slow, choose the granularity; buy capacity only when the lower bound of a good plan is above the window.

## 20.9 Problem: 06:51 Against 06:00

**Problem 20.1.**

Weekend problem — a grid that misses its window

The chapter’s book, costs, cluster and plans.

**Part I — The grid.**

1. What is the grid’s work, in trades, scenarios and pricing calls?
2. Which trades dominate its cost, and why?
3. What does a task cost besides its pricing?
4. What is the window, and on how many cores?
5. What is the lower bound on 32 cores?

**Part II — Granularity.**

6. What is the [makespan](https://one-course.com/books/quant/15/en/chapter/13-distributed-compute-and-schedulers#def-pl-distributed-compute-and-schedulers-makespan) at 1 000, 100 and 5 scenarios a task?
7. Which granularities meet the window with the naive plan?
8. What does the slow trade do to coarse tasks, and why does the scheduler not help?
9. What does the cost-aware plan change?
10. What is the best cost-aware finish?

**Part III — Failures and savings.**

11. What happens to a failing trade per task and per trade?
12. How many results and core-seconds does each policy lose?
13. How fast are adjoint swap deltas against bumped ones, here?
14. How many pricing calls does the adjoint save over the book’s swaps?
15. What does an incremental intraday rerun cost?

**Part IV — The verdict.**

16. State the *named result* : the batch’s [makespan](https://one-course.com/books/quant/15/en/chapter/13-distributed-compute-and-schedulers#def-pl-distributed-compute-and-schedulers-makespan) as a function of [task granularity](#def-pl-the-risk-grid-fanout) and cluster size, the granularity that meets the window, and the pricing calls saved by adjoint sensitivities and by incremental intraday revaluation.
17. What would you change first in the grid’s planning?
18. What would you monitor every night?
19. When would you buy more machines?
20. In one sentence: what decides when a nightly batch ends?

**Solution of Problem 20.1.**

1. 105 000 trades by 1 000 scenarios: 105 million pricing calls (plus the base valuations).
2. The 5 000 Monte Carlo autocallables: 22.2 milliseconds a call, 99% of the 112.2 seconds of pricing per scenario.
3. Five seconds: loading the snapshot and the trades, writing the results.
4. 90 minutes, 04:30 to 06:00, on 32 cores.
5. About 4 000 seconds at 100 scenarios a task: the work over 32 cores.
6. 2.35, 1.19 and 2.93 hours.
7. 20, 50, 100 and 250 scenarios a task (20 by sixteen seconds).
8. It makes one task of more than 5 300 seconds, started late because its estimate is small.
9. It uses last night’s measured cost to give the slow trade tasks of their own, fanned out finely.
10. 05:34, with 250 scenarios a task.
11. Per task: two retries, then the task’s results are lost; per trade: the trade is quarantined with a report and the task completes.
12. 38.3 million results and 1 100 core-seconds per task; 1 000 results and nothing per trade.
13. 5.9 times (388 against 66 microseconds).
14. 1.02 million bumped repricings replaced by 60 000 recorded valuations.
15. 1.5 million calls for 1 500 changed trades, against 105 million.
16. **Named result.** On 32 cores the grid takes 2.35 hours with a task per trade batch across all scenarios, 1.19 with 100 scenarios a task and 2.93 with five; only 20 to 250 scenarios meet the 90-minute window with the naive plan, and a cost-aware plan meets it at every granularity from 20 to 1 000 (best 1.07 hours at 250); doubling the cores moves the naive plan from 06:51 to 06:25; adjoints replace 1.02 million bumped repricings of the swaps by 60 000 valuations, and an incremental intraday rerun makes 1.5 million calls instead of 105 million.
17. Plan by measured cost, with the scenario batch chosen from the fixed cost and the cluster, and quarantine failing trades.
18. Each task’s duration against its estimate, the slowest trades, failures and quarantines, the finish against the window, and the lower bound.
19. When a good plan’s lower bound, not its [makespan](https://one-course.com/books/quant/15/en/chapter/13-distributed-compute-and-schedulers#def-pl-distributed-compute-and-schedulers-makespan) , approaches the window.
20. Its longest tasks and when they start.

## 20.10 Interview questions

**Interview question 20.1 ★ developer.**

Your nightly risk batch misses its window. What do you look at first?

**Solution of Interview question 20.1.**

The tasks on the critical path: which ran last and longest, and whether their durations matched their estimates; then failures and retries; then the granularity against the fixed cost of a task; capacity last.

*What the interviewer is looking for: Critical path before capacity.*

**Interview question 20.2 ★★ developer.**

How would you split a million trades by a thousand scenarios into tasks?

**Solution of Interview question 20.2.**

Batch trades by measured cost to a target per scenario, batch scenarios so that tasks last minutes (many times the fixed cost, much less than the window), isolate and fan out expensive trades, and schedule longest first.

*What the interviewer is looking for: Cost-based batching, two dimensions.*

**Interview question 20.3 ★★ developer, researcher.**

Why use adjoint differentiation for risk, and where does it not help?

**Solution of Interview question 20.3.**

Because it gives every sensitivity of a price for a few times the cost of the price, where bumping costs two prices per sensitivity; it does not help when there are few sensitivities, when the code cannot be taped (a closed-source pricer), or for non-smooth payoffs without smoothing.

*What the interviewer is looking for: Cost independent of the number of inputs.*

**Interview question 20.4 ★★ developer.**

One trade fails to price every night. What should the grid do with it?

**Solution of Interview question 20.4.**

Quarantine it: the pricing loop catches its failure, records it in an error report with its reason, completes the task without it, and the report goes to the owner of the trade’s data before the next run.

*What the interviewer is looking for: Isolate the failure, never the task.*

**Interview question 20.5 ★★★ developer.**

Design the nightly and intraday revaluation infrastructure of a bank’s trading book.

**Solution of Interview question 20.5.**

A grid planned from measured costs on a scheduler with fair share and longest-first ordering; tasks sized from their fixed costs; results stored by trade, trade version and market version; failures quarantined with reports; adjoint sensitivities where the library supports them; incremental intraday reruns for changed trades and sensitivities carried forward for market moves.

*What the interviewer is looking for: Planning, robustness, reuse.*
