---
title: "The Platform Map"
book: "Research, Data and Risk Platforms"
subject: quant
language: en
chapter: 1
exercises: 8
source: https://one-course.com/books/quant/15/en/chapter/1-the-platform-map
---

# Chapter 1 — The Platform Map

At 07:40 the overnight risk report, the P&L and a researcher’s backtest disagree about one position. Three teams spend the morning finding out why. The vendor’s corporate-action file, which normally lands at 19:30, came at 21:30 the night before; the risk batch, which starts at 21:00 by the clock, had already read the old security master and the positions built on it; the P&L job, which starts when its inputs are ready, read the new ones twenty minutes later. Every system did what it was built to do. What was missing was a picture, kept as data and checked by a program, of which system reads what, from whom, and when. This chapter draws that picture for the miniature firm built across the series, and turns it into a tool (`firm.platmap`) that answers the questions the three teams spent their morning on: what a late file reaches, which job reads an input before it is complete, and how late each input may arrive before the morning’s reports miss their deadline.

## 1.1 The systems of a trading firm

A trading firm is a chain of systems that turn market data and decisions into positions, and positions into money that has been counted, checked and reported. Books 1 to 14 of this series built most of the links one at a time: the exchange simulator and its feed (Book 10), the feed handler, order gateway and risk gate (Book 13), the backtesters and the research workflow (Book 7), the pricing library (Book 5), the risk engine (Book 6), the machine-learning platform (Book 12). This book is about the platforms that connect them, and it starts with the map.

**Definition 1.1 (Front-to-back flow).**

The *front-to-back flow* of a trading firm is the path that a trade and the data about it take through the firm’s systems: market and reference data; the decision and the order; the execution; the position and its P&L; the risk it adds; its booking, confirmation, settlement and reconciliation; and the controls and reports that close the day.

The industry names the stretches of that path after the teams that historically ran them: the front, middle and back office, which One Quant Book 16 (chapter 13) defines as organisational units. This book cares about the systems, whoever runs them. [Figure 1.1](#fig-pl-the-platform-map-layers) shows the miniature firm’s eighteen systems in eight layers, and the direction in which the day’s data move through them.

![The miniature firm’s systems by layer, from the venues at the top to the controls at the bottom; the research platform reads the tick store and the reference data beside the trading path. Arrows show the main direction of the day’s data; the full set of flows is the map’s data (pl_platmap), not the drawing.](https://one-course.com/images/onecourse/chapters/quant-15/pl-the-platform-map/fig-562402596525.svg)

***Figure 1.1.** The miniature firm’s systems by layer, from the venues at the top to the controls at the bottom; the research platform reads the tick store and the reference data beside the trading path. Arrows show the main direction of the day’s data; the full set of flows is the map’s data (`pl_platmap`), not the drawing.*

Read from the top, the layers are the order in which a trade’s data are born. Market data and reference data exist before any decision; the order gateway sends orders that the venue fills; the position service keeps what the fills add up to; the P&L service values it with the pricing library’s marks; the risk grid revalues it under scenarios; trade capture books it and post-trade systems confirm, settle and reconcile it; product control, surveillance and regulatory reporting check and report all of the above. A research platform sits to the side: it reads the same tick store and reference data to build the next strategy, and its outputs come back as parameters and models that the trading layer loads (chapter 16).

**Example 1.2 (The miniature firm in numbers).**

The map of the miniature firm (`pl_platmap`) has 18 systems, 22 datasets and 15 end-of-day jobs. Seven datasets arrive from outside: the day’s fills, the day’s market-data capture, the official closing prices, a market-data snapshot of curves and volatilities, the vendor’s reference-data and corporate-action files, and the clearing broker’s statement. Every other dataset is produced by exactly one job from others.

## 1.2 Data flows and their clocks

The difficulty of a platform is not the number of systems but the number of clocks. Market data and fills arrive continuously during the session; the official closing prices come some minutes after the close; vendor files come in the evening at times the vendor does not guarantee; the clearing broker’s statement comes when the clearing house has finished its own day. Each job of the end-of-day chain runs on one of two kinds of clock: it starts when its inputs are ready, or it starts at a time on the wall.

**Definition 1.3 (Data-flow map).**

A *data-flow map* is a description, kept as data, of a firm’s systems, the datasets each one holds, and the jobs that read datasets and write others, each job with its duration and its clock (dependency-driven or scheduled at a fixed time), each dataset with its [system of record](#def-pl-the-platform-map-sor) and, for external inputs, its arrival time.

Because it is data, a map can be checked (every dataset has one owner, the jobs form no cycle), traversed (what does this file reach?) and simulated (when will each report be ready tonight?). A diagram drawn once for a presentation does none of these, and is out of date by the next release.

![The clocks of a trading day. Ticks and fills stream during the session; the end-of-day batch starts at the close and waits for files whose arrival times it does not control; overnight research runs after the batch has produced its inputs; the morning reports are due at 07:00. Times are the miniature firm’s nominal ones.](https://one-course.com/images/onecourse/chapters/quant-15/pl-the-platform-map/fig-255960a2e67c.svg)

***Figure 1.2.** The clocks of a trading day. Ticks and fills stream during the session; the [end-of-day batch](#def-pl-the-platform-map-batch) starts at the close and waits for files whose arrival times it does not control; overnight research runs after the batch has produced its inputs; the morning reports are due at 07:00. Times are the miniature firm’s nominal ones.*

[Listing 1.1](#lst-pl-the-platform-map-jobs) is the end-of-day chain of the miniature firm as the map records it: each job with its inputs, outputs and duration in minutes. The risk batch takes an optional fixed start; everything else starts on its inputs.

```python
def jobs(risk_start=None) -> list[Job]:
    return [
        Job("compact ticks", ("capture",), ("tick history",), 35),
        Job("golden copy", ("vendor reference", "vendor corporate actions"),
            ("security master",), 25),
        Job("end-of-day marks", ("closing prices", "snapshot"), ("marks",), 20),
        Job("positions", ("fills", "security master"), ("positions",), 20),
        Job("pnl", ("positions", "marks", "security master"), ("pnl",), 25),
        Job("scenarios", ("snapshot",), ("scenarios",), 40),
        Job("risk batch", ("positions", "marks", "scenarios", "security master"),
            ("risk report",), 360, start=risk_start),
        Job("trade capture", ("fills",), ("booked trades",), 30),
        Job("confirmations", ("booked trades",), ("confirmations",), 45),
        Job("reconciliation", ("positions", "clearing statement"), ("breaks",), 40),
        Job("pnl sign-off", ("pnl", "breaks"), ("signed pnl",), 30),
        Job("surveillance", ("tick history", "fills"), ("alerts",), 90),
        Job("transaction reports", ("booked trades", "security master"),
            ("transaction reports",), 45),
        Job("features", ("tick history", "security master"), ("features",), 120),
        Job("backtests", ("features",), ("backtests",), 360),
    ]
```

***Listing 1.1.** The end-of-day jobs of the miniature firm: inputs, outputs, duration in minutes and, for the risk batch, an optional fixed start. code/platforms/01-the-platform-map/python/pl_platmap.py*

## 1.3 Systems of record and golden sources

A number exists in several places at once: the closing price of a stock is in the exchange’s feed, in the tick store, in the pricing library’s snapshot, in the P&L service’s marks and in every spreadsheet that copied it. When two of those copies differ, the firm needs a rule that says which one is right.

**Definition 1.4 (System of record, golden source).**

The *system of record* of a dataset is the one system whose copy is authoritative: corrections are made there and nowhere else, and every other copy is derived from it. A *golden source* is a system of record that the firm designates for a kind of data used across many systems (instruments, counterparties, closing marks), so that every consumer reads from it or from a copy that the map shows to be derived from it.

**Proposition 1.5 (One owner, one answer).**

If every dataset has exactly one [system of record](#def-pl-the-platform-map-sor), every copy is derived through the map from that system, and two reports read the same version of a dataset, then they agree about it.

**Proof.** Each report’s value of the dataset is a function of the version it read, applied along the map’s derivations; the same version through the same derivations gives the same value. A disagreement therefore requires two systems of record (two versions that are not derived from one another), a copy made outside the map (a derivation nobody declared), or two versions read at different times. ∎

The proof is short because the proposition is almost a definition, and its value is in the three ways it can fail. The first two are structural and a program can find them in the map; the third is about time, and the rest of this chapter is about it. [Listing 1.2](#lst-pl-the-platform-map-validate) is the structural check. On the map as first drawn it finds that two systems claim the marks, the pricing library and product control.

**Example 1.6 (Two claims on the marks).**

The pricing library computes the end-of-day marks from the closing prices and the curve and volatility snapshot; product control independently verifies them against broker quotes and consensus data (One Quant Book 6, chapter 27). Both call their output “the marks”. They are two datasets: the marks, owned by the pricing library, and the verified marks with their adjustments, owned by product control. The map’s fix is to name the second dataset and make its dependence on the first explicit; the check then passes, and a disagreement between the two becomes a number product control reports rather than a discovery.

```python
    def validate(self) -> list[str]:
        out = []
        sysnames = {s.name for s in self.systems}
        for d in self.datasets:
            owners = d.system if isinstance(d.system, tuple) else (d.system,)
            if len(owners) != 1:
                names = ", ".join(owners)
                out.append(f"{d.name}: {len(owners)} systems of record ({names})")
            for o in owners:
                if o not in sysnames:
                    out.append(f"{d.name}: unknown system {o}")
            makers = [j.name for j in self.jobs if d.name in j.outputs]
            if d.ready is None and not makers:
                out.append(f"{d.name}: no producer and no arrival time")
            if len(makers) > 1:
                out.append(f"{d.name}: written by {len(makers)} jobs ({', '.join(makers)})")
        for j in self.jobs:
            for x in j.inputs + j.outputs:
                if x not in self._ds:
                    out.append(f"{j.name}: unknown dataset {x}")
        if self._cycle():
            out.append("cycle in the batch: " + " -> ".join(self._cycle()))
        return out
```

***Listing 1.2.** Structural checks: one system of record per dataset, known names, one producer per dataset, and no cycle. code/firm/platmap/firm_platmap.py*

**As of September 2026 — How far banks are from one owner per number.**

The Basel Committee’s principles for effective risk data aggregation and risk reporting (BCBS 239, published on 9 January 2013) ask banks to generate accurate, complete and timely risk data from a data architecture that supports it. In its progress report of 28 November 2023 (consulted September 2026), the Committee found that of the 31 global systemically important banks it assessed, only two were fully compliant with all the principles.

## 1.4 Batch and stream

**Definition 1.7 (Batch processing, stream processing, end-of-day batch).**

*Batch processing* runs a job on a complete, bounded dataset (a day of fills, a vendor file) and produces its output once. *Stream processing* runs a job continuously on an unbounded sequence of events, updating its output as each event arrives. The *end-of-day batch* is the set of batch jobs that run after the close to produce the day’s official positions, marks, P&L, risk and reports.

A trading firm runs both, often for the same quantity. The position service of chapter 17 keeps positions as a stream during the day, from each fill as it arrives, because traders and the risk gate need them now; the end-of-day positions job recomputes them as a batch from the complete day, the clearing statement and the security master, because the books are closed on those. The two must agree, and when they do not the difference is a break (chapter 22).

In a batch chain the questions are about time: when will each output be ready, which job decides when the chain ends, and how late may each input be. With unlimited machines (chapter 13 removes that assumption), two passes over the map answer them.

**Method 1.8 (Forward and backward passes).**

Order the jobs so that every job comes after the producers of its inputs. *Forward:* for each job, its start is the latest ready time of its inputs (or its fixed start), its end is start plus duration, and its outputs are ready at its end; external inputs are ready at their arrival times. *Backward:* for each job in reverse order, its latest end is the earliest of its own deadline and the latest starts of the jobs that read its outputs; its latest start is that minus its duration. The *slack* of a job is its latest start minus its scheduled start. The latest safe arrival of an external input is the earliest latest start among the jobs that read it.

**Proposition 1.9 (The dependency-driven schedule is the earliest).**

With unlimited capacity and fixed durations, the forward pass of a map whose jobs all start on their inputs gives each job the earliest start any schedule can give it, and the chain ends at the end of its longest chain of jobs: the chain that ends last, traced back through the input that was ready last at each step.

**Proof.** By induction in dependency order. A job cannot start before all its inputs are ready; if each input is ready as early as possible (the induction hypothesis, or its arrival time for an external input), the job starts as early as possible by starting at the latest of those times. Following the last-ready input back from the job that ends last gives a chain of jobs with no gap between one’s end and the next’s start, whose total is the end time. ∎

[Listing 1.3](#lst-pl-the-platform-map-passes) is both passes. The forward pass also records every read of an input that was not yet complete when the job started, which is only possible for a job with a fixed start: that is the check the hook needed.

```python
    def _forward(self, arrivals):
        arrivals = arrivals or {}
        ready = {d.name: arrivals.get(d.name, d.ready) for d in self.datasets
                 if d.ready is not None}
        sched = {}
        for j in self._order():
            deps = max((ready[i] for i in j.inputs), default=0.0)
            start = deps if j.start is None else j.start
            end = start + j.duration
            sched[j.name] = (start, end)
            for o in j.outputs:
                ready[o] = end
        return sched, ready

    def stale_reads(self, arrivals: dict | None = None) -> list[tuple]:
        sched, ready = self._forward(arrivals)
        out = []
        for j in self._order():
            s = sched[j.name][0]
            for i in j.inputs:
                if ready[i] > s + 1e-9:
                    out.append((j.name, i, s, ready[i]))
        return out

    def latest_starts(self, deadline: float) -> dict[str, float]:
        """Backward pass: a job must end by its own deadline (default `deadline`) and
        before the latest start of every job that reads one of its outputs."""
        late: dict[str, float] = {}
        for j in reversed(self._order()):
            end = j.deadline if j.deadline is not None else deadline
            for k in self.jobs:
                if k.name in late and set(j.outputs) & set(k.inputs):
                    end = min(end, late[k.name])
            late[j.name] = end - j.duration
        return late
```

***Listing 1.3.** The forward pass (start on the last input, or at the fixed time), the stale reads it implies, and the backward pass of latest starts. code/firm/platmap/firm_platmap.py*

![The miniature firm’s end-of-day batch on a nominal night, every job started on its inputs. The longest chain (red) runs from the corporate-action file through the golden copy of the reference data, the research features and the overnight backtests, and ends at 03:55; the dashed line is the 07:00 deadline. Data: fig_platmap.py.](https://one-course.com/images/onecourse/chapters/quant-15/pl-the-platform-map/fig-e20ec9c08297.svg)

***Figure 1.3.** The miniature firm’s [end-of-day batch](#def-pl-the-platform-map-batch) on a nominal night, every job started on its inputs. The longest chain (red) runs from the corporate-action file through the golden copy of the reference data, the research features and the overnight backtests, and ends at 03:55; the dashed line is the 07:00 deadline. Data: `fig_platmap.py`.*

On a nominal night ([Figure 1.3](#fig-pl-the-platform-map-gantt)) the chain ends at 03:55, 185 minutes before its deadline. Its longest chain is not the risk batch, which is the longest job (six hours, from 20:15 to 02:15), but the research branch: the golden copy waits for the 19:30 corporate-action file, the features wait for the golden copy, and six hours of backtests follow. A firm that bought more machines for the risk grid would not move the end of the night by a minute. [Table 1.1](#tab-pl-the-platform-map-latest) turns the backward pass into the question operations actually ask the vendor: how late can your file be?

| input | arrives | latest safe arrival | margin (min) |
| --- | --- | --- | --- |
| fills, capture | 16:05 | 00:40, 22:25 | 515, 380 |
| closing prices | 16:20 | 00:40 | 500 |
| curve and volatility snapshot | 16:45 | 00:20 | 455 |
| vendor reference data | 18:00 | 22:35 | 275 |
| vendor corporate actions | 19:30 | 22:35 | 185 |
| clearing statement | 21:30 | 05:50 | 500 |

***Table 1.1.** Arrival and latest safe arrival of each external input for every morning report to be ready by 07:00, from the backward pass on the nominal map. The corporate-action file has the smallest margin. Data: `pl_platmap.latest_arrivals`.*

## 1.5 Reading a platform map

Two more questions complete the reading. The first is structural: what does an input reach?

**Definition 1.10 (Blast radius).**

The *blast radius* of a dataset is the set of datasets derived from it, directly or through other jobs: everything that is wrong if it is wrong and late if it is late.

On the miniature firm’s map the day’s fills, the vendor reference file and the corporate-action file each reach nine of the fifteen derived datasets; the clearing statement reaches two (the reconciliation breaks and the signed P&L); the market-data capture reaches four. A wrong corporate action is as dangerous as a wrong fill, and much less watched: fills are reconciled with the venue’s drop copy (Book 13, chapter 21) within seconds, while a corporate action is checked, if at all, when a report looks odd.

The second question is about time, and it is the hook’s. A job that starts on a fixed clock time reads whatever version of its inputs exists at that moment. If an input’s producer runs late, the job reads the previous version without an error, and its output silently disagrees with outputs computed from the new one.

**Proposition 1.11 (When a fixed start reads stale data).**

A job with fixed start $s$ reads a complete version of input $x$ if and only if $x$ is ready by $s$. If $x$ is the output of a chain of dependency-driven jobs of total duration $D_x$ that starts on an external input arriving at time $a$, the job reads stale data exactly when $a > s - D_x$.

**Proof.** The first part is the definition of the forward pass. For the second, the forward pass gives $x$ the ready time $a + D_x$ when $a$ is the last input of the chain, and the read is stale exactly when $a + D_x > s$. ∎

The risk batch of the hook starts at 21:00 ($s = 300$ minutes after the close). Its inputs include the positions, ready 45 minutes after the corporate-action file arrives (25 minutes of golden copy, 20 of positions), so it reads stale positions as soon as the file arrives after 20:15. On a nominal night the file comes at 19:30, 45 minutes of margin, and nobody notices the fixed start. On the night of the hook it came at 21:30: the golden copy was ready at 21:55, the positions at 22:15, and the risk batch had been running on the old versions since 21:00 ([Figure 1.4](#fig-pl-the-platform-map-hook)). The P&L job, dependency-driven, read the new ones at 22:15. The three teams’ morning was the price of one fixed start.

![The night of the hook: the corporate-action file arrives at 21:30. The risk batch, started by the clock at 21:00, has already read the security master and the positions (their new versions are ready at the dashed lines); the P&L job, started on its inputs, reads the new ones. Data: fig_platmap.py (firm_platmap.stale_reads).](https://one-course.com/images/onecourse/chapters/quant-15/pl-the-platform-map/fig-f85a2d815c01.svg)

***Figure 1.4.** The night of the hook: the corporate-action file arrives at 21:30. The risk batch, started by the clock at 21:00, has already read the security master and the positions (their new versions are ready at the dashed lines); the P&L job, started on its inputs, reads the new ones. Data: `fig_platmap.py` (`firm_platmap.stale_reads`).*

The fix is not to move the fixed start to 22:00, which only moves the threshold; it is to make the risk batch start on its inputs. Dependency-driven, on the night of the hook it starts at 22:15 and ends at 04:15, well inside the deadline, and the map reports no stale read. Fixed starts exist for good reasons (a machine is free only after 21:00; a vendor file has no completion signal), and where one is kept, the map’s check turns the silent disagreement into a failed precondition: the job refuses to start on an incomplete input, or starts and marks its output as built on a stale version.

**Remark 1.12 (What the map does not know).**

The passes assume fixed durations and unlimited machines. Durations vary from night to night, jobs share a cluster, and a job can fail and be rerun; chapter 13 adds queues and failures, and chapter 20 sizes the risk batch itself. The map also knows only the flows someone declared: a spreadsheet that copies the marks at 18:00 is invisible to it, which is why the map is enforced by making systems read their inputs through it (chapters 6, 7 and 16) rather than drawn beside them.

## 1.6 Tutorial: the miniature firm as a map

**Goal.** Describe the miniature firm as data, check it, schedule a night, and find the hook’s defect by program. **End state:** Figures [1.3](#fig-pl-the-platform-map-gantt) and [1.4](#fig-pl-the-platform-map-hook), [Table 1.1](#tab-pl-the-platform-map-latest) and the [blast radius](#def-pl-the-platform-map-blast) of every input.

1. **Declare.** Read `pl_platmap` : systems by layer, datasets with their [system of record](#def-pl-the-platform-map-sor) and, for external inputs, their arrival times, and the jobs of [Listing 1.1](#lst-pl-the-platform-map-jobs) .
2. **Validate.** `draft().validate()` reports that the marks have two systems of record; `firm().validate()` reports nothing.
3. **Schedule.** `firm().schedule()` and `longest_chain()` give [Figure 1.3](#fig-pl-the-platform-map-gantt) : the chain golden copy, features, backtests, ending at 03:55. `slack(900)` and `latest_arrivals` give [Table 1.1](#tab-pl-the-platform-map-latest) .
4. **Replay the hook.** `firm(risk_start=300).stale_reads(HOOK_NIGHT)` returns the two stale reads of the risk batch; `firm().stale_reads(HOOK_NIGHT)` returns none.
5. **[Blast radius](#def-pl-the-platform-map-blast).** `firm().blast_radius()` and `downstream` of the corporate-action file.

**What to change next.** Give the backtests a deadline of 06:00 and find the new latest arrival of the corporate-action file; add a margin-forecast job that reads the positions and the clearing statement and must be ready by 06:00, and see which input’s latest arrival it moves.

## 1.7 Build: the platform map

**Purpose.** The firm’s map of systems, datasets and jobs as data, with the checks and the time questions every later chapter uses: the tick store (chapter 4) and the reference data (chapter 6) declare their outputs in it, the data-quality rules of chapter 7 and the observability of chapter 28 attach to its datasets, and the scheduler of chapter 13 runs its jobs.

**Interface.** `System(name, layer, owner)`, `Dataset(name, system, cadence, ready)`, `Job(name, inputs, outputs, duration, start, deadline)`; `PlatformMap(systems, datasets, jobs)` with `validate`, `producer`, `downstream`, `upstream`, `schedule`, `ready_times`, `stale_reads`, `latest_starts`, `slack`, `longest_chain`, `blast_radius`; times in minutes after the close.

**Rules.** One [system of record](#def-pl-the-platform-map-sor) per dataset and one producer per derived dataset; no cycles; a fixed-start job is allowed but every read before its input is ready is reported; the map is the only place a flow is declared.

**Acceptance tests.** `code/firm/platmap/tests/`: the forward pass and the longest chain on a small map, and how a late input moves both; the backward pass and slack by hand; a stale read found and absent; two systems of record, an unknown system, an orphan dataset and a cycle reported; closures and [blast radius](#def-pl-the-platform-map-blast).

**Stretch.** Durations as distributions and the probability of meeting the deadline by simulation; versions of datasets and the version each job read; a map generated from the systems’ own declarations rather than written by hand.

Sources and further reading

- Basel Committee on Banking Supervision, *Principles for effective risk data aggregation and risk reporting* (BCBS 239), January 2013, and *Progress in adopting the Principles* , November 2023.
- T. Akidau et al., “The dataflow model”, *Proceedings of the VLDB Endowment* 8(12), 2015 (batch as a special case of stream processing).

## 1.8 Exercises

**Exercise 1.1 ★.**

The P&L job reads the positions (ready at 20:15), the marks (17:05) and the security master (19:55), and takes 25 minutes. When does it start and end on a nominal night, and which input decides?

**Solution of Exercise 1.1.**

It starts when its last input is ready: the positions, at 20:15. It ends at 20:40. The marks (17:05) and the security master (19:55) were ready earlier; the positions decide.

**Exercise 1.2 ★.**

The vendor’s reference-data file arrives at 22:50 instead of 18:00. Using [Table 1.1](#tab-pl-the-platform-map-latest), which report is late, and by how many minutes?

**Solution of Exercise 1.2.**

The file’s latest safe arrival is 22:35; at 22:50 it is 15 minutes late. The golden copy runs from 22:50 to 23:15, the features until 01:15, and the backtests end at 07:15: the research report is 15 minutes late. Every other report still has slack (the risk batch, for example, runs from 23:35 to 05:35).

**Exercise 1.3 ★.**

Why does the clearing statement have the smallest [blast radius](#def-pl-the-platform-map-blast) of the seven inputs, and does a small [blast radius](#def-pl-the-platform-map-blast) mean the input needs less checking?

**Solution of Exercise 1.3.**

Only two datasets derive from it: the reconciliation breaks and the signed P&L, because it is read only by the reconciliation. A small [blast radius](#def-pl-the-platform-map-blast) says how far an error travels through the batch, not how much it matters: a wrong clearing statement can hide a real break in the firm’s positions, which is exactly what the reconciliation exists to find. It needs checking at the point of use (chapter 22), even though little is built on it.

**Exercise 1.4 ★★.**

To finish earlier, the risk team moves the risk batch’s fixed start from 21:00 to 20:30. Up to what arrival time of the corporate-action file does it read complete inputs, and what margin does that leave on a nominal night?

**Solution of Exercise 1.4.**

The positions are ready 45 minutes after the file (25 minutes of golden copy, 20 of positions), so the batch reads complete inputs only if the file arrives by $270 - 45 = 225$ minutes after the close, 19:45. The nominal file comes at 19:30: 15 minutes of margin instead of 45. The earlier start makes the stale read three times more likely to happen on a slightly late night.

**Exercise 1.5 ★★.**

Product control’s independently verified marks differ from the pricing library’s by a few basis points on some positions. Explain why these are two datasets rather than two copies of one, and which of the two the P&L sign-off should read.

**Solution of Exercise 1.5.**

The verified marks are computed from other inputs (broker quotes, consensus data) by another system and carry adjustments; they are not derived from the library’s marks by the map, so they are a different dataset with its own [system of record](#def-pl-the-platform-map-sor), product control. The sign-off reads the verified marks, and reports the difference between the two as a number, because the purpose of independent verification is that the desk’s own marks do not close the books.

**Exercise 1.6 ★★.**

During the day the position service keeps positions as a stream; after the close a batch job recomputes them. Give two reasons the two can differ at 17:00, and say which one the books are closed on.

**Solution of Exercise 1.6.**

(a) Late information: busts, corrections and late fills arrive after the stream has applied the originals, and a drop copy may show fills the trading session missed (chapter 17). (b) Reference data: the batch applies the evening’s corporate actions from the security master (a split changes quantities), the stream does not. The books are closed on the batch, reconciled with the clearing statement; the difference with the stream is investigated as a break.

**Exercise 1.7 ★★★.**

*Coding.* Add a job `margin forecast` that reads the positions and the clearing statement, takes 60 minutes and must be ready by 06:00. What is its latest start, does it change any other job’s latest start, and what is the clearing statement’s new latest safe arrival?

**Solution of Exercise 1.7.**

Its latest start is $840 - 60 = 780$ minutes after the close, 05:00. No other job’s latest start moves: the positions already had to be ready by 01:00 for the risk batch, well before 05:00. The clearing statement’s latest safe arrival becomes 05:00 instead of 05:50, because the forecast now reads it with an earlier deadline than the reconciliation’s.

**Exercise 1.8 ★★★.**

*Find the flaw.* “The overnight batch is too close to its deadline. We will double the risk grid’s machines, which halves the risk batch to three hours, and the whole batch will end three hours earlier.”

**Solution of Exercise 1.8.**

The risk batch is not on the longest chain. Halved to three hours it ends at 23:15 instead of 02:15, but the batch still ends at 03:55, with the backtests: the longest chain is the corporate-action file, the golden copy, the features and the backtests. And doubling machines halves a job only if it parallelises perfectly and nothing else limits it (chapters 13 and 20).

## 1.9 Problem: The Morning the Numbers Disagreed

**Problem 1.1.**

Weekend problem — a late file and a fixed start

The miniature firm’s map (`pl_platmap`), its nominal night and the night the corporate-action file arrived at 21:30.

**Part I — The nominal night.**

1. When do the golden copy, the positions and the risk batch start and end, with every job started on its inputs?
2. Which job ends last, and when?
3. What is the longest chain, and why is the risk batch, the longest job, not on it?
4. How many minutes of slack does the batch have before 07:00?
5. What is the risk batch’s slack?

**Part II — How late can the files be?**

6. What is the latest safe arrival of the vendor’s corporate-action file, and which job sets it?
7. What is the latest safe arrival of the market-data capture, and why is it earlier than that of the fills?
8. Which input has the largest margin between its nominal and its latest safe arrival?
9. How many datasets does the corporate-action file reach?
10. Which datasets does a wrong clearing statement reach?

**Part III — The night of the hook.**

11. The file arrives at 21:30. When are the security master and the positions ready?
12. The risk batch starts at 21:00 by the clock. Which of its inputs does it read stale?
13. When does the P&L job run, and which versions does it read?
14. Up to what arrival time of the file would the fixed-start risk batch have read complete inputs?
15. Started on its inputs, when would the risk batch have run on that night, and would it have met the deadline?

**Part IV — The verdict.**

16. State the *named result* : the end and slack of the nominal night, the latest safe arrival of the corporate-action file, and the arrival time beyond which the fixed-start risk batch reads stale data.
17. Why did the disagreement produce no error in any system?
18. Which of the three failure modes of [Proposition 1.5](#prop-pl-the-platform-map-oneowner) occurred?
19. Give two ways to keep a fixed start without the silent disagreement.
20. In one sentence: what does a platform map give that a system diagram does not?

**Solution of Problem 1.1.**

1. Golden copy 19:30–19:55; positions 19:55–20:15; risk batch 20:15–02:15.
2. The backtests, at 03:55.
3. The corporate-action file, the golden copy, the features and the backtests: $25 + 120 + 360$ minutes from 19:30. The risk batch’s own chain (file, golden copy, positions, risk batch) ends at 02:15, 100 minutes earlier, because the research branch carries two long jobs.
4. $900 - 715 = 185$ minutes.
5. 285 minutes: its latest start is 01:00 and it starts at 20:15.
6. 22:35, set by the golden copy, whose latest start is fixed by the features (latest start 23:00) and the backtests (01:00).
7. 22:25: the capture feeds the tick compaction, which feeds the features and so the backtests; the fills feed the positions, whose latest start (00:40) is set by the risk batch, a later deadline in the chain.
8. The fills: 515 minutes (16:05 to 00:40).
9. Nine: the security master, the positions, the P&L, the risk report, the reconciliation breaks, the signed P&L, the transaction reports, the features and the backtests.
10. The reconciliation breaks and the signed P&L.
11. 21:55 and 22:15.
12. The security master and the positions (the scenarios and the marks were ready long before).
13. From 22:15 to 22:40, on the new security master and the new positions.
14. 20:15: the file must arrive 45 minutes before the 21:00 start.
15. From 22:15 to 04:15, 165 minutes before the deadline.
16. **Named result.** On a nominal night the batch ends at 03:55, 185 minutes before 07:00, on the research chain; the corporate-action file may arrive until 22:35; a risk batch started by the clock at 21:00 reads a stale security master and stale positions whenever the file arrives after 20:15, only 45 minutes after its nominal time, while started on its inputs it would have finished at 04:15 on the night of the hook.
17. The stale versions exist and are valid; a fixed-start job has no way to know that a newer version is coming.
18. The third: two versions of the same datasets read at different times.
19. Make the job check that each input’s version for the business date is complete before it starts (a completion marker written by the producer), and refuse or wait otherwise; or let it start but record the version of each input it read, so that the report is marked stale and rerun when the new version lands.
20. It can be checked and simulated: a program answers what a file reaches, when each report is ready, and which job reads an incomplete input.

## 1.10 Interview questions

**Interview question 1.1 ★ developer, risk.**

What is a [system of record](#def-pl-the-platform-map-sor), and why does a firm need one per dataset?

**Solution of Interview question 1.1.**

The one system whose copy of a dataset is authoritative: corrections are made there, and everything else is derived from it. Without one per dataset, two systems can each be right by their own rules and disagree, and nobody can say which report is correct.

*What the interviewer is looking for: Corrections in one place; copies derived, not edited; disagreement as a symptom of two owners.*

**Interview question 1.2 ★ developer.**

When would you compute positions as a stream and when as a batch?

**Solution of Interview question 1.2.**

As a stream when a decision needs it now (pre-trade risk, a trader’s screen, a kill switch); as a batch when the answer must be complete and final (the books, the clearing reconciliation, regulatory reports), because late fills, busts and corporate actions are only known after the close. Most firms do both and reconcile them.

*What the interviewer is looking for: Latency against completeness; both kept, and reconciled.*

**Interview question 1.3 ★★ developer, risk.**

The morning risk report and the P&L disagree on one position. How do you find out why?

**Solution of Interview question 1.3.**

Trace both back through the map to the datasets they read and the versions and times they read them. Check the three causes in order: two systems of record for one input, a copy made outside the declared flows (a spreadsheet, a cache), or the same dataset read at two different times (a job that ran before a late input was complete). Then fix the cause, not the number.

*What the interviewer is looking for: Lineage and versions rather than guessing; the timing cause.*

**Interview question 1.4 ★★ developer.**

Should a batch job start at a fixed time or when its inputs are ready? What goes wrong with each?

**Solution of Interview question 1.4.**

On its inputs, where possible: it starts as early as it can and never reads an incomplete input, but it needs a reliable completion signal from every producer and can start at an awkward time. A fixed start is predictable for machines and people, but reads stale data silently when an input is late. If a fixed start is kept, check completeness before starting and record the versions read.

*What the interviewer is looking for: The stale-read failure named; completion markers.*

**Interview question 1.5 ★★ developer.**

How would you find out how late a vendor file can arrive before the morning reports are late?

**Solution of Interview question 1.5.**

A backward pass over the dependency graph: from each report’s deadline, subtract job durations back along every path to the file; the earliest of the resulting latest starts among the jobs that read it is the file’s latest safe arrival. Measure the durations’ spread and keep a margin, because the pass assumes fixed durations and free machines.

*What the interviewer is looking for: Latest start, backward over the graph, with a margin for variable durations.*

**Interview question 1.6 ★★★ developer, risk.**

You join a firm with no diagram of its systems. How do you build a map that stays correct, and what do you check first?

**Solution of Interview question 1.6.**

Start from the outputs people rely on (the P&L, the risk report, regulatory reports) and trace back what each reads, from where and when, recording it as data rather than as a drawing; make systems declare their inputs so the map is generated and checked, not maintained by hand. First checks: one [system of record](#def-pl-the-platform-map-sor) per dataset, no undeclared copies, and every scheduled job’s inputs complete at its start.

*What the interviewer is looking for: Map as data, derived from the systems; ownership and timing checks first.*
