---
title: "Signal Serving and Parameter Management"
book: "Research, Data and Risk Platforms"
subject: quant
language: en
chapter: 16
exercises: 8
source: https://one-course.com/books/quant/15/en/chapter/16-signal-serving-and-parameter-management
---

# Chapter 16 — Signal Serving and Parameter Management

A risk parameter entered as 50 meant fifty basis points to the person who typed it and fifty percent to the strategy that read it. Nothing checked the unit; the change went to every strategy at once; the first sign was the fill rate. A parameter is code that nobody compiles: it changes behaviour as much as a release does, it is edited by more people, and it usually has no types, no tests and no history. This chapter gives parameters and research signals the machinery code already has — types, review, history, gradual rollout — and measures, on ten simulated strategies, what each piece of that machinery saves when fifty means the wrong thing.

## 16.1 From research signal to production input

A research signal becomes a production input when a strategy reads it on a schedule and trades on it. Between the two sit the questions a research notebook never asks: which version is live, how old the value may be before it is not to be used, what the strategy does when it is missing, and who changed it. Book 12 (chapters 24 and 25) answered them for machine-learning features and models — the feature store, feature freshness, the model registry and its audit trail. This chapter applies the same answers to the plainer inputs every strategy has: signals computed by research jobs and parameters typed by people.

## 16.2 The signal service

**Definition 16.1 (Signal service).**

A *signal service* serves research signals to production strategies: each value with its version and publication time, served only while it is fresher than a stated bound, and replaced by a declared fallback when it is stale or missing, with the status reported to the reader.

The fallback is part of the signal’s contract, chosen when the signal is designed, not by the strategy’s author at three in the morning: for a directional signal, usually zero (no view); for a risk input, usually the most conservative value.

```python
class SignalService:
    """Versioned values; a reader gets the latest if fresh, else the declared fallback."""

    def __init__(self, max_age: float, fallback: float):
        self.max_age, self.fallback = max_age, fallback
        self.values: dict = {}                 # name -> [(ts, version, value)]

    def publish(self, name: str, value: float, version: str, ts: float) -> None:
        self.values.setdefault(name, []).append((ts, version, value))

    def get(self, name: str, now: float):
        vals = [v for v in self.values.get(name, []) if v[0] <= now]
        if not vals:
            return self.fallback, None, "missing-fallback"
        ts, version, value = max(vals)
        if now - ts > self.max_age:
            return self.fallback, version, "stale-fallback"
        return value, version, "fresh"
```

***Listing 16.1.** The signal service: the latest value if it is fresh, else the declared fallback, always with the version and a status. code/firm/paramstore/firm_paramstore.py*

With a freshness bound of five minutes and a fallback of zero, a momentum signal published at time 0 and again at 60 seconds is served as 0.7 (version 3, fresh) at 200 seconds and as 0 with the status “stale-fallback” at 600 seconds; a signal never published is served as the fallback with the status “missing-fallback”. The strategy knows which it got.

## 16.3 Configuration as data

**Definition 16.2 (Parameter store, configuration as data).**

A *parameter store* holds every production parameter of every strategy, each value with the time from which it applies and the time it was recorded, the change that set it and its approvals. *Configuration as data* is the practice of keeping configuration in such a store, as typed, versioned records that tools can validate, compare and query, rather than as text edited in place on production machines.

A value’s two times are the valid time and knowledge time of Book 7 (chapter 3): the time from which it applies and the time the store learned it. They let the store answer the question every post-mortem asks — what was live at 10:31 last Tuesday — in both of its senses: as the firm knew it then, and as it is known now that a correction has been recorded.

## 16.4 Schemas, approvals and history

**Definition 16.3 (Configuration schema).**

A *configuration schema* declares, for each parameter, its canonical unit and the units it may be entered in, its bounds, the largest change allowed in one step, and the size of change that needs a second approver; every proposed value is converted and checked against it before anyone approves it.

The chapter’s parameter is an inventory limit, as a share of the strategy’s capital: canonical unit a fraction, entered in basis points, percent or as a fraction, bounded by 2%, changed by at most 50 basis points at a time, and needing two approvers for any change above 20 basis points — the four-eyes principle of One Quant Book 16 (chapter 12) applied to a number.

```python
    def propose(self, strategy, name, value, unit, effective, recorded, author) -> Change:
        s = self.schemas[name]
        c = Change(len(self.changes), strategy, name, math.nan, f"{value} {unit}",
                   effective, recorded, author)
        self.changes.append(c)
        if unit not in s.units:
            c.reasons.append(f"unit {unit!r} not accepted for {name} "
                             f"(accepted: {', '.join(s.units)})")
        else:
            c.value = float(value) * s.units[unit]
            if not s.lo <= c.value <= s.hi:
                c.reasons.append(f"{c.value:g} {s.unit} outside [{s.lo:g}, {s.hi:g}]")
            old = self.as_of(strategy, name, recorded, recorded)
            step = abs(c.value - old) if old is not None else 0.0
            if step > s.max_step:
                c.reasons.append(f"step {step:g} {s.unit} above the allowed {s.max_step:g}")
            c.needed = 2 if step > s.two_approvals_above else 1
        c.status = "refused" if c.reasons else "pending"
        body = {"cid": c.cid, "strategy": strategy, "name": name, "entered": c.entered,
                "status": c.status, "reasons": c.reasons}
        self.audit.append("propose", body, recorded)
        return c
```

***Listing 16.2.** A proposal: converted to the canonical unit, checked against bounds and the step limit, and told how many approvals it needs; every proposal, refused or not, goes to the audit trail. code/firm/paramstore/firm_paramstore.py*

The fifty of the hook, entered without a unit, is refused: a value with no unit is not accepted. Entered as “50 percent”, it is refused twice over: 0.5 is outside $[0, 0.02]$, and a step of 0.495 exceeds the allowed 0.005. Entered as “50 bp” it is accepted as 0.005. A raise to 75 basis points (a step of 25) waits after one approval, cannot be approved by its author, and applies after the second. Every proposal, approval and application is a line of Book 12’s hash-chained audit trail (`firm.exptrack`): editing any line breaks the chain, and `verify` says where.

```python
    def approve(self, cid: int, approver: str, ts: float) -> Change:
        c = self.changes[cid]
        if c.status != "pending" or approver == c.author or approver in c.approvals:
            self.audit.append("approve-refused", {"cid": cid, "approver": approver}, ts)
            return c
        c.approvals.append(approver)
        self.audit.append("approve", {"cid": cid, "approver": approver}, ts)
        if len(c.approvals) >= c.needed:
            c.status = "applied"
            rec = max(ts, c.recorded)
            row = (c.effective, rec, c.value, cid)
            self.rows.setdefault((c.strategy, c.name), []).append(row)
            body = {"cid": cid, "value": c.value, "effective": c.effective}
            self.audit.append("apply", body, rec)
        return c

    def as_of(self, strategy, name, valid, known=None):
        """The value in force at `valid`, as the store knew it at `known` (default: now)."""
        rows = self.rows.get((strategy, name), [])
        rows = [r for r in rows if r[0] <= valid and (known is None or r[1] <= known)]
        if not rows:
            return None
        return max(rows, key=lambda r: (r[0], r[1]))[2]
```

***Listing 16.3.** Approval by someone other than the author, as many as the change needs, and the bitemporal query: the value in force at a valid time, as known at a knowledge time. code/firm/paramstore/firm_paramstore.py*

![The life of a parameter change in the store: typed on entry, checked by the schema, approved by others, rolled out behind a flag to one strategy and then to all, every step recorded with its times in a history that cannot be edited unnoticed.](https://one-course.com/images/onecourse/chapters/quant-15/pl-signal-serving-and-parameter-management/fig-d4359f913f5c.svg)

***Figure 16.1.** The life of a parameter change in the store: typed on entry, checked by the schema, approved by others, rolled out behind a flag to one strategy and then to all, every step recorded with its times in a history that cannot be edited unnoticed.*

**Example 16.4 (What was live at 10:31 last Tuesday).**

The raise to 75 basis points was recorded on Tuesday at 09:43 and took effect at 10:00. On Thursday a correction was recorded: from Tuesday 10:30 the limit should have been 60. Asked on Wednesday what was live at 10:31 on Tuesday, the store answers 75 basis points; asked now, it answers 60. Both answers are right, and a post-mortem needs both: what the strategy used, and what it should have used.

## 16.5 Rolling out a change

**Definition 16.5 (Feature flag).**

A *feature flag* is a named switch in configuration that turns a behaviour or a new value on for chosen strategies, users or a share of traffic without a new release, so that a change can be rolled out to one first and withdrawn by switching the flag off.

**Definition 16.6 (Configuration drift).**

*Configuration drift* is a difference between the configuration a system is declared to run and the one it actually runs: a server not updated, a value edited in place, a flag left from an old release.

Flags are how the canary deployment and staged rollout of Book 7 (chapter 21) apply to parameters: the new value is behind a flag switched on for one strategy, watched, and extended. Drift is what a flag, left in place, becomes: the chapter’s `drift` compares, strategy by strategy, the declared value with the one each strategy reports it is running, and on the demonstration finds the one strategy running 0.5 where the store says 0.005. The best-known case of both is Knight Capital’s: on 1 August 2012 a flag once used for an old function was repurposed for new code, the new code had not reached one of eight servers, and orders carrying the flag triggered the old code there; the loss exceeded $460 million (Book 7, chapter 21, and Book 13, chapter 22, tell it in full).

## 16.6 Fifty of what?

The chapter’s incident runs ten quoting strategies (chapter 11’s quoter, one lot a side) for an hour each on Book 10’s simulator, each on its own `firm.tape` session, with $4 million of capital each. The intended limit, 50 basis points of capital, is two lots at $100; read as 50 percent it is two hundred. The change arrives ten minutes into the session. A monitor compares each strategy’s fills in every five-minute window with its usual count for that window — here, the same session without the change, a better baseline than any real monitor has — and fires above 1.3 times; a firing rolls the parameter back. Exposure is the inventory beyond what the intended limit allows (three lots, with one order pending), in dollars.

```python
def detection(base: dict, wrong: dict, start: float) -> float | None:
    """End of the first window after `start` whose fills exceed RATIO times the usual count."""
    b, w = windows(base), windows(wrong)
    for k in range(len(w)):
        if k * WINDOW >= start and w[k] > RATIO * max(b[k], 1):
            return (k + 1) * WINDOW
    return None
```

***Listing 16.4.** The monitor: the end of the first five-minute window after the change in which fills exceed 1.3 times the usual count. code/platforms/16-signal-serving-and-parameter-management/python/pl_paramstore.py*

![Inventory beyond the intended limit, summed over the strategies that received the wrong value, from the change until the monitor fires and the value is rolled back, on a ten-second grid. All at once, ten strategies accrue it for five minutes; staged, one strategy for fifteen; with a schema, the value never enters. Data: fig_paramstore.py.](https://one-course.com/images/onecourse/chapters/quant-15/pl-signal-serving-and-parameter-management/fig-779ece98127e.svg)

***Figure 16.2.** Inventory beyond the intended limit, summed over the strategies that received the wrong value, from the change until the monitor fires and the value is rolled back, on a ten-second grid. All at once, ten strategies accrue it for five minutes; staged, one strategy for fifteen; with a schema, the value never enters. Data: `fig_paramstore.py`.*

| policy | detection (minutes) | peak excess ($k) | excess $-minutes (k) |
| --- | --- | --- | --- |
| all at once | 5 | 410 | 1 348 |
| staged (one canary for 30 min) | 15 | 30 | 53 |
| schema check | 0 (refused) | 0 | 0 |

***Table 16.1.** The units error under three policies: minutes from the change to detection, the peak of the excess inventory summed over affected strategies, and its integral over time.*

[Table 16.1](#tab-pl-signal-serving-and-parameter-management-policies) is the result. All at once is detected soonest — ten monitors watch instead of one, and the first fires at five minutes — and costs most: $410 000 of excess inventory at the peak and 1.35 million dollar-minutes. The canary’s own monitor takes fifteen minutes to fire (its session is quiet), but only the canary holds the wrong value: $30 000 at the peak and 53 000 dollar-minutes, 25 times less. The schema check costs nothing, because the error never becomes a value. The strategies differ widely in how soon their fill rate shows the error: five of the ten after five minutes, three after fifteen, one after thirty and one after fifty.

![Minutes from the units error to the fill-rate monitor’s first firing, strategy by strategy, each on its own one-hour session. How soon an error shows depends on the market the strategy happens to trade: five strategies show it in the first window, one only at the end of the hour. The staged rollout’s canary is q01. Data: fig_paramstore.py.](https://one-course.com/images/onecourse/chapters/quant-15/pl-signal-serving-and-parameter-management/fig-b2ef7ce337a7.svg)

***Figure 16.3.** Minutes from the units error to the fill-rate monitor’s first firing, strategy by strategy, each on its own one-hour session. How soon an error shows depends on the market the strategy happens to trade: five strategies show it in the first window, one only at the end of the hour. The staged rollout’s canary is q01. Data: `fig_paramstore.py`.*

**Remark 16.7 (Detection is not protection).**

Faster detection did not make all at once safer: it had ten chances to notice and ten strategies accruing exposure while it did. A staged rollout limits the damage to the canary whatever the monitor’s speed; a schema prevents the damage whatever the rollout. The monitor here is optimistic — it knows each strategy’s counterfactual fill count — and a real one would be slower, which widens every gap.

**As of September 2026 — Configuration at scale.**

Facebook’s configuration-management paper (Tang and others, SOSP 2015) described a system handling thousands of online configuration changes a day and trillions of configuration checks, used for product rollouts, experiments, load balancing and model deployment. The SEC’s order on Knight Capital (Release 34-70694, 2013) records the repurposed flag, the server not updated, about 45 minutes of unintended orders and a loss of more than $460 million.

## 16.7 Tutorial: a units error, three ways

**Goal.** Serve a signal with a freshness bound, put a parameter under a schema with approvals and history, and measure what a units error costs under three policies. **End state:** [Figure 16.2](#fig-pl-signal-serving-and-parameter-management-timeline) and [Table 16.1](#tab-pl-signal-serving-and-parameter-management-policies).

1. **Signals** : `pl_paramstore.signals()` .
2. **Parameters** : `governance(dir)` : the refused entries, the two approvals, the bitemporal answers, flags, drift, the audit.
3. **Incident** : `incident()` runs twenty hours of simulated trading (ten strategies, with and without the error).
4. **Policies** : read the detection times and exposures from the result.

**What to change next.** Lower the monitor’s threshold to 1.2 and count its false alarms on the unchanged runs; stage the rollout in three steps (one, three, all) instead of two.

## 16.8 Build: the parameter store

**Purpose.** Every production input a strategy reads — signal or parameter — typed, versioned, approved, auditable and rolled out gradually.

**Interface.** `Schema`; `ParamStore(root)` with `register`, `propose`, `approve`, `as_of`, `history`, `audit`; `Flags`; `drift`; `SignalService(max_age, fallback)` with `publish` and `get`.

**Rules.** A value enters with its unit or not at all; the author never approves; changes above a threshold need two approvers; values carry valid and knowledge times; the audit trail is hash-chained; signals carry a version, a freshness bound and a declared fallback.

**Acceptance tests.** `code/firm/paramstore/tests/`: units, bounds and steps; approvals (author refused, two distinct approvers); bitemporal queries and history; the audit trail’s verification and its detection of an edited line; flags, drift and the [signal service](#def-pl-signal-serving-and-parameter-management-service)’s three statuses.

**Stretch.** Percentage rollouts by hash of the strategy name; automatic rollback wired to the monitor; drift checks against each strategy’s own report at start-up.

Sources and further reading

- C. Tang et al., “Holistic configuration management at Facebook”, SOSP 2015.
- US Securities and Exchange Commission, *In the Matter of Knight Capital Americas LLC* , Release No. 34-70694, 2013.
- One Quant Book 7, chapters 3 and 21; One Quant Book 12, chapters 24 and 25; One Quant Book 16, chapter 12.

## 16.9 Exercises

**Exercise 16.1 ★.**

How many lots is an inventory limit of 50 basis points of $4 million at $100 a share, and of 50 percent?

**Solution of Exercise 16.1.**

50 basis points of $4 million is $20 000, 200 shares at $100: two lots. 50 percent is $2 million, 20 000 shares: two hundred lots.

**Exercise 16.2 ★.**

Why does the raise from 50 to 75 basis points need two approvers under the chapter’s schema, and the correction to 60 only one?

**Solution of Exercise 16.2.**

The raise is a step of 25 basis points (0.0025), above the 20 that need a second approver; the correction from 75 to 60 is a step of 15, below it.

**Exercise 16.3 ★.**

What does the [signal service](#def-pl-signal-serving-and-parameter-management-service) return at 400 seconds in the chapter’s example, and why?

**Solution of Exercise 16.3.**

0.7, version 3, fresh: the latest value was published at 60 seconds, 340 seconds earlier, within the five-minute bound.

**Exercise 16.4 ★★.**

Why was the all-at-once rollout detected sooner than the staged one, and why is it still worse?

**Solution of Exercise 16.4.**

All at once, ten strategies had the error and the first of ten monitors fired after five minutes; staged, only the canary’s monitor could fire, and its quiet session took fifteen. But all at once, ten strategies accrued exposure until detection, staged only one: 1.35 million dollar-minutes against 53 000.

**Exercise 16.5 ★★.**

A post-mortem asks what limit a strategy used at 10:31 on Tuesday. Which of the store’s two answers does it want, and when does it want the other?

**Solution of Exercise 16.5.**

What the strategy used is the value as known then (75 basis points): that is what explains its behaviour. What it should have used is the value as known now (60): that is what the correction and any restatement of risk or P&L need.

**Exercise 16.6 ★★.**

Which of the chapter’s controls would have stopped the Knight Capital incident as the SEC describes it, and which would not?

**Solution of Exercise 16.6.**

A drift check comparing what each server runs with what is declared would have shown the eighth server’s old code; retiring flags instead of repurposing them, and a staged rollout watched per server, would have limited the damage. A schema on parameter values would not: the flag’s value was valid, its meaning on that server was not.

**Exercise 16.7 ★★★.**

*Coding.* Rerun `incident` with the canary chosen as the strategy whose monitor fires last (q10). What happens to the staged rollout’s detection time and exposure?

**Solution of Exercise 16.7.**

With q10 as the canary, its monitor does not fire within the thirty minutes, the other nine receive the value at forty minutes, and the first of their monitors fires five minutes later: detection 35 minutes after the change, a peak of $1.31 million and 5.44 million dollar-minutes — worse than all at once. A canary whose monitor cannot see the error only delays the rollout; the canary must be a strategy on which the error would show.

**Exercise 16.8 ★★★.**

*Find the flaw.* “Our parameters live in a YAML file in the repository, so they are versioned, reviewed and safe.”

**Solution of Exercise 16.8.**

Versioned and reviewed, yes; typed and checked, no: a reviewer sees “50” in a diff and nothing says in what unit, whether it is in bounds or how far it moved; the file says what is declared, not what each strategy runs; and a merge goes everywhere at once. A schema, approvals by size, a drift check and a flag-based rollout are still needed.

## 16.10 Problem: Fifty of What?

**Problem 16.1.**

Weekend problem — a units error under three policies

The ten strategies, the schema and the monitor of the chapter.

**Part I — The store.**

1. What does the schema of the inventory limit declare?
2. What happens to 50 entered without a unit, as percent, and as basis points?
3. Who may approve a change, and how many approvals does it need?
4. What are a value’s two times, and what question do they answer?
5. What does the audit trail guarantee?

**Part II — Signals and flags.**

6. What does the [signal service](#def-pl-signal-serving-and-parameter-management-service) return, and with what status, when a signal is stale?
7. Who chooses a signal’s fallback, and when?
8. What is a [feature flag](#def-pl-signal-serving-and-parameter-management-flag) for, and what does it become if left in place?
9. What does the drift check find in the demonstration?
10. What did a repurposed flag do at Knight Capital?

**Part III — The incident.**

11. What do 50 basis points and 50 percent mean in lots?
12. How does the monitor decide, and why is it optimistic?
13. When does each strategy’s monitor fire?
14. What are the exposure and time to detection all at once?
15. And staged?

**Part IV — The verdict.**

16. State the *named result* : the exposure accrued before detection by the units error under an all-at-once change, a staged rollout and a schema check, and the minutes to detection in each case.
17. Why is detection not protection?
18. Which control would you build first?
19. What would you add to catch a units error that stays within bounds?
20. In one sentence: what does a parameter need that a variable does not?

**Solution of Problem 16.1.**

1. A canonical unit (a fraction of capital), accepted units (basis points, percent, fraction), bounds 0 to 2%, a largest step of 50 basis points, and two approvers above 20.
2. Refused (no unit); refused (outside the bounds and too large a step); accepted as 0.005.
3. Anyone but the author; one, or two above the threshold, distinct.
4. Valid time and knowledge time: what was in force at a moment, as known at another moment.
5. That any edit of a recorded line is detected, and where.
6. The declared fallback, with the status “stale-fallback” (or “missing-fallback” if never published).
7. The signal’s owner, when the signal is designed.
8. Rolling a change out to chosen strategies and back; a source of drift and of latent behaviour.
9. One strategy running 0.5 where the store declares 0.005.
10. Orders carrying a flag repurposed for new code triggered old code on a server the new code had not reached.
11. Two lots and two hundred lots.
12. It fires in the first five-minute window after the change whose fills exceed 1.3 times the usual count; it knows each strategy’s counterfactual count, which a real monitor does not.
13. After 5 minutes for five strategies, 15 for three, 30 for one and 50 for one.
14. Five minutes; $410 000 at the peak and 1.35 million dollar-minutes.
15. Fifteen minutes; $30 000 and 53 000 dollar-minutes.
16. **Named result.** All at once: detected after 5 minutes, with $410 000 of excess inventory at the peak and 1.35 million dollar-minutes; staged: 15 minutes, $30 000 and 53 000 dollar-minutes; schema check: refused at entry, no exposure.
17. Detection counts only after the damage; what limits damage is how many strategies hold the error until then, and whether it can enter at all.
18. The schema with units: it costs least and stops the whole class of error.
19. Tighter bounds per strategy, a maximum step relative to the current value, and a comparison of the new value’s effect in a simulator before the rollout.
20. A unit, bounds, an owner, approvals, a history and a rollout plan.

## 16.11 Interview questions

**Interview question 16.1 ★ developer.**

How do you stop a units error in a production parameter?

**Solution of Interview question 16.1.**

Make the unit part of the value: typed parameters with a canonical unit, conversion on entry, bounds and step limits in a schema, and refusal of anything without a unit.

*What the interviewer is looking for: Types and schemas, not care.*

**Interview question 16.2 ★★ developer, researcher.**

A research signal feeds live strategies. What happens when the job that computes it fails?

**Solution of Interview question 16.2.**

The signal goes stale; the service serves the declared fallback with a status the strategy can act on, and an alert goes to the signal’s owner. The fallback and the freshness bound are part of the signal’s contract.

*What the interviewer is looking for: Freshness bounds and declared fallbacks.*

**Interview question 16.3 ★★ developer.**

How would you roll out a change to a risk limit across a hundred strategies?

**Solution of Interview question 16.3.**

Through the [parameter store](#def-pl-signal-serving-and-parameter-management-store) with two approvals, behind a flag: to one canary strategy on which the change’s effect is visible, watched against its usual behaviour, then to a few, then to all, with automatic rollback when a monitor fires.

*What the interviewer is looking for: Staged rollout with a meaningful canary and rollback.*

**Interview question 16.4 ★★ developer.**

What is [configuration drift](#def-pl-signal-serving-and-parameter-management-drift), and how would you detect it?

**Solution of Interview question 16.4.**

The difference between declared and running configuration; detect it by having every process report the configuration it loaded (with a hash) and comparing with the store, continuously and at start-up.

*What the interviewer is looking for: Running configuration reported and compared.*

**Interview question 16.5 ★★★ developer.**

Design a parameter management system for a trading firm.

**Solution of Interview question 16.5.**

A store of typed, schema-checked, bitemporal values with owners; proposals and approvals by change size; a hash-chained audit trail; flags for staged rollouts with rollback; drift detection against running processes; a [signal service](#def-pl-signal-serving-and-parameter-management-service) with versions, freshness and fallbacks; and post-mortem queries of what was live when.

*What the interviewer is looking for: Types, approvals, history, rollout, drift.*
