Research, Data and Risk Platforms · Technology
16Signal Serving and Parameter Management
A risk parameter entered as 50 meant fifty basis points to the person who typed it and fifty percent to the strategy that read it. Nothing checked the unit; the change went to every strategy at once; the first sign was the fill rate. A parameter is code that nobody compiles: it changes behaviour as much as a release does, it is edited by more people, and it usually has no types, no tests and no history. This chapter gives parameters and research signals the machinery code already has — types, review, history, gradual rollout — and measures, on ten simulated strategies, what each piece of that machinery saves when fifty means the wrong thing.
16.1 From research signal to production input
A research signal becomes a production input when a strategy reads it on a schedule and trades on it. Between the two sit the questions a research notebook never asks: which version is live, how old the value may be before it is not to be used, what the strategy does when it is missing, and who changed it. Book 12 (chapters 24 and 25) answered them for machine-learning features and models — the feature store, feature freshness, the model registry and its audit trail. This chapter applies the same answers to the plainer inputs every strategy has: signals computed by research jobs and parameters typed by people.
16.2 The signal service
Definition 16.1 (Signal service)
A signal service serves research signals to production strategies: each value with its version and publication time, served only while it is fresher than a stated bound, and replaced by a declared fallback when it is stale or missing, with the status reported to the reader.
The fallback is part of the signal’s contract, chosen when the signal is designed, not by the strategy’s author at three in the morning: for a directional signal, usually zero (no view); for a risk input, usually the most conservative value.
class SignalService:
"""Versioned values; a reader gets the latest if fresh, else the declared fallback."""
def __init__(self, max_age: float, fallback: float):
self.max_age, self.fallback = max_age, fallback
self.values: dict = {} # name -> [(ts, version, value)]
def publish(self, name: str, value: float, version: str, ts: float) -> None:
self.values.setdefault(name, []).append((ts, version, value))
def get(self, name: str, now: float):
vals = [v for v in self.values.get(name, []) if v[0] <= now]
if not vals:
return self.fallback, None, "missing-fallback"
ts, version, value = max(vals)
if now - ts > self.max_age:
return self.fallback, version, "stale-fallback"
return value, version, "fresh"
With a freshness bound of five minutes and a fallback of zero, a momentum signal published at time 0 and again at 60 seconds is served as 0.7 (version 3, fresh) at 200 seconds and as 0 with the status “stale-fallback” at 600 seconds; a signal never published is served as the fallback with the status “missing-fallback”. The strategy knows which it got.
16.3 Configuration as data
Definition 16.2 (Parameter store, configuration as data)
A parameter store holds every production parameter of every strategy, each value with the time from which it applies and the time it was recorded, the change that set it and its approvals. Configuration as data is the practice of keeping configuration in such a store, as typed, versioned records that tools can validate, compare and query, rather than as text edited in place on production machines.
A value’s two times are the valid time and knowledge time of Book 7 (chapter 3): the time from which it applies and the time the store learned it. They let the store answer the question every post-mortem asks — what was live at 10:31 last Tuesday — in both of its senses: as the firm knew it then, and as it is known now that a correction has been recorded.
16.4 Schemas, approvals and history
Definition 16.3 (Configuration schema)
A configuration schema declares, for each parameter, its canonical unit and the units it may be entered in, its bounds, the largest change allowed in one step, and the size of change that needs a second approver; every proposed value is converted and checked against it before anyone approves it.
The chapter’s parameter is an inventory limit, as a share of the strategy’s capital: canonical unit a fraction, entered in basis points, percent or as a fraction, bounded by 2%, changed by at most 50 basis points at a time, and needing two approvers for any change above 20 basis points — the four-eyes principle of One Quant Book 16 (chapter 12) applied to a number.
def propose(self, strategy, name, value, unit, effective, recorded, author) -> Change:
s = self.schemas[name]
c = Change(len(self.changes), strategy, name, math.nan, f"{value} {unit}",
effective, recorded, author)
self.changes.append(c)
if unit not in s.units:
c.reasons.append(f"unit {unit!r} not accepted for {name} "
f"(accepted: {', '.join(s.units)})")
else:
c.value = float(value) * s.units[unit]
if not s.lo <= c.value <= s.hi:
c.reasons.append(f"{c.value:g} {s.unit} outside [{s.lo:g}, {s.hi:g}]")
old = self.as_of(strategy, name, recorded, recorded)
step = abs(c.value - old) if old is not None else 0.0
if step > s.max_step:
c.reasons.append(f"step {step:g} {s.unit} above the allowed {s.max_step:g}")
c.needed = 2 if step > s.two_approvals_above else 1
c.status = "refused" if c.reasons else "pending"
body = {"cid": c.cid, "strategy": strategy, "name": name, "entered": c.entered,
"status": c.status, "reasons": c.reasons}
self.audit.append("propose", body, recorded)
return c
The fifty of the hook, entered without a unit, is refused: a value with no unit is not accepted. Entered as “50 percent”, it is refused twice over: 0.5 is outside , and a step of 0.495 exceeds the allowed 0.005. Entered as “50 bp” it is accepted as 0.005. A raise to 75 basis points (a step of 25) waits after one approval, cannot be approved by its author, and applies after the second. Every proposal, approval and application is a line of Book 12’s hash-chained audit trail (firm.exptrack): editing any line breaks the chain, and verify says where.
def approve(self, cid: int, approver: str, ts: float) -> Change:
c = self.changes[cid]
if c.status != "pending" or approver == c.author or approver in c.approvals:
self.audit.append("approve-refused", {"cid": cid, "approver": approver}, ts)
return c
c.approvals.append(approver)
self.audit.append("approve", {"cid": cid, "approver": approver}, ts)
if len(c.approvals) >= c.needed:
c.status = "applied"
rec = max(ts, c.recorded)
row = (c.effective, rec, c.value, cid)
self.rows.setdefault((c.strategy, c.name), []).append(row)
body = {"cid": cid, "value": c.value, "effective": c.effective}
self.audit.append("apply", body, rec)
return c
def as_of(self, strategy, name, valid, known=None):
"""The value in force at `valid`, as the store knew it at `known` (default: now)."""
rows = self.rows.get((strategy, name), [])
rows = [r for r in rows if r[0] <= valid and (known is None or r[1] <= known)]
if not rows:
return None
return max(rows, key=lambda r: (r[0], r[1]))[2]
Example 16.4 (What was live at 10:31 last Tuesday)
The raise to 75 basis points was recorded on Tuesday at 09:43 and took effect at 10:00. On Thursday a correction was recorded: from Tuesday 10:30 the limit should have been 60. Asked on Wednesday what was live at 10:31 on Tuesday, the store answers 75 basis points; asked now, it answers 60. Both answers are right, and a post-mortem needs both: what the strategy used, and what it should have used.
16.5 Rolling out a change
Definition 16.5 (Feature flag)
A feature flag is a named switch in configuration that turns a behaviour or a new value on for chosen strategies, users or a share of traffic without a new release, so that a change can be rolled out to one first and withdrawn by switching the flag off.
Definition 16.6 (Configuration drift)
Configuration drift is a difference between the configuration a system is declared to run and the one it actually runs: a server not updated, a value edited in place, a flag left from an old release.
Flags are how the canary deployment and staged rollout of Book 7 (chapter 21) apply to parameters: the new value is behind a flag switched on for one strategy, watched, and extended. Drift is what a flag, left in place, becomes: the chapter’s drift compares, strategy by strategy, the declared value with the one each strategy reports it is running, and on the demonstration finds the one strategy running 0.5 where the store says 0.005. The best-known case of both is Knight Capital’s: on 1 August 2012 a flag once used for an old function was repurposed for new code, the new code had not reached one of eight servers, and orders carrying the flag triggered the old code there; the loss exceeded $460 million (Book 7, chapter 21, and Book 13, chapter 22, tell it in full).
16.6 Fifty of what?
The chapter’s incident runs ten quoting strategies (chapter 11’s quoter, one lot a side) for an hour each on Book 10’s simulator, each on its own firm.tape session, with $4 million of capital each. The intended limit, 50 basis points of capital, is two lots at $100; read as 50 percent it is two hundred. The change arrives ten minutes into the session. A monitor compares each strategy’s fills in every five-minute window with its usual count for that window — here, the same session without the change, a better baseline than any real monitor has — and fires above 1.3 times; a firing rolls the parameter back. Exposure is the inventory beyond what the intended limit allows (three lots, with one order pending), in dollars.
def detection(base: dict, wrong: dict, start: float) -> float | None:
"""End of the first window after `start` whose fills exceed RATIO times the usual count."""
b, w = windows(base), windows(wrong)
for k in range(len(w)):
if k * WINDOW >= start and w[k] > RATIO * max(b[k], 1):
return (k + 1) * WINDOW
return None
fig_paramstore.py.| policy | detection (minutes) | peak excess ($k) | excess $-minutes (k) |
|---|---|---|---|
| all at once | 5 | 410 | 1 348 |
| staged (one canary for 30 min) | 15 | 30 | 53 |
| schema check | 0 (refused) | 0 | 0 |
Table 16.1 is the result. All at once is detected soonest — ten monitors watch instead of one, and the first fires at five minutes — and costs most: $410 000 of excess inventory at the peak and 1.35 million dollar-minutes. The canary’s own monitor takes fifteen minutes to fire (its session is quiet), but only the canary holds the wrong value: $30 000 at the peak and 53 000 dollar-minutes, 25 times less. The schema check costs nothing, because the error never becomes a value. The strategies differ widely in how soon their fill rate shows the error: five of the ten after five minutes, three after fifteen, one after thirty and one after fifty.
fig_paramstore.py.Remark 16.7 (Detection is not protection)
Faster detection did not make all at once safer: it had ten chances to notice and ten strategies accruing exposure while it did. A staged rollout limits the damage to the canary whatever the monitor’s speed; a schema prevents the damage whatever the rollout. The monitor here is optimistic — it knows each strategy’s counterfactual fill count — and a real one would be slower, which widens every gap.
As of September 2026 — Configuration at scale
Facebook’s configuration-management paper (Tang and others, SOSP 2015) described a system handling thousands of online configuration changes a day and trillions of configuration checks, used for product rollouts, experiments, load balancing and model deployment. The SEC’s order on Knight Capital (Release 34-70694, 2013) records the repurposed flag, the server not updated, about 45 minutes of unintended orders and a loss of more than $460 million.
16.7 Tutorial: a units error, three ways
Goal. Serve a signal with a freshness bound, put a parameter under a schema with approvals and history, and measure what a units error costs under three policies. End state: Figure 16.2 and Table 16.1.
- Signals:
pl_paramstore.signals(). - Parameters:
governance(dir): the refused entries, the two approvals, the bitemporal answers, flags, drift, the audit. - Incident:
incident()runs twenty hours of simulated trading (ten strategies, with and without the error). - Policies: read the detection times and exposures from the result.
What to change next. Lower the monitor’s threshold to 1.2 and count its false alarms on the unchanged runs; stage the rollout in three steps (one, three, all) instead of two.
16.8 Build: the parameter store
Purpose. Every production input a strategy reads — signal or parameter — typed, versioned, approved, auditable and rolled out gradually.
Interface. Schema; ParamStore(root) with register, propose, approve, as_of, history, audit; Flags; drift; SignalService(max_age, fallback) with publish and get.
Rules. A value enters with its unit or not at all; the author never approves; changes above a threshold need two approvers; values carry valid and knowledge times; the audit trail is hash-chained; signals carry a version, a freshness bound and a declared fallback.
Acceptance tests. code/firm/paramstore/tests/: units, bounds and steps; approvals (author refused, two distinct approvers); bitemporal queries and history; the audit trail’s verification and its detection of an edited line; flags, drift and the signal service’s three statuses.
Stretch. Percentage rollouts by hash of the strategy name; automatic rollback wired to the monitor; drift checks against each strategy’s own report at start-up.
Sources and further reading
- C. Tang et al., “Holistic configuration management at Facebook”, SOSP 2015.
- US Securities and Exchange Commission, In the Matter of Knight Capital Americas LLC, Release No. 34-70694, 2013.
- One Quant Book 7, chapters 3 and 21; One Quant Book 12, chapters 24 and 25; One Quant Book 16, chapter 12.
16.9 Exercises
Exercise 16.1 ★
How many lots is an inventory limit of 50 basis points of $4 million at $100 a share, and of 50 percent?
Solution
Solution of Exercise 16.1.
50 basis points of $4 million is $20 000, 200 shares at $100: two lots. 50 percent is $2 million, 20 000 shares: two hundred lots.
Exercise 16.2 ★
Why does the raise from 50 to 75 basis points need two approvers under the chapter’s schema, and the correction to 60 only one?
Solution
Solution of Exercise 16.2.
The raise is a step of 25 basis points (0.0025), above the 20 that need a second approver; the correction from 75 to 60 is a step of 15, below it.
Exercise 16.3 ★
What does the signal service return at 400 seconds in the chapter’s example, and why?
Solution
Solution of Exercise 16.3.
0.7, version 3, fresh: the latest value was published at 60 seconds, 340 seconds earlier, within the five-minute bound.
Exercise 16.4 ★★
Why was the all-at-once rollout detected sooner than the staged one, and why is it still worse?
Solution
Solution of Exercise 16.4.
All at once, ten strategies had the error and the first of ten monitors fired after five minutes; staged, only the canary’s monitor could fire, and its quiet session took fifteen. But all at once, ten strategies accrued exposure until detection, staged only one: 1.35 million dollar-minutes against 53 000.
Exercise 16.5 ★★
A post-mortem asks what limit a strategy used at 10:31 on Tuesday. Which of the store’s two answers does it want, and when does it want the other?
Solution
Solution of Exercise 16.5.
What the strategy used is the value as known then (75 basis points): that is what explains its behaviour. What it should have used is the value as known now (60): that is what the correction and any restatement of risk or P&L need.
Exercise 16.6 ★★
Which of the chapter’s controls would have stopped the Knight Capital incident as the SEC describes it, and which would not?
Solution
Solution of Exercise 16.6.
A drift check comparing what each server runs with what is declared would have shown the eighth server’s old code; retiring flags instead of repurposing them, and a staged rollout watched per server, would have limited the damage. A schema on parameter values would not: the flag’s value was valid, its meaning on that server was not.
Exercise 16.7 ★★★
Coding. Rerun incident with the canary chosen as the strategy whose monitor fires last (q10). What happens to the staged rollout’s detection time and exposure?
Solution
Solution of Exercise 16.7.
With q10 as the canary, its monitor does not fire within the thirty minutes, the other nine receive the value at forty minutes, and the first of their monitors fires five minutes later: detection 35 minutes after the change, a peak of $1.31 million and 5.44 million dollar-minutes — worse than all at once. A canary whose monitor cannot see the error only delays the rollout; the canary must be a strategy on which the error would show.
Exercise 16.8 ★★★
Find the flaw. “Our parameters live in a YAML file in the repository, so they are versioned, reviewed and safe.”
Solution
Solution of Exercise 16.8.
Versioned and reviewed, yes; typed and checked, no: a reviewer sees “50” in a diff and nothing says in what unit, whether it is in bounds or how far it moved; the file says what is declared, not what each strategy runs; and a merge goes everywhere at once. A schema, approvals by size, a drift check and a flag-based rollout are still needed.
16.10 Problem: Fifty of What?
Problem 16.1
Weekend problem — a units error under three policies
The ten strategies, the schema and the monitor of the chapter.
Part I — The store.
- What does the schema of the inventory limit declare?
- What happens to 50 entered without a unit, as percent, and as basis points?
- Who may approve a change, and how many approvals does it need?
- What are a value’s two times, and what question do they answer?
- What does the audit trail guarantee?
Part II — Signals and flags.
- What does the signal service return, and with what status, when a signal is stale?
- Who chooses a signal’s fallback, and when?
- What is a feature flag for, and what does it become if left in place?
- What does the drift check find in the demonstration?
- What did a repurposed flag do at Knight Capital?
Part III — The incident.
- What do 50 basis points and 50 percent mean in lots?
- How does the monitor decide, and why is it optimistic?
- When does each strategy’s monitor fire?
- What are the exposure and time to detection all at once?
- And staged?
Part IV — The verdict.
- State the named result: the exposure accrued before detection by the units error under an all-at-once change, a staged rollout and a schema check, and the minutes to detection in each case.
- Why is detection not protection?
- Which control would you build first?
- What would you add to catch a units error that stays within bounds?
- In one sentence: what does a parameter need that a variable does not?
Solution
Solution of Problem 16.1.
- A canonical unit (a fraction of capital), accepted units (basis points, percent, fraction), bounds 0 to 2%, a largest step of 50 basis points, and two approvers above 20.
- Refused (no unit); refused (outside the bounds and too large a step); accepted as 0.005.
- Anyone but the author; one, or two above the threshold, distinct.
- Valid time and knowledge time: what was in force at a moment, as known at another moment.
- That any edit of a recorded line is detected, and where.
- The declared fallback, with the status “stale-fallback” (or “missing-fallback” if never published).
- The signal’s owner, when the signal is designed.
- Rolling a change out to chosen strategies and back; a source of drift and of latent behaviour.
- One strategy running 0.5 where the store declares 0.005.
- Orders carrying a flag repurposed for new code triggered old code on a server the new code had not reached.
- Two lots and two hundred lots.
- It fires in the first five-minute window after the change whose fills exceed 1.3 times the usual count; it knows each strategy’s counterfactual count, which a real monitor does not.
- After 5 minutes for five strategies, 15 for three, 30 for one and 50 for one.
- Five minutes; $410 000 at the peak and 1.35 million dollar-minutes.
- Fifteen minutes; $30 000 and 53 000 dollar-minutes.
- Named result. All at once: detected after 5 minutes, with $410 000 of excess inventory at the peak and 1.35 million dollar-minutes; staged: 15 minutes, $30 000 and 53 000 dollar-minutes; schema check: refused at entry, no exposure.
- Detection counts only after the damage; what limits damage is how many strategies hold the error until then, and whether it can enter at all.
- The schema with units: it costs least and stops the whole class of error.
- Tighter bounds per strategy, a maximum step relative to the current value, and a comparison of the new value’s effect in a simulator before the rollout.
- A unit, bounds, an owner, approvals, a history and a rollout plan.
16.11 Interview questions
Interview question 16.1 ★ developer
How do you stop a units error in a production parameter?
Solution
Solution of Interview question 16.1.
Make the unit part of the value: typed parameters with a canonical unit, conversion on entry, bounds and step limits in a schema, and refusal of anything without a unit.
What the interviewer is looking for: Types and schemas, not care.
Interview question 16.2 ★★ developer, researcher
A research signal feeds live strategies. What happens when the job that computes it fails?
Solution
Solution of Interview question 16.2.
The signal goes stale; the service serves the declared fallback with a status the strategy can act on, and an alert goes to the signal’s owner. The fallback and the freshness bound are part of the signal’s contract.
What the interviewer is looking for: Freshness bounds and declared fallbacks.
Interview question 16.3 ★★ developer
How would you roll out a change to a risk limit across a hundred strategies?
Solution
Solution of Interview question 16.3.
Through the parameter store with two approvals, behind a flag: to one canary strategy on which the change’s effect is visible, watched against its usual behaviour, then to a few, then to all, with automatic rollback when a monitor fires.
What the interviewer is looking for: Staged rollout with a meaningful canary and rollback.
Interview question 16.4 ★★ developer
What is configuration drift, and how would you detect it?
Solution
Solution of Interview question 16.4.
The difference between declared and running configuration; detect it by having every process report the configuration it loaded (with a hash) and comparing with the store, continuously and at start-up.
What the interviewer is looking for: Running configuration reported and compared.
Interview question 16.5 ★★★ developer
Design a parameter management system for a trading firm.
Solution
Solution of Interview question 16.5.
A store of typed, schema-checked, bitemporal values with owners; proposals and approvals by change size; a hash-chained audit trail; flags for staged rollouts with rollback; drift detection against running processes; a signal service with versions, freshness and fallbacks; and post-mortem queries of what was live when.
What the interviewer is looking for: Types, approvals, history, rollout, drift.