Research, Data and Risk Platforms · Technology
30Surveillance and Compliance Technology
On 27 September 2022 the SEC announced that sixteen firms had agreed to pay more than $1.1 billion in penalties, and the CFTC the same day $710 million against eleven: from 2018 to 2021 their employees had discussed business in text messages and messaging applications on personal phones, and the firms had not kept the substantial majority of those messages. The rules were old; the phones were new; the archive had a hole the size of a pocket. Book 9, chapter 29, built the detectors that find spoofing, marking the close and wash trades in account-days. This chapter builds what surrounds them in a firm: the pipeline from orders to alerts, the queue in which people review alerts, the archive of what people said, the reports a firm sends its regulator about every trade, and the measurements that say whether any of it works.
30.1 The surveillance pipeline
Definition 30.1 (Trade surveillance system, surveillance alert)
A trade surveillance system turns a firm’s orders and trades into features per account and day, scores them with detectors for known forms of abuse, and raises a surveillance alert — an account, a day, a detector, a score — for each score above its threshold, for a person to review.
firm.survpipe extracts, from a stream of order events, the account-day features Book 9’s detectors use: orders placed and filled by size, large orders cancelled, and large cancellations that follow a fill on the account’s other side within a second. The chapter’s quarter then runs on Book 9’s own populations of account-days — market makers, quote refreshers, deep liquidity providers, directional traders, index funds, multi-algorithm firms — 2 200 legitimate account-days a day, with planted spoofing, marking-the-close and wash episodes arriving at random: 45 in the quarter’s 63 days. Each detector’s threshold is set on a separate calibration quarter at a false-positive rate of 0.2%, 0.5% or 1% of legitimate account-days.
30.2 From alert to case
Definition 30.2 (Alert triage, case management system)
Alert triage is the first review of an alert: dismissed with a reason, or opened as a case. A case management system holds alerts and cases with their states, owners, evidence, decisions and ages, so that every alert has an outcome and a record of how it was reached — an audit trail (Book 12, chapter 25) of the surveillance itself.
The team is two analysts who can each review eight alerts a day. Each day’s alerts join the queue, and the team takes sixteen: the oldest first, or the highest score first.
def work_queue(alerts_by_day: list, per_day: int, policy: str = "oldest") -> QueueRun:
"""Each day the day's alerts join the queue and the team reviews `per_day` of them:
the oldest first, or the highest score first."""
run, queue = QueueRun(), []
for day, new in enumerate(alerts_by_day):
queue.extend(new)
if policy == "score":
queue.sort(key=lambda a: (-a.score, a.day, a.aid))
else:
queue.sort(key=lambda a: (a.day, a.aid))
take, queue = queue[:per_day], queue[per_day:]
run.reviewed += [(a, day) for a in take]
run.backlog.append(len(queue))
return run
fig_survpipe.py.The threshold is a staffing decision (Figure 30.2). At 0.2% the detectors raise 32 of the 45 planted episodes and the team reviews every one on the day it appears; at 0.5%, 36. At 1% they raise 38, but the alerts outnumber the team’s capacity, and the order of the queue decides what that is worth (Figure 30.3). Oldest first, the team reviews 28 of the 38 by the quarter’s end and only 11 within five days, with a median wait of seven days; highest score first, 37 of 38 within five days, while the queue’s backlog is made of low scores. A queue that is always behind is not a surveillance system: an alert reviewed a month late finds a trader who has done it thirty more times.
fig_survpipe.py.30.3 Communications archiving and surveillance
Definition 30.3 (Communications archive, communications surveillance)
A communications archive keeps every business communication — e-mail, chat, voice — in a form that cannot be altered, for the period the rules require, with holds that suspend deletion for matters under investigation. Communications surveillance searches and scores the archive for signs of abuse — collusion, leaks, pressure — with lexicons and models, and raises alerts like trade surveillance.
Definition 30.4 (Off-channel communication)
An off-channel communication is a business communication on a channel the firm does not capture — a personal phone, a consumer messaging application — and is therefore missing from its archive.
The chapter’s archive holds a quarter of messages from 200 people, 252 025 of them. Forty conversations move off channel; 25 leave a pointer in the archive ("let’s take this to whatsapp", "call my cell") and 15 leave nothing. Two lexicons search it (Table 30.1).
| lexicon | messages flagged | pointers found (of 25) | conversations (of 40) |
|---|---|---|---|
| broad (single words) | 464 | 10 | 10 |
| narrow (phrases) | 10 | 10 | 10 |
The table is the hook in numbers. A lexicon finds only phrasings it knows, so a quarter’s catch is ten of forty conversations, and precision comes from writing phrases rather than words. And the fifteen conversations with no pointer are outside any search: an archive cannot surveil what it does not hold. The control that matters is upstream — the firm’s devices and approved channels capturing everything, and attestations and phone reviews for the rest — which is what the 2022 orders required the firms to fix.
30.4 Transaction reporting
Definition 30.5 (Transaction reporting)
Transaction reporting sends a regulator, for every transaction in reportable instruments, a report in a fixed field format — in the European Union the fields of RTS 22 — identifying the firm, the buyer and seller (by legal entity identifier, Book 2, chapter 28), the instrument (by ISIN), the time, quantity, price and venue; a correction is a cancellation of the report and a new one. Unlike the post-trade publication of Book 2’s trade reporting, it is not public: it feeds the regulator’s own surveillance.
A week of 2 500 equity executions for five clients is reported twice. The legacy generator copies each trade when it is booked and has four faults: a client LEI left empty on 1% of reports, a mistyped ISIN on 0.5%, amendments never reported (41 trades, 1.6%, are amended later in the week) and a batch occasionally sent twice. The fixed generator reads chapter 21’s event-sourced trade store, so an amendment becomes a cancellation of the first report and a new one, and takes identifiers from validated reference data.
def event_reports(trades: dict) -> list[dict]:
"""The fixed generator reads the event-sourced trade store: an amendment cancels the
first report and sends a new one (CANC, then NEWT); identifiers from reference data."""
out = []
for trn, t in trades.items():
if t["amended"]:
out.append(_report(trn, t, qty=t["quantity"] + 100)) # as first booked
out.append(_report(trn, t, status="CANC", qty=t["quantity"] + 100))
out.append(_report(trn, t))
return out
def reconcile(reports, trades) -> dict:
"""Net each reference's reports (a CANC removes the report it cancels) and compare with
the firm's trades by reference: missing, extra, and mismatched quantity or price."""
live: dict = {}
for r in reports:
if r["status"] == "CANC":
live.pop(r["trn"], None)
elif r["trn"] in live:
live[r["trn"] + "#dup"] = r
else:
live[r["trn"]] = r
out = {"missing": [], "extra": [], "mismatched": []}
for trn, t in trades.items():
r = live.get(trn)
if r is None:
out["missing"].append(trn)
elif r["quantity"] != t["quantity"] or abs(r["price"] - t["price"]) > 1e-9:
out["mismatched"].append(trn)
out["extra"] = sorted(k for k in live if k not in trades)
return out
fig_survpipe.py.Of the legacy generator’s 2 506 reports, 87 are wrong in one of three ways (Figure 30.4): 81 of the week’s 2 500 trades (3.2%) are either unreported or misreported, and 6 are reported twice. Validation catches the malformed ones before sending; only reconciliation with the trades catches the amendments, because each unreported amendment leaves a perfectly valid report of the trade as first booked. The event-sourced generator sends 2 582 reports — more, because each amendment costs two — and reconciles exactly.
30.5 Measuring a surveillance system
Every component of this chapter has a measure, and a surveillance function that cannot state them cannot say whether it works. For detectors, the true-positive rate at a false-positive rate, on labelled or planted episodes (Book 9), and the precision–recall curve (Book 12, chapter 20) as the base rate falls. For the queue, alerts a day against capacity, backlog, and the age at review of what matters. For the archive, the share of channels captured — which no search statistic can show — and a lexicon’s hits against known cases. For reporting, the shares that fail validation and reconciliation, every day. Planted episodes, run through the whole pipeline on a schedule like the replayed incidents of chapter 28, are the end-to-end test.
As of September 2026 — Off-channel penalties and transaction reports
On 27 September 2022 the SEC announced that sixteen firms (fifteen broker-dealers and one investment adviser) had agreed to pay combined penalties of more than $1.1 billion for failures to maintain and preserve electronic communications, their employees having used text messaging applications on personal devices from January 2018 to September 2021; the CFTC announced orders against eleven institutions with $710 million in penalties the same day. In the European Union, Delegated Regulation (EU) 2017/590 (RTS 22) sets out the transaction-report fields, among them the report status (NEWT or CANC), a transaction reference number unique to the executing firm, the executing entity, buyer and seller identified by LEI, the trading date and time, quantity, price, the venue by ISO 10383 MIC and the instrument by ISIN.
30.6 Tutorial: a hole the size of a pocket
Goal. Run a quarter of surveillance with a fixed team, search an archive, and report a week of trades two ways. End state: Figures 30.2 and 30.4.
- Features:
firm_survpipe.featureson a small stream of order events. - Alerts:
pl_survpipe.thresholds(fpr)andalerts(fpr)for the three rates. - Queue:
queue_study(fpr, policy), oldest first and highest score first. - Archive:
archive()andcomms_study()with both lexicons. - Reports:
report_study(), legacy against event-sourced.
What to change next. Give the team a third analyst and find the threshold at which the queue again empties daily; add a detector that looks for accounts whose alerts keep being dismissed.
30.7 Build: the surveillance pipeline
Purpose. Turn orders, trades and messages into reviewed cases and correct regulatory reports, with measured coverage, capacity and quality.
Interface. features, Alert, work_queue, search, retention, lei_ok, isin_ok, make_lei, make_isin, validate, reconcile.
Rules. Thresholds set against analyst capacity and reviewed quarterly; queues ordered by risk; every alert closed with a reason; archives immutable with holds; reports generated from the event-sourced trade store, validated before sending and reconciled every day; planted episodes run end to end.
Acceptance tests. code/firm/survpipe/tests/: features from an event stream, including a linked cancellation; the two queue orders; lexicon search and retention with a hold; LEI and ISIN check digits, report validation, and reconciliation with a cancellation, a duplicate and an unknown reference.
Stretch. Features from firm.exchsim runs with planted spoofing agents; a model-based message classifier against the lexicon; the full RTS 22 field set.
Sources and further reading
- US Securities and Exchange Commission, press release 2022-174, 27 September 2022; Commodity Futures Trading Commission, press release 8599-22, 27 September 2022.
- Commission Delegated Regulation (EU) 2017/590 (RTS 22), Annex I.
- One Quant Book 9, chapter 29 (detectors); Book 2, chapters 22 and 28 (trade reporting, LEIs); Book 12, chapter 20 (precision–recall).
30.8 Exercises
Exercise 30.1 ★
How many alerts a day can the team review, and at which false-positive rates does the queue stay empty?
Solution
Solution of Exercise 30.1.
Sixteen: two analysts reviewing eight each. The queue stays empty at 0.2% (4.8 alerts a day) and 0.5% (10.7); at 1% (21.2) it grows.
Exercise 30.2 ★
Why does the legacy generator’s unreported amendment pass validation?
Solution
Solution of Exercise 30.2.
Because the report itself is well formed: it describes the trade as it was first booked, with valid identifiers, a positive quantity and price. It is wrong only against the trade as it now stands, which only reconciliation compares.
Exercise 30.3 ★
Why does the fixed generator send more reports than there are trades?
Solution
Solution of Exercise 30.3.
Each of the 41 amended trades is reported three times — as first booked, cancelled, and as amended — so 2 500 trades need 2 582 reports; netting the cancellations leaves one live report per trade.
Exercise 30.4 ★★
At 1%, why does ordering the queue by score catch almost every planted episode in time, and what does it cost?
Solution
Solution of Exercise 30.4.
Planted episodes score far above most false positives, so a queue sorted by score reviews them the day they appear while the backlog fills with low-scoring alerts. The cost is that those low-scoring alerts may never be reviewed; a firm must still close them, by rule or by sampling, and watch that the backlog does not hide a new pattern with modest scores.
Exercise 30.5 ★★
Why does the narrow lexicon find as many pointers as the broad one with a fraction of the hits?
Solution
Solution of Exercise 30.5.
Because every pointer both find is a phrase that moves a conversation ("take this to whatsapp"), which the narrow lexicon matches exactly; the broad one also matches every ordinary message that happens to contain the word — 454 false hits for the same ten pointers.
Exercise 30.6 ★★
Why is the share of channels captured the measure that matters for communications surveillance?
Solution
Solution of Exercise 30.6.
Because a message that was never captured cannot be found by any search: fifteen of the forty conversations left no trace at all. Search statistics measure what is in the archive; only capture measures what is not.
Exercise 30.7 ★★★
Coding. Give the team a third analyst: rerun queue_study(0.01, "oldest", analysts=3). What happens to the backlog and to the planted episodes reviewed within five days?
Solution
Solution of Exercise 30.7.
With twenty-four reviews a day against 21.2 alerts, the queue empties every day: no backlog at the quarter’s end, and all 38 alerted episodes reviewed within five days, even oldest first.
Exercise 30.8 ★★★
Find the flaw. "Our lexicon found every off-channel conversation in last quarter’s archive, so the problem is under control."
Solution
Solution of Exercise 30.8.
The lexicon found the conversations that left a pointer phrased as it expects; it cannot find those that left a pointer phrased otherwise or no pointer at all. In the chapter’s quarter it found ten of forty. The claim confuses recall within the archive with coverage of the conversations.
30.9 Problem: A Hole the Size of a Pocket
Problem 30.1
Weekend problem — a hole the size of a pocket
The chapter’s quarter, archive and week of reports.
Part I — Alerts.
- What does the pipeline turn into what?
- How are the detectors’ thresholds set?
- How many planted episodes are there, and how many does each threshold alert?
- How many alerts a day does each threshold raise?
- What is the team’s capacity?
Part II — Cases.
- What happens to the queue at each threshold?
- At 1%, how many planted episodes are reviewed, and how many within five days, oldest first?
- And highest score first?
- What is the median wait oldest first at 1%?
- What does a case management system record?
Part III — Archive and reports.
- How many conversations move off channel, and how many leave a pointer?
- What do the two lexicons find?
- What were the legacy generator’s four faults?
- What do validation and reconciliation find in the legacy reports?
- What does the event-sourced generator change?
Part IV — The verdict.
- State the named result: the backlog after a quarter at each alert threshold for a team of fixed size, the planted episodes caught, and the share of transaction reports that fail validation or reconciliation before and after the fixes.
- Which threshold and queue order would you choose for this team?
- What would you tell the board about the archive?
- How would you test the whole pipeline every month?
- In one sentence: what makes a surveillance system work?
Solution
Solution of Problem 30.1.
- Orders and trades into account-day features, scores, alerts and cases; messages into lexicon hits; trades into validated, reconciled transaction reports.
- On a separate calibration quarter, at a false-positive rate of legitimate account-days: 0.2%, 0.5% or 1%.
- 45 planted; 32, 36 and 38 alerted.
- 4.8, 10.7 and 21.2.
- Sixteen alerts a day.
- Empty every day at 0.2% and 0.5%; at 1% it grows to 330.
- 28 reviewed by the quarter’s end, 11 within five days.
- 37 of 38, all within five days.
- Seven days.
- Each alert’s state, owner, evidence, decision with reason, and age.
- Forty; twenty-five leave a pointer.
- Both find the same ten pointers; the broad one flags 464 messages, the narrow one 10.
- Missing client LEIs, mistyped ISINs, unreported amendments, and batches sent twice.
- 41 reports fail validation, 6 are duplicates, and 40 misstate an amended trade: 81 of 2 500 trades unreported or misreported.
- Amendments become cancellations and new reports, identifiers come from reference data: nothing fails.
- Named result. With two analysts reviewing sixteen alerts a day, thresholds at false-positive rates of 0.2%, 0.5% and 1% leave backlogs of 0, 0 and 330 alerts after a quarter and alert 32, 36 and 38 of 45 planted episodes; at 1%, 11 are reviewed within five days oldest first and 37 highest score first. The legacy report generator leaves 3.2% of trades unreported or misreported (41 failing validation, 40 failing reconciliation) and 6 duplicates; the event-sourced generator, none.
- 0.5% with the queue ordered by score, or 1% with a third analyst.
- That no search can cover channels the firm does not capture, and what capture and attestation cover.
- Plant episodes, off-channel pointers and report faults on a schedule and check that each comes out of the pipeline where it should.
- Coverage, capacity and quality measured, not assumed.
30.10 Interview questions
Interview question 30.1 ★ developer
What is the difference between trade reporting and transaction reporting?
Solution
Solution of Interview question 30.1.
Trade reporting publishes trades to the market after execution (prices and sizes, for transparency); transaction reporting sends the regulator a confidential record of every transaction with the parties’ identities, for surveillance.
What the interviewer is looking for: Public transparency against confidential surveillance.
Interview question 30.2 ★★ developer
Your surveillance team is three months behind on alerts. What do you change?
Solution
Solution of Interview question 30.2.
Measure alerts a day against capacity per detector; raise thresholds or improve detectors where false positives dominate; order the queue by risk; close low-risk alerts by documented rules or sampling; staff to the chosen threshold.
What the interviewer is looking for: Threshold as a capacity decision; risk-ordered queue.
Interview question 30.3 ★★ developer
How do you make sure every amended trade is correctly reported to the regulator?
Solution
Solution of Interview question 30.3.
Generate reports from the event-sourced trade store, so an amendment is a cancellation and a new report; validate before sending; reconcile the net live reports with the trades every day.
What the interviewer is looking for: Events and daily reconciliation.
Interview question 30.4 ★★ developer
How would you measure whether a spoofing detector works?
Solution
Solution of Interview question 30.4.
True-positive rate at a fixed false-positive rate on planted or labelled episodes, the share of each legitimate account type it flags, and precision at the real base rate; then end-to-end, whether its alerts are reviewed in time.
What the interviewer is looking for: Detection at a stated false-alarm rate, then in the queue.
Interview question 30.5 ★★★ developer
Design the surveillance and regulatory reporting platform of a trading firm.
Solution
Solution of Interview question 30.5.
Feature extraction from orders and trades; detectors with thresholds set against capacity; a case management system with risk-ordered queues and ageing; a communications archive capturing all approved channels with holds and lexicon or model search; transaction reports from the event-sourced trade store with validation and daily reconciliation; planted end-to-end tests and published measures.
What the interviewer is looking for: Pipeline, people, archive, reports, measures.