Research, Data and Risk Platforms · Technology
29Security
On 21 February 2025 a large crypto exchange signed a transaction from one of its wallets that required several signatures. The signers approved what their screens showed them: a legitimate transaction. The web application that prepared it had been altered two days earlier to display legitimate transaction data while changing what was actually signed, and the signatures gave the attackers the wallet’s ether. The FBI attributed the theft of about $1.5 billion to North Korea. No cryptography was broken; the attackers changed what people were shown. Security in a trading firm is mostly about who and what is allowed to do things, and about making sure that what they see is what they sign. This chapter writes the firm’s access policy as data and checks it, keeps secrets and venue keys under rules, segments the network, and looks for the insider who copies a little too much every day.
29.1 What a trading firm protects, and from whom
Definition 29.1 (Threat model)
A threat model lists what a system must protect (its assets), from whom (the adversaries and their means), through which paths (the attack surface), and which controls address each path, so that security effort goes where the losses would be.
A trading firm’s assets are unusual in two ways. Some are money that moves by itself on a signature — exchange balances, wallets (Book 3, chapter 24), payment instructions — so a stolen credential is a loss within minutes. Others are information whose value is its secrecy — strategy code, signals, parameters, positions — so a copy is a loss even when nothing is deleted. The adversaries follow: criminals and states for the first, competitors and departing employees for the second. The operational risk of Book 6, chapter 28, covers both; this chapter builds the controls.
29.2 Identity, access and least privilege
Definition 29.2 (Least privilege, role-based access control)
Least privilege gives each person and program only the permissions its task needs, for only as long as it needs them. Role-based access control grants permissions — an action on a resource — to roles and roles to people, so that access follows the job rather than the individual.
Definition 29.3 (Multi-factor authentication)
Multi-factor authentication requires more than one distinct kind of authentication factor — something one knows, has or is — before access; NIST’s current guidance requires, at its middle assurance level, two distinct factors and the offer of a phishing-resistant one.
The chapter’s firm has twelve roles and a hundred people. Its policy is data (firm.accessctl): roles grant (resource, action) pairs, people hold roles, deny rules override grants, and five segregation-of-duties constraints (Book 6, chapter 28) name pairs of permissions no single person may hold.
P = lambda r, a: (r, a) # noqa: E731 - a permission
SOD = [(P("orders", "submit"), P("trades", "amend")), # no correcting one's own trades
(P("params", "propose"), P("params", "approve")), # chapter 16's four eyes
(P("code", "merge"), P("prod", "deploy")), # chapter 27: nobody ships alone
(P("venue_keys", "create"), P("wallet", "withdraw")), # key makers move no money
(P("payments", "create"), P("payments", "release"))]
@dataclass
class Policy:
roles: dict
users: dict
deny: set = field(default_factory=set)
sod: list = field(default_factory=list)
def permissions(self, user: str) -> set:
perms = set().union(*(self.roles[r] for r in self.users.get(user, ())))
return {p for p in perms if (user, *p) not in self.deny}
def allowed(self, user: str, resource: str, action: str) -> bool:
return (resource, action) in self.permissions(user)
def toxic_roles(self) -> list[tuple]:
"""Roles that alone hold both halves of a segregation-of-duties constraint."""
return sorted((r, a, b) for r, perms in self.roles.items() for a, b in self.sod
if a in perms and b in perms)
def toxic_users(self) -> list[tuple]:
"""Users who hold both halves, through one role or by combining several."""
return sorted((u, a, b) for u in self.users for a, b in self.sod
if a in self.permissions(u) and b in self.permissions(u))
| role | first permission | second permission |
|---|---|---|
| developer | merge code | deploy to production |
| head of desk | propose parameters | approve parameters |
| operations | create payments | release payments |
| site reliability | create venue keys | withdraw from wallets |
| trader | submit orders | amend trades |
Each toxic role was created for convenience and each is how a known kind of loss happens: a trader who amends his own bookings can hide losses (the rogue-trading cases of Book 6); a developer who deploys his own merge skips chapter 27’s review; a head of desk who approves his own parameters skips chapter 16’s four eyes. The check is a few lines of code over the policy data and belongs in the policy’s own continuous integration: a change that creates a toxic combination fails like a failing test.
29.3 Secrets and keys
Definition 29.4 (Secrets management, secret rotation)
Secrets management keeps credentials — passwords, API keys, private keys, database passwords — in a vault that issues them to authenticated programs under leases and logs every access, never in code or configuration files. Secret rotation replaces a secret on a schedule and on any suspicion, so that a leaked secret is valid only until its next rotation: its cryptoperiod, in NIST’s terms.
Definition 29.5 (Hardware security module)
A hardware security module (HSM) is a tamper-resistant device that generates and stores keys and performs operations with them, so that private keys never leave it in plain form; signing happens inside, on data the host presents.
Rotation bounds exposure: a secret that leaks at a random time and is rotated every 90 days stays valid 45 days on average; rotated weekly, three and a half. An HSM, or the multi-party computation of Book 3, chapter 24, removes the key from the host — but not the question the hook asks: the device signs whatever it is shown. The defence is to make what is signed verifiable independently of the screen that prepared it: a second channel that shows the transaction’s actual content, a policy engine that refuses a transaction type (a contract upgrade, a new destination) whatever the signers approve.
29.4 Exchange and wallet credentials
Definition 29.6 (Withdrawal allow-list)
A withdrawal allow-list restricts the destinations to which an account or key may send assets to a list fixed in advance, changed only through a separate, slower procedure, so that a stolen credential cannot move assets to the thief.
Venue API keys (Book 3, chapter 15) are the firm’s most exposed secrets: they sit on the trading servers and sign every order. The chapter’s keys carry scopes — a trading key can read and trade but not withdraw; a withdrawal key can only withdraw, to allow-listed addresses — and every request is signed with Book 13’s HMAC signer and stamped with its time. The venue’s checks decide what a stolen key can do (Table 29.2).
def verify(keys: dict, req: dict, at: float, max_age: float = 5.0) -> tuple[bool, str]:
"""What the venue checks: a known key, a valid signature, a fresh request, a scope that
covers the action, and for withdrawals an address on the allow-list."""
key = keys.get(req["key"])
if key is None:
return False, "unknown key"
if Signer(key.secret).sign(req["body"]) != req["sig"]:
return False, "bad signature"
body = json.loads(req["body"])
if not 0 <= at - body["at"] <= max_age:
return False, "stale request"
if body["action"] not in key.scopes:
return False, "action outside the key's scope"
if body["action"] == "withdraw" and body.get("address") not in key.allow_list:
return False, "address not on the allow-list"
return True, "ok"
| attempt | accepted | reason |
|---|---|---|
| trade with the trading key | yes | ok |
| withdraw with the trading key | no | action outside the key’s scope |
| withdraw to an outside address | no | address not on the allow-list |
| withdraw to the custodian | yes | ok |
| forged signature | no | bad signature |
| replayed a minute later | no | stale request |
29.5 Segmentation and insider risk
Definition 29.7 (Network segmentation, zero-trust architecture)
Network segmentation divides a firm’s network into zones — trading, research, corporate, the outside — with traffic between zones allowed only on stated paths, so that a compromise in one zone does not reach the others. A zero-trust architecture grants no implicit trust for being inside a network: every request is authenticated and authorised on its own, by identity and device, before a session to a resource is established.
Definition 29.8 (Insider threat, data-loss prevention)
An insider threat is harm done by someone with legitimate access — an employee, a contractor — through theft, sabotage or negligence. Data-loss prevention is the set of controls that detect and block the movement of protected data out of the places it is allowed to be: copies to removable media, uploads, mail, and unusual volumes of reading.
The best-known case in trading is also a warning about the law’s limits: in 2009 a programmer uploaded more than 500 000 lines of his employer’s high-frequency trading code to a server in Germany on his last day; his federal conviction was reversed on appeal in 2012, on the reach of the statutes rather than the facts. A firm cannot rely on the courts (Book 16, chapter 11, on trade secrets); it must see the copying. The difficult insider is not the one who takes everything on the last day but the one who takes a little more than usual every day.
The chapter’s model gives each of the hundred people a year of 250 trading days of access logs — the strategy-code files each reads a day, at each person’s own level, with a weekly rhythm and noise — and plants an insider who, from a random day, reads a stated share more than usual, every day. Four detectors score each person against their own baseline, the median and median absolute deviation of 30 days: a single-day threshold, or a CUSUM of the daily scores (Book 7, chapter 13), each with a baseline that trails today or one that ends 60 days earlier. All four are calibrated on clean people to one false alarm a month per hundred people.
def robust_scores(x: np.ndarray, window: int = 30, gap: int = 0) -> np.ndarray:
"""Each user-day against the user's own baseline: (x - median) / (1.4826 MAD) over the
`window` days ending `gap` days before; a gap keeps a slow change out of its baseline.
Days without a full baseline score 0."""
n_users, n_days = x.shape
z = np.zeros_like(x, dtype=float)
for d in range(window + gap, n_days):
past = x[:, d - gap - window:d - gap]
med = np.median(past, axis=1)
mad = 1.4826 * np.median(np.abs(past - med[:, None]), axis=1)
z[:, d] = (x[:, d] - med) / np.maximum(mad, 1e-9)
return z
def cusum(z: np.ndarray, k: float = 0.5) -> np.ndarray:
"""One-sided CUSUM per user: S_d = max(0, S_{d-1} + z_d - k)."""
s = np.zeros_like(z, dtype=float)
for d in range(1, z.shape[1]):
s[:, d] = np.maximum(0.0, s[:, d - 1] + z[:, d] - k)
return s
fig_accessctl.py.Two things matter, and Figure 29.3 separates them. A single day of reading 30% more than usual is well inside a person’s noise; only accumulation — the CUSUM — turns thirty such days into a signal. And a baseline that trails today learns the new habit: after thirty days the insider’s copying is the person’s normal, and the score returns to zero. Lagging the baseline by sixty days keeps the copying out of its own reference long enough to be seen. With both, an insider reading 30% more than usual is flagged in a median of 22 days (19 of 20), one reading 50% more in 12 (all 20); a single-day threshold on a trailing baseline flags one of twenty.
fig_accessctl.py.As of September 2026 — The cases and the standards
The FBI’s public service announcement of 26 February 2025 attributed to North Korea the theft of about $1.5 billion in virtual assets from the exchange Bybit on or about 21 February; the investigation published by Sygnia describes JavaScript altered on 19 February in the Safe{Wallet} web application so that its interface displayed legitimate transaction data while the transaction itself was altered. In United States v. Aleynikov (2d Cir. 2012) the court reversed a conviction for uploading more than 500 000 lines of trading source code. NIST SP 800-207 (2020) defines zero-trust architecture; SP 800-63B (revision of August 2025) the authentication assurance levels; SP 800-57 Part 1 (2020) the cryptoperiod of a key.
29.6 Tutorial: the slow copy
Goal. Check the access policy for toxic combinations, issue and test scoped venue keys, and find a slow insider at a fixed false-alarm rate. End state: Table 29.1 and Figure 29.3.
- Policy:
pl_accessctl.policies()andtoxic_summary(), before and after the fix. - Keys:
stolen_key_attempts(); add a key with a trade scope and an allow-list and try it. - Vault:
firm_accessctl.Vaultwith a 90-day lease; rotate and list what expires. - Logs:
access_logs()andstatistics(x);thresholds(). - Insider:
delays(extra)for each share of extra reading.
What to change next. Score people against their peers in the same role as well as against themselves; add file-level features (files never read before) and see how much sooner the insider is flagged.
29.7 Build: access control
Purpose. Grant each person and program only what its job needs, keep secrets and venue keys under rules that bound a leak, and see a slow copy at a stated false-alarm rate.
Interface. Policy (allowed, permissions, toxic_roles, toxic_users), Vault, VenueKey, request, verify, robust_scores, cusum, onsets, calibrate, first_alarm.
Rules. Policy as reviewed data, checked for toxic combinations in continuous integration; secrets only from the vault, under leases, rotated; venue keys scoped, withdrawal only to allow-listed addresses; every request signed and time-stamped; detectors calibrated on clean populations and reviewed monthly.
Acceptance tests. code/firm/accessctl/tests/: grants, denies, toxic roles and users; vault leases, rotation and expiry; each of the venue’s checks; lagged robust scores, CUSUM, calibration to a false-alarm rate, and a persistent shift found.
Stretch. Just-in-time access grants that expire; peer-group scoring; an independent transaction decoder for signers.
Sources and further reading
- Federal Bureau of Investigation, North Korea Responsible for $1.5 Billion Bybit Hack, PSA 250226, 26 February 2025; Sygnia, investigation of the Bybit hack, 2025.
- United States v. Aleynikov, 676 F.3d 71 (2d Cir. 2012).
- NIST SP 800-207, Zero Trust Architecture, 2020; SP 800-63B, Authentication and Authenticator Management; SP 800-57 Part 1 Rev. 5, Recommendation for Key Management, 2020.
- One Quant Book 3, chapters 14, 15 and 24 (keys, venue APIs, wallets); Book 6, chapter 28 (segregation of duties); Book 16, chapter 11 (trade secrets).
29.8 Exercises
Exercise 29.1 ★
Why does the check for toxic combinations look at people as well as roles?
Solution
Solution of Exercise 29.1.
Because harmless roles combine into toxic people. In the chapter’s policy three quantitative researchers who were also risk managers could propose parameters and approve them, a combination neither role held alone.
Exercise 29.2 ★
A secret leaks at a random moment. How long does it stay valid on average under 90-day and weekly rotation?
Solution
Solution of Exercise 29.2.
Half the rotation period on average: 45 days under 90-day rotation, three and a half days under weekly rotation.
Exercise 29.3 ★
What can a thief do with the firm’s stolen trading key, and what can he not?
Solution
Solution of Exercise 29.3.
Read and trade — which can still lose money, for example by trading against the firm’s own positions — but not withdraw: the key’s scope excludes it, and even a withdrawal key can send only to allow-listed addresses. Forged or replayed requests are refused.
Exercise 29.4 ★★
Why does a baseline that trails today stop seeing a slow copy after about a month?
Solution
Solution of Exercise 29.4.
The baseline is the median of the last thirty days; after thirty days of copying, the copying days are the baseline, the person’s usual level has risen to include them, and the daily scores fall back to zero.
Exercise 29.5 ★★
Why is a single-day threshold the wrong tool for a slow copy, whatever its level?
Solution
Solution of Exercise 29.5.
A day of 30% more reading is about one standard deviation of a person’s daily noise; a threshold high enough to allow only one false alarm a month per hundred people sits above four. Any single day of the copy is ordinary; only the sum of many is not.
Exercise 29.6 ★★
The signers of the hook each checked the transaction on their screen. What control would have caught the attack?
Solution
Solution of Exercise 29.6.
A check of what was actually to be signed, independent of the screen that prepared it: the transaction decoded on the signing device or a second system, or a policy engine that refuses a transaction type (a change to the wallet’s own contract, a new destination) whatever the signatures.
Exercise 29.7 ★★★
Coding. Rerun delays(0.3) with the CUSUM’s reference value K at 0.25 instead of 0.5 (recalibrating the thresholds). What happens to the detections and why?
Solution
Solution of Exercise 29.7.
With the thresholds recalibrated, the lagged CUSUM flags all 20 insiders at +30% a day, in a median of 20.5 days instead of 22 (19 of 20). A smaller reference value lets smaller daily excesses accumulate: it is tuned to a smaller shift, at the price of a higher threshold for the same false-alarm rate.
Exercise 29.8 ★★★
Find the flaw. "Our trading network is behind the firewall, so the systems inside it can trust each other."
Solution
Solution of Exercise 29.8.
The firewall is one control, not a reason for trust: a compromised laptop, a stolen credential or a malicious insider is already inside. Every request should be authenticated and authorised on its own, by identity and device, with paths between zones stated and minimal.
29.9 Problem: The Slow Copy
Problem 29.1
Weekend problem — the slow copy
The chapter’s policy, keys, vault and access logs.
Part I — Access.
- What does the firm protect, and from whom?
- What are the five segregation-of-duties constraints?
- Which roles were toxic before the fix, and how many people held a toxic combination?
- What did the fix change?
- Where should the toxic-combination check run?
Part II — Secrets and keys.
- What does a vault add to a secret in a configuration file?
- How does rotation bound a leak?
- What are the venue’s checks on a request?
- What does each stolen key allow?
- What does an HSM protect, and what does it not?
Part III — Insiders.
- How are the four detectors built and calibrated?
- What does each flag at +30% a day?
- Why does the CUSUM matter?
- Why does the lag of the baseline matter?
- When does the example insider’s lagged CUSUM cross its threshold?
Part IV — The verdict.
- State the named result: the days until a planted insider’s slow exfiltration is flagged at one false alarm a month per hundred users, and the number of toxic role combinations in the firm’s policy before and after the fix.
- What would you add to the detectors first?
- How should the trading zone be segmented?
- What is the lesson of the hook for any firm that signs transactions?
- In one sentence: what is security in a trading firm about?
Solution
Solution of Problem 29.1.
- Money that moves on a signature, from criminals and states; information whose value is its secrecy, from competitors and departing staff.
- No submitting orders and amending trades; no proposing and approving parameters; no merging code and deploying it; no creating venue keys and withdrawing; no creating and releasing payments.
- Developer, head of desk, operations, site reliability and trader; 60 of the 100 people.
- Split the five roles, named a release manager, removed the three second roles: no toxic role and no toxic person.
- In the policy’s own continuous integration, failing any change that creates a combination.
- Leases, access logs, rotation and revocation, and no copy in code or configuration.
- A leaked secret is valid only until the next rotation: half the period on average.
- A known key, a valid signature, a fresh timestamp, an action within the key’s scope, and an allow-listed withdrawal address.
- The trading key: reads and trades; the withdrawal key: withdrawals to the custodian only.
- It keeps private keys from ever leaving it; it does not check what it is asked to sign.
- Robust scores against each person’s own 30-day baseline, trailing or lagged by 60 days, used alone or accumulated in a CUSUM; each threshold set on clean people for one false alarm a month per hundred.
- Single day, trailing: 1 of 20; CUSUM, trailing: 5; single day, lagged: 8; CUSUM, lagged: 19, in a median of 22 days.
- It adds small daily excesses that no single day shows.
- A trailing baseline absorbs the copying within a month; a lagged one keeps it out of its own reference.
- On day 131, eleven days after the copying began.
- Named result. At one false alarm a month per hundred people, an insider reading 30% more strategy code than usual every day is flagged in a median of 22 days (19 of 20) by a CUSUM on a lagged baseline, and 50% more in 12 days (20 of 20); a single-day threshold on a trailing baseline flags one of twenty. The policy held 5 toxic roles and 60 toxic people before the fix, none after.
- Peer comparison within the role and features of novelty (files never read before).
- Trading, research, corporate and outside zones, with orders only to venues, fills to research, parameters back only through review, and no path from laptops or the internet to trading.
- Verify what is signed independently of what prepares it, and let policy refuse what no signer should approve.
- Who and what may do what, and whether what they see is what they do.
29.10 Interview questions
Interview question 29.1 ★ developer
Where should an exchange API secret live, and how should a strategy get it?
Solution
Solution of Interview question 29.1.
In a secrets vault, issued to the strategy’s service identity under a lease at start-up, rotated on a schedule, never in code, images or configuration, and every access logged.
What the interviewer is looking for: Vault, leases, rotation.
Interview question 29.2 ★★ developer
What permissions should an exchange API key used for trading have?
Solution
Solution of Interview question 29.2.
Read and trade only, never withdraw; restricted to the firm’s IP addresses; a separate key per strategy or server; withdrawals, if needed at all, on a separate key limited to allow-listed addresses.
What the interviewer is looking for: Scopes and allow-lists.
Interview question 29.3 ★★ developer
How would you detect an employee who copies strategy code slowly over months?
Solution
Solution of Interview question 29.3.
Log reads of protected code and data per person; score each day against the person’s own lagged baseline and their peers; accumulate the scores with a CUSUM; calibrate the threshold on clean data to a false-alarm rate the team can investigate; add data-loss controls on copies and uploads.
What the interviewer is looking for: Accumulation and a baseline that does not learn the theft.
Interview question 29.4 ★★ developer
What is zero trust, and what does it change inside a trading network?
Solution
Solution of Interview question 29.4.
No implicit trust from being inside the network: every request authenticated and authorised by identity and device. Inside a trading network it means service identities, mutual authentication between services, and paths between zones stated and minimal.
What the interviewer is looking for: Identity per request.
Interview question 29.5 ★★★ developer
Design access control, secrets and monitoring for a trading firm of a hundred people with venue keys and a crypto treasury.
Solution
Solution of Interview question 29.5.
Roles as reviewed data with segregation-of-duties checks in CI; multi-factor, phishing-resistant authentication; a vault with leases and rotation; scoped venue keys and allow-listed withdrawals on separate keys; treasury signing with independent transaction verification and policy limits; network zones; access logging with calibrated insider detection; a threat model reviewed yearly.
What the interviewer is looking for: Least privilege, bounded secrets, verified signing, detection.