---
title: "Clustering, Regimes and Anomaly Detection"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 20
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/20-clustering-regimes-and-anomaly-detection
---

# Chapter 20 — Clustering, Regimes and Anomaly Detection

A risk dashboard shows the market in its calm regime on the Friday before a crash. On Tuesday, refitted, the same model shows the crash beginning on Friday. Both were right about what they could know: the first was a filtered probability, computed from the data up to Friday; the second a smoothed one, computed with Monday in hand. On the chapter’s planted regimes the filtered probability of stress on the day before a switch averages 0.03, no higher than on any calm day; the smoothed one averages 0.35. This chapter is about learning without labels: grouping assets and days ([clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster)), describing the market by hidden states (regime models), and finding what does not fit ([anomaly detection](#def-ml-clustering-regimes-and-anomaly-detection-anomaly)). Each is useful, and each invites a particular mistake: clusters that are noise, regimes known only in hindsight, anomalies that are merely rare.

## 20.1 Clustering assets and days

**Definition 20.1 (Clustering, k-means, hierarchical clustering).**

*Clustering* partitions objects into groups whose members are more similar to each other than to others, without labels. *k-means* chooses $k$ centres and assigns each object to the nearest, alternating the two steps to minimise the within-group sum of squared distances (Lloyd’s algorithm). *Hierarchical clustering* merges the closest groups step by step into a tree and cuts it at the desired number of groups; for assets the usual distance is $\sqrt{2(1-\rho_{ij})}$ between return series (Mantegna, 1999), the one behind hierarchical risk parity (Book 7, chapter 26).

The chapter clusters the names of the firm’s synthetic market (Book 7’s `firm.synthmkt`) that stay listed for five years, whose returns load on a market factor, ten planted industry factors, style factors and fat-tailed specific noise. The market factor is removed first (each name’s return minus its beta times the equal-weighted average), since it would dominate every correlation. [Figure 20.1](#fig-ml-reg-clusters) scores [k-means](#def-ml-clustering-regimes-and-anomaly-detection-cluster) and average-linkage [hierarchical clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) by the adjusted Rand index against the planted industries (1 is perfect, 0 is chance), for windows of three months to five years and for two strengths of the industry factor: the default 8% a year, and 15%.

![Agreement of return-based clusters with the planted industries, by the length of the return history and the strength of the industry factor. Data: ml_regimes.clustering.](https://one-course.com/images/onecourse/chapters/quant-12/ml-clustering-regimes-and-anomaly-detection/fig-7617e0125796.svg)

***Figure 20.1.** Agreement of return-based clusters with the planted industries, by the length of the return history and the strength of the industry factor. Data: `ml_regimes.clustering`.*

With strong industries, [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) recovers them from five years of data (index 0.99 and 1.00) and half of them from one year; with the default strength, which is closer to what specific risk leaves of real industry effects, it finds almost nothing in a year (0.06) and 0.15 to 0.21 in five. The half-sample stability score measures the same thing without labels: the index between [clusterings](#def-ml-clustering-regimes-and-anomaly-detection-cluster) of the two halves of a year is 0.32 to 0.35 with strong industries and 0.06 to 0.08 with weak ones. A [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) that does not survive a split of its own sample is describing noise, and a desk that uses it to group risk or build hierarchical risk parity is rebalancing on noise.

## 20.2 Regime models

**Definition 20.2 (Hidden Markov model, regime-switching model, Viterbi algorithm).**

A *hidden Markov model* (HMM) has an unobserved state that follows a Markov chain (Book 4, chapter 8) and observations whose distribution depends on the current state; it is fitted by expectation-maximisation (Book 4, chapter 19), the Baum–Welch algorithm (Rabiner, 1989). A *regime-switching model* is an HMM for economic series, each state a regime with its own mean, volatility or dynamics (Hamilton, 1989; Ang and Timmermann, 2012). The *Viterbi algorithm* finds the single most likely path of states given all the observations, by dynamic programming.

The chapter’s index switches between a calm regime (daily mean 5 basis points, volatility 0.8%) and a stress regime (mean $-10$ basis points, volatility 2.5%), staying in each with probability 0.99 and 0.95 a day; twenty years are simulated, and stress takes a fifth of the time. A two-state Gaussian HMM fitted on the first ten years finds calm with volatility 0.80% and stress with 2.60%, staying probabilities 0.988 and 0.958. A three-state model adds 5.7 to the log-likelihood with six more parameters, less than the Bayesian information criterion’s penalty of 23.5, and spends its third state splitting stress in two (2.28% and 2.65%).

## 20.3 Filtering against smoothing: what a regime model knows in real time

An HMM gives three answers about the state at date $t$. The *filtered* probability uses the data up to $t$ (the forward recursion of [Listing 20.2](#lst-ml-reg-hmm)); the *smoothed* probability uses all the data, including what came after $t$ (the backward pass); the Viterbi path is the best sequence given all the data. Only the first existed at $t$. On the ten test years, the filtered estimate (stress when its probability exceeds one half) is right on 95.9% of days, the smoothed on 97.7% and the Viterbi path on 97.5%; a rule on the trailing 20-day volatility is right on 89.8%. After each of the 20 switches into stress, the filter says stress after a median of one day (mean 1.5, one short spell missed), the smoothed estimate on the day itself (mean 0.8). [Figure 20.2](#fig-ml-reg-switch) shows one switch: the smoothed probability rises before the regime changes, because the days after the switch make the days before it look like stress.

![Filtered and smoothed probabilities of stress around one switch in the test years (the first switch the filter took at least two days to see). Data: ml_regimes (fig_regimes.py).](https://one-course.com/images/onecourse/chapters/quant-12/ml-clustering-regimes-and-anomaly-detection/fig-e5af73ac4607.svg)

***Figure 20.2.** Filtered and smoothed probabilities of stress around one switch in the test years (the first switch the filter took at least two days to see). Data: `ml_regimes` (`fig_regimes.py`).*

The difference is a look-ahead bias (Book 7, chapter 3) that no timestamp reveals. On the day before a switch, the filtered probability of stress averages 0.032, below its average of 0.041 over all calm days: the filter cannot see a switch coming. The smoothed probability averages 0.353 on those days against 0.019 on calm days. A backtest that scales exposure to a target volatility using the regime probabilities earns a Sharpe ratio of 1.00 with yesterday’s filtered probability carried forward by the transition matrix, and 1.21 with the same day’s smoothed probability; buying and holding earns 0.91. A regime model used for trading or risk must be run as a filter, refitted only on data available at each date, and its statistics (the volatility of a regime, its average duration) computed from filtered probabilities.

## 20.4 Anomaly detection

**Definition 20.3 (Anomaly detection, isolation forest, precision–recall curve).**

*Anomaly detection* scores observations by how unlike the bulk of the data they are, without labels for the anomalies. An *isolation forest* builds random trees that split on random features at random values; points that are isolated after few splits are anomalous (Liu, Ting and Zhou, 2008). The *precision–recall curve* plots, as the alarm threshold moves, the share of alarms that are true (precision) against the share of true cases that raise an alarm (recall); when positives are rare it is more informative than the ROC curve (Davis and Goadrich, 2006).

Book 9 (chapter 29) built surveillance detectors for spoofing on account-days: 100 planted episodes among 11 500 legitimate account-days, including market makers who cancel almost everything and deep-book liquidity providers whose large orders rarely fill. [Table 20.1](#tab-ml-reg-anomaly) adds two unsupervised scores on five features of the same account-days (the log counts of small and large orders, their fill rates, and the share of large cancellations that follow a fill on the other side): an [isolation forest](#def-ml-clustering-regimes-and-anomaly-detection-anomaly) ([Listing 20.3](#lst-ml-reg-iforest)) and the reconstruction error of a small [autoencoder](https://one-course.com/books/quant/12/en/chapter/9-cross-sectional-deep-models#def-ml-cross-sectional-deep-models-ae) (chapter 9).

| detector | recall at 1% false positives | precision |
| --- | --- | --- |
| order-to-trade ratio (rule) | 0.00 | 0.00 |
| fill-rate gap (rule) | 0.37 | 0.24 |
| cancellations after a fill (rule) | 0.51 | 0.31 |
| gap $\times$ cancellations (rule) | 0.77 | 0.40 |
| [isolation forest](#def-ml-clustering-regimes-and-anomaly-detection-anomaly) | 1.00 | 0.47 |
| [autoencoder](https://one-course.com/books/quant/12/en/chapter/9-cross-sectional-deep-models#def-ml-cross-sectional-deep-models-ae) | 0.32 | 0.22 |

***Table 20.1.** Spoofing detectors on 11 600 account-days (100 planted episodes), at the threshold that flags 1% of legitimate account-days. Data: `ml_regimes.anomalies`.*

## 20.5 Surveillance detectors

The [isolation forest](#def-ml-clustering-regimes-and-anomaly-detection-anomaly) finds every planted episode at the 1% false-positive rate, and its precision, 0.47, is the most any detector can have there: 115 legitimate account-days are flagged whatever the detector, against at most 100 true ones. The [autoencoder](https://one-course.com/books/quant/12/en/chapter/9-cross-sectional-deep-models#def-ml-cross-sectional-deep-models-ae) does worse than the rules. Three cautions keep this result in its place. The forest was given the features the rules’ authors chose, which encode what spoofing looks like; on raw order logs it would find the unusual, not the manipulative. The planted spoofers are a small, distinct population; when a manipulative pattern is common, or when legitimate traders who look like it are many, an unsupervised score ranks them together. And precision at a fixed false-positive rate is what a compliance team lives with: a detector with perfect recall still hands them more innocent account-days than guilty ones. Unsupervised scores earn their place as a second net under the rules, reviewed by people, and as a way to find patterns no rule names yet.

**Method 20.4 (Unsupervised learning on market data).**

1. Remove the common factor before [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) returns, and report a stability score across sub-samples with every [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) .
2. Fit regime models on data available at each date and use filtered probabilities only; report smoothed results as descriptions of history, never as backtests.
3. Choose the number of regimes by an out-of-sample likelihood or an information criterion, and check that the regimes differ in something the desk acts on.
4. Evaluate anomaly scores like detectors: recall and precision at the false-positive rate reviewers can handle, against the rules, on data with planted and look-alike cases.

## 20.6 Tutorial: which regime are we in?

**Goal.** Cluster the synthetic names, fit and run a regime model as a filter and as a smoother, and score unsupervised detectors against the surveillance rules. **End state:** Figures [20.1](#fig-ml-reg-clusters) and [20.2](#fig-ml-reg-switch), [Table 20.1](#tab-ml-reg-anomaly).

1. **[Hierarchical clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) on the correlation distance.** `def hier_clusters (R, k): """Average-linkage hierarchical clustering on the correlation distance (Mantegna's metric).""" from scipy.cluster.hierarchy import fcluster, linkage from scipy.spatial.distance import squareform D = corr_distance(R) np.fill_diagonal(D, 0.0 ) return fcluster(linkage(squareform(D, checks=False ), " average " ), k, " maxclust " ) - 1` **Listing 20.1.** Average linkage on Mantegna’s distance. code/firm/regimes/firm_regimes.py
2. **Filtering and smoothing.** `def _forward (self , x): B = self ._dens(x) a = np.zeros((len (x), self .k)) c = np.zeros(len (x)) a[0 ] = self .pi * B[0 ] c[0 ] = a[0 ].sum() a[0 ] /= c[0 ] for t in range (1 , len (x)): a[t] = (a[t - 1 ] @ self .P) * B[t] c[t] = a[t].sum() a[t] /= c[t] return a, c, B def filter (self , x): """P(s_t | x_1, ..., x_t): the regime as known at date t.""" return self ._forward(np.asarray(x, float ))[0 ] def smooth (self , x): """P(s_t | x_1, ..., x_T): the regime as known at the end of the sample.""" x = np.asarray(x, float ) a, c, B = self ._forward(x) b = np.ones((len (x), self .k)) for t in range (len (x) - 2 , -1 , -1 ): b[t] = (self .P @ (B[t + 1 ] * b[t + 1 ])) / c[t + 1 ] g = a * b return g / g.sum(1 , keepdims=True )` **Listing 20.2.** The forward recursion (filter) and the backward pass (smoother). code/firm/regimes/firm_regimes.py
3. **An [isolation forest](#def-ml-clustering-regimes-and-anomaly-detection-anomaly).** `def iforest_scores (X, seed=0 ): """Isolation forest (Liu, Ting and Zhou): anomalies are isolated by fewer random splits.""" from sklearn.ensemble import IsolationForest f = IsolationForest(n_estimators=200 , random_state=seed, n_jobs=1 ).fit(X) return -f.score_samples(X)` **Listing 20.3.** Isolation-forest anomaly scores. code/firm/regimes/firm_regimes.py
4. **Run** `ml_regimes.clustering()` , `regimes()` , `vol_targeting()` , `anomalies()` and `fig_regimes.py` .

**What to change next.** Refit the HMM every year on an expanding window and filter the next year; double the number of legitimate deep-book providers and rerun the anomaly scores.

## 20.7 Build: clustering, regimes and anomaly scores

**Purpose.** Unsupervised tools whose outputs are known in real time and scored like detectors.

**Interface.** `market_residuals`, `corr_distance`, `kmeans_clusters`, `hier_clusters`, `ari`, `stability`; `regime_series`; `GaussianHMM(k)` with `fit`, `filter`, `smooth`, `viterbi`; `iforest_scores`, `autoencoder_scores`; with Book 9’s `firm.surveil.tpr_at_fpr`.

**Rules.** Filtered probabilities only in anything that trades or reports risk; every [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) with its stability score; anomaly scores with recall and precision at a fixed false-positive rate.

**Acceptance tests.** `code/firm/regimes/tests/`: [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) recovers well-separated planted groups; the HMM recovers planted parameters and its filter uses no future data (it is unchanged when later data change); smoothing and filtering agree on the last day; Viterbi matches brute force on a short series; the [isolation forest](#def-ml-clustering-regimes-and-anomaly-detection-anomaly) ranks planted outliers first.

**Stretch.** Regimes with their own dynamics (autoregressive or GARCH within states); online refitting; [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) days rather than assets.

Sources and further reading

- J. D. Hamilton, “A new approach to the economic analysis of nonstationary time series and the business cycle”, *Econometrica* 57(2), 1989.
- A. Ang and A. Timmermann, “Regime changes and financial markets”, *Annual Review of Financial Economics* 4, 2012.
- L. R. Rabiner, “A tutorial on hidden Markov models and selected applications in speech recognition”, *Proceedings of the IEEE* 77(2), 1989.
- F. T. Liu, K. M. Ting and Z.-H. Zhou, “Isolation forest”, *IEEE International Conference on Data Mining* , 2008.
- S. Lloyd, “Least squares quantization in PCM”, *IEEE Transactions on Information Theory* 28(2), 1982.
- R. N. Mantegna, “Hierarchical structure in financial markets”, *European Physical Journal B* 11, 1999.
- J. Davis and M. Goadrich, “The relationship between precision-recall and ROC curves”, *ICML* , 2006.

## 20.8 Exercises

**Exercise 20.1 ★.**

Two names have return correlation 0.5. What is their Mantegna distance? What distance do perfectly correlated and uncorrelated names have?

**Solution of Exercise 20.1.**

$\sqrt{2(1 - 0.5)} = 1$. Perfectly correlated names are at distance 0, uncorrelated ones at $\sqrt2 = 1.41$ (and perfectly anti-correlated at 2).

**Exercise 20.2 ★.**

With staying probabilities 0.988 and 0.958, what are the expected durations of the calm and stress regimes, and the long-run share of time in stress?

**Solution of Exercise 20.2.**

Durations $1/(1 - p)$: 82 days calm and 24 days stress (the fitted matrix, 0.988 and 0.958 rounded, gives 82.2 and 23.9). The long-run share of stress is $P_{12}/(P_{12} + P_{21}) = 0.225$, against 0.215 on the test years.

**Exercise 20.3 ★.**

At a 1% false-positive rate on 11 500 legitimate account-days, what is the highest precision any detector can reach with 100 true cases?

**Solution of Exercise 20.3.**

115 legitimate account-days are flagged at 1%; with all 100 true cases flagged too, precision is $100/215 = 0.47$.

**Exercise 20.4 ★★.**

Why is the filtered probability of stress on the day before a switch no higher than on an average calm day?

**Solution of Exercise 20.4.**

The filter uses only data up to that day, and the day before a switch is a calm day: its return is drawn from the calm distribution. A switch is announced by the stress days that follow it, which the filter has not seen; its probability on that day is just the calm-day probability of a switch, about 1%, pushed around by the day’s return.

**Exercise 20.5 ★★.**

Why must the market factor be removed before [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) returns?

**Solution of Exercise 20.5.**

Every stock loads on the market, so raw correlations are dominated by one common factor: all names look alike, and the [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) picks up differences in beta or specific volatility rather than industries. Removing the market leaves the industry correlations to drive the distances.

**Exercise 20.6 ★★.**

*Find the flaw.* “Our two-regime model’s stress probability, fitted on 2000–2024, turned up weeks before each of the three big drawdowns, so it is an early-warning signal.”

**Solution of Exercise 20.6.**

The probabilities are smoothed, or the model was fitted on the whole sample, drawdowns included: its states and parameters know where the drawdowns were, and the smoothed probability rises before each switch by construction (0.35 against 0.02 on calm days here). Rerun the model as a filter, refitted only on data before each date, and look again.

**Exercise 20.7 ★★★.**

*Coding.* Compare the three-state model with the two-state one on the test years by the log-likelihood of the test data under each fitted model (filtering forward from the fitted parameters). Which wins out of sample?

**Solution of Exercise 20.7.**

The test years’ log-likelihood is 7 866.6 under the two-state model and 7 864.4 under the three-state one: the extra state, which improved the fit in sample by 5.7, loses 2.2 out of sample. The information criterion’s verdict holds.

**Exercise 20.8 ★★★.**

Derive the forward recursion $\alpha_t(j) = \big(\sum_i\alpha_{t-1}(i)P_{ij}\big)b_j(x_t)$ for the joint probability of the observations up to $t$ and the state at $t$, and show that normalising it gives the filtered probability.

**Solution of Exercise 20.8.**

$\alpha_t(j) = p(x_1,\dots,x_t, s_t = j) = \sum_ip(x_1,\dots,x_{t-1}, s_{t-1} = i)\P(s_t = j\mid s_{t-1} = i)p(x_t\mid s_t = j)$, using the Markov property of the state and the conditional independence of $x_t$ given $s_t$; that is $\big(\sum_i\alpha_{t-1}(i)P_{ij}\big)b_j(x_t)$. Dividing by $\sum_j\alpha_t(j) = p(x_1,\dots,x_t)$ gives $\P(s_t = j\mid
x_1,\dots,x_t)$, the filtered probability; the code normalises at every step, which also keeps the numbers from underflowing.

## 20.9 Problem: Which Regime Are We In?

**Problem 20.1.**

Weekend problem — filtered, smoothed, and rare

The chapter’s clusters, regimes and detectors.

**Part I — Clusters.**

1. How well do clusters recover industries, by window and strength?
2. What does the stability score say, and why is it the one to report on real data?
3. What would you use clusters for on a desk, and what not?
4. Which [clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) method would you choose, and why?

**Part II — Regimes.**

5. What does the fitted two-state model find, against the planted regimes?
6. Why not three states?
7. What are the accuracies of the four regime estimates on the test years?
8. What are the delays after switches into stress?

**Part III — Hindsight.**

9. What do the filtered and smoothed probabilities say on the day before a switch?
10. How much does the smoothed probability flatter a volatility-targeting backtest?
11. How should regime statistics be computed?
12. Where else does smoothing hide in a research pipeline?

**Part IV — The verdict.**

13. State the *named result* : the filtered regime’s accuracy and detection delay against the smoothed one, and the anomaly detectors’ precision at a fixed false-positive rate against the rule-based detectors.
14. What is a regime model for, if not prediction?
15. Would you let the [isolation forest](#def-ml-clustering-regimes-and-anomaly-detection-anomaly) replace the rules?
16. How would you review the [isolation forest](#def-ml-clustering-regimes-and-anomaly-detection-anomaly) ’s alarms?
17. What would you monitor to know that a regime model has broken?
18. How do clusters feed hierarchical risk parity, and what happens when they are noise?
19. What labels would turn the anomaly problem into a supervised one, and at what cost?
20. In one sentence: what does a model know about today?

**Solution of Problem 20.1.**

**Part I.**

1. With industries at 15%: adjusted Rand index 0.38 to 0.42 on three months, 0.60 to 0.75 on a year, 0.99 to 1.00 on five years. At 8%: 0.03 to 0.06, 0.06, and 0.15 to 0.21.
2. Half-sample agreement over a year: 0.32 to 0.35 (strong industries), 0.06 to 0.08 (weak). On real data there are no planted labels; stability is the only check.
3. Grouping for risk reports and hierarchical risk parity when stable; not as a trading signal, and not when they do not survive a split.
4. [Hierarchical clustering](#def-ml-clustering-regimes-and-anomaly-detection-cluster) on the correlation distance is slightly better here and gives a tree that shows how groups nest; [k-means](#def-ml-clustering-regimes-and-anomaly-detection-cluster) needs a fixed number of groups and a Euclidean representation.

**Part II.**

1. Calm volatility 0.80% and stress 2.60% (planted 0.8% and 2.5%), staying probabilities 0.988 and 0.958 (0.99 and 0.95).
2. The third state adds 5.7 to the in-sample log-likelihood for six parameters, below the information criterion’s 23.5, and loses out of sample (exercise 7).
3. Filtered 95.9%, smoothed 97.7%, Viterbi 97.5%, volatility threshold 89.8%.
4. Median one day for the filter (mean 1.5), zero for the smoother (mean 0.8); the volatility threshold’s mean is 3.8 and it misses five of 20 spells.

**Part III.**

1. Filtered 0.032 (calm-day average 0.041); smoothed 0.353 (calm-day average 0.019).
2. Sharpe ratio 1.21 with the smoothed probability against 1.00 with the filtered one (buy and hold 0.91).
3. From filtered probabilities, with parameters fitted on data before each date.
4. In any statistic computed with a full-sample fit: normalisations, regime labels, detrending, outlier cleaning, and Kalman smoothers used where a filter was meant (Book 4, chapter 19).

**Part IV.**

1. *Which regime are we in?* The filtered regime is right on 95.9% of days and sees a switch into stress after a median of one day; the smoothed one is right on 97.7% and sees it on the day, because it knows the future. At a 1% false-positive rate the [isolation forest](#def-ml-clustering-regimes-and-anomaly-detection-anomaly) ’s precision is 0.47 (recall 1.00), the best rule’s 0.40 (0.77) and the [autoencoder](https://one-course.com/books/quant/12/en/chapter/9-cross-sectional-deep-models#def-ml-cross-sectional-deep-models-ae) ’s 0.22 (0.32).
2. Describing the present state for risk scaling and for conditioning other models, quickly and honestly, not predicting switches.
3. No: as a second net, reviewed, with the rules as the primary detectors whose logic can be explained to a regulator.
4. By the features that made each alarm unusual, grouped by account type, against the rules’ verdicts, with outcomes fed back as labels.
5. Its filtered probabilities’ calibration (the frequency of stress when it says 0.8), parameter drift across refits, and out-of-sample likelihood.
6. The tree sets the allocation’s hierarchy; noise clusters make the weights jump with each refit and add turnover without diversification.
7. Confirmed cases from investigations: slow, few, biased towards what the rules already find.
8. What yesterday’s data say, and no more.

## 20.10 Interview questions

**Interview question 20.1 ★ researcher.**

What is the difference between filtered and smoothed state probabilities in an HMM?

**Solution of Interview question 20.1.**

Filtered: the state’s probability given observations up to $t$, available at $t$. Smoothed: given all observations, including after $t$; better for describing history, look-ahead if used for decisions.

*What the interviewer is looking for: the information sets and the look-ahead implication.*

**Interview question 20.2 ★★ researcher, mle.**

How would you cluster 500 stocks by their returns, and how would you know the clusters are real?

**Solution of Interview question 20.2.**

Remove the market (and possibly styles), compute correlations over a window long enough for the effects, cluster on the correlation distance, and check stability across sub-samples and against known classifications.

*What the interviewer is looking for: factor removal, a distance, and a stability check.*

**Interview question 20.3 ★★ researcher.**

Explain the Baum–Welch algorithm in a few sentences.

**Solution of Interview question 20.3.**

EM for HMMs: the E-step runs forward-backward to get each state’s and each transition’s posterior probability; the M-step re-estimates transition probabilities and emission parameters as posterior-weighted averages; repeat until the likelihood stops rising.

*What the interviewer is looking for: forward-backward, posterior weights, re-estimation, likelihood monotonicity.*

**Interview question 20.4 ★★ mle.**

How does an [isolation forest](#def-ml-clustering-regimes-and-anomaly-detection-anomaly) work, and what are its blind spots?

**Solution of Interview question 20.4.**

Random trees split on random features at random values; anomalies need fewer splits to be isolated. Blind spots: anomalies that are not rare (clusters of manipulators), anomalies visible only in combinations the random splits rarely separate, and irrelevant features that add noise.

*What the interviewer is looking for: the mechanism and at least two blind spots.*

**Interview question 20.5 ★★ researcher, trader.**

A regime-switching overlay improved a strategy’s backtested Sharpe ratio from 0.9 to 1.2. What do you check?

**Solution of Interview question 20.5.**

Whether the regime probabilities were filtered and fitted only on past data; whether the number of regimes and parameters were chosen on the whole sample; the result’s significance; and its stability across sub-periods.

*What the interviewer is looking for: smoothing and full-sample fitting as the first suspects.*

**Interview question 20.6 ★★★ mle.**

Design an anomaly-detection system for trade surveillance. How do you evaluate it without many labels?

**Solution of Interview question 20.6.**

Rules for known patterns plus unsupervised scores on engineered features; thresholds set by reviewer capacity; evaluation on planted scenarios and look-alikes, on the few confirmed cases, and by reviewers’ dispositions over time; an audit trail of every alarm and decision.

*What the interviewer is looking for: rules and scores together, capacity-based thresholds, planted-case evaluation, feedback loop.*
