Machine Learning for Markets · Machine learning
16Generative Models and Synthetic Data
A generative adversarial network trained on twenty years of daily index returns produces windows on which a simple strategy, trained and tested on generated data, earns a Sharpe ratio of 10.7. Trained on the same generated data and tested on the ten real years that followed, it earns 0.5. The network had invented a predictability the market did not have, and it had not learned the one the market did. Synthetic data are one of machine learning’s most useful tools in finance: they let a desk stress a strategy on crashes that happened once, test a backtester on paths with a known truth, and train models where data are scarce. They also mislead in ways that no stylised-fact checklist catches. This chapter builds five generators, from a block bootstrap to a diffusion model, and judges them by four tests; its data are synthetic, so that the real series has a known structure to learn.
16.1 What a generative model is for
Definition 16.1 (Generative model)
A generative model is a model of the distribution of the data that can draw new samples from it. For market data the samples are paths (of returns, order-book states, curves) or windows of them, and the model is judged by whether the samples are useful for the purpose, not only by whether they look real.
The chapter’s real series is thirty years of daily index returns from a GJR-GARCH model with Student-t shocks (Book 4, chapter 18): fat tails, volatility clustering, a leverage effect, six planted crashes of 8 to 11 standard deviations, and one piece of predictability, a conditional mean of 0.1 times the current volatility in the direction of the last twenty days’ return. The first twenty years train; the last ten test. Every generator produces windows of 32 trading days. Its purposes set the tests: stress testing needs the tails and the clustering; training a predictor needs the predictability, and needs it not to be invented; privacy or data sharing needs samples that are not copies.
Two classical generators set the baseline. A block bootstrap (Book 4, chapter 13) cuts the training series into 20-day blocks and glues random blocks together: it keeps everything inside a block and nothing across blocks. A GARCH(1,1) model with Student-t shocks, fitted by maximum likelihood (, , 3.6 degrees of freedom), keeps what its equations say and nothing else: it has no leverage effect and no momentum.
16.2 Variational autoencoders and adversarial networks
Definition 16.2 (Variational autoencoder)
A variational autoencoder (VAE) is an autoencoder (chapter 9) whose encoder outputs a distribution over a low-dimensional latent code and whose decoder outputs a distribution over the data; it is trained to maximise a lower bound on the likelihood, the reconstruction log-likelihood minus the Kullback–Leibler divergence of the encoder’s distribution from a standard normal prior, and it generates by decoding codes drawn from the prior (Kingma and Welling, 2014).
Definition 16.3 (Generative adversarial network, mode collapse)
A generative adversarial network (GAN) trains two networks against each other: a generator that maps noise to samples, and a discriminator that tells generated samples from real ones; the generator is trained to make the discriminator fail (Goodfellow and co-authors, 2014). Mode collapse is the failure in which the generator covers only part of the data’s distribution (a few typical patterns, no extremes) because that is enough to fool the current discriminator.
The chapter’s VAE has a four-dimensional latent code and outputs a mean and a variance for each of the 32 days; its GAN maps 16-dimensional noise through two layers of 128 units. Both are multilayer perceptrons on windows standardised by the training returns’ volatility, trained for 4 000 steps of Adam on one core in a few seconds. The VAE’s windows have excess kurtosis 11.1 against the real training windows’ 13.0; the GAN’s have 2.0: it produces calm and moderately volatile weeks and never a crash, the signature of mode collapse on a fat-tailed distribution. Quant GANs (Wiese and co-authors, 2020) and the small-data market simulator of Buehler and co-authors (2020) are designs aimed at these failures.
16.3 Diffusion models
Definition 16.4 (Diffusion model)
A diffusion model learns to reverse a process that gradually adds Gaussian noise to the data: a network is trained to predict the noise added to a sample at each of steps, and new samples are generated by starting from pure noise and removing the predicted noise step by step (Ho, Jain and Abbeel, 2020).
Listing 16.2 is a complete denoising diffusion model on windows: 100 noise levels, a two-layer network that takes the noisy window and the step, and the reverse chain. It is the most stable of the three networks to train, and on this data the weakest at the tails (kurtosis 4.0): the Gaussian noise it starts from and the few thousand training steps leave the extremes under-represented.
16.4 Judging synthetic data
Definition 16.5 (Classifier two-sample test, train-on-synthetic test-on-real)
The classifier two-sample test trains a classifier to tell generated samples from real ones and reports its accuracy on held-out samples: 50% means the classifier cannot tell them apart (Lopez-Paz and Oquab, 2017). Train-on-synthetic test-on-real (TSTR) trains a downstream model on generated data and evaluates it on real data, comparing with the same model trained on real data (Esteban, Hyland and Rätsch, 2017).
Four tests judge the five generators (Table 16.1). The scorecard checks six stylised facts (Cont, 2001) within windows: heavy tails (excess kurtosis above 5), no linear autocorrelation, volatility clustering in absolute returns at lags 1 and 10, the leverage effect (negative correlation of a return with the next absolute return), gain-loss asymmetry (negative skewness) and aggregational Gaussianity (kurtosis of 20-day sums below half the daily one). The real training windows pass all six. The two-sample test uses a gradient-boosted classifier on 13 window features; real windows overlap, so its halves are split by year, or the classifier recognises a test window by its neighbours in the training half (a random split scores the even training years against the odd ones at 0.83). TSTR fits a linear forecast of a window’s last day from features of its first 31 on 20 000 generated windows and trades the sign of the forecast on the ten real test years; it is averaged over five generated samples. Train-on-synthetic test-on-synthetic (TSTS) does the same with a second generated sample as the test set. The last column is the memorisation check (Listing 16.3): the share of generated windows closer to a training window than 95% of the real test windows are.
| facts | two-sample | TSTR | TSTS | closer than | ||
| generator | passed | accuracy | IC | Sharpe | Sharpe | real (%) |
| real (train on real, test on real) | 6 | 0.545 | 0.078 | 0.86 | 5.0 | |
| block bootstrap | 6 | 0.491 | 0.062 | 0.71 | 1.19 | 14.5 |
| GARCH-t | 4 | 0.539 | 5.7 | |||
| VAE | 6 | 0.619 | 0.040 | 0.76 | 0.79 | 5.1 |
| GAN | 3 | 0.749 | 0.052 | 0.52 | 10.68 | 11.2 |
| diffusion | 4 | 0.693 | 0.030 | 0.23 | 1.45 | 1.4 |
ml_gen.judge.ml_gen.abs_acf.The table holds the chapter’s lessons. The scorecard is the weakest test: the VAE passes all six facts and is told apart from real windows 62% of the time; the block bootstrap passes all six and is indistinguishable (0.491) because its windows are pieces of real ones. The two-sample test sees what the scorecard misses, but not what matters for a strategy. Only TSTR measures that: the real model earns 0.86; the bootstrap’s windows keep most of the momentum inside their 20-day blocks (0.71); the VAE learned a noisy version of it (0.76, with a standard deviation of 0.28 across samples); GARCH, which cannot represent it, teaches nothing (, standard deviation 0.65: pure noise); the diffusion model teaches little (0.23).
16.5 The ways synthetic data mislead
Figure 16.2 is the hook’s result. On the GAN’s own windows a strategy earns a Sharpe ratio of 10.7: the generator produces far more linear autocorrelation (0.099 at lag 1, against 0.078 in the real windows) and persistent volatility patterns that a linear model can trade, an artefact of a network that has collapsed onto a few smooth shapes. Tested on real data, the same strategy earns 0.52. Every generator except GARCH shows a gap between TSTS and TSTR, and nothing inside the synthetic world reveals it.
ml_gen.judge.The memorisation column shows the other two failures. The block bootstrap’s windows are closer to a training window than the real test windows are three times as often as chance (14.5% against 5%): they are copies by construction, which is fine for stress tests and useless where synthetic data must protect the real data’s owner. The GAN’s 11.2% has a different cause: its windows are all calm, and calm windows have close neighbours among the many calm training windows. Nearest-neighbour distance flags copying and collapse alike, and has to be read with the tails. The memorisation of chapter 14 applies to generators as to language models: a large generator trained long enough on a small market history reproduces it.
Method 16.6 (Using synthetic market data)
- Name the purpose (stress, training, testing a backtester, sharing), and choose the tests it needs.
- Start from the classical generators (bootstrap, GARCH, a simulator with known mechanisms) and require a network to beat them on those tests.
- Judge by TSTR against train-on-real, never by results inside the synthetic world; add the two-sample test with a blocked split and the memorisation check.
- Treat any edge found only in synthetic data as an artefact until it is found in real data.
16.6 Tutorial: paths that never happened
Goal. Fit five generators to the training years, and judge them by scorecard, two-sample test, TSTR, TSTS and nearest-neighbour distance. End state: Table 16.1, Figures 16.1 and 16.2.
A generative adversarial network.
class GAN(_Scaled): def __init__(self, L=32, noise=16, hidden=(128, 128)): self.L, self.noise = L, noise self.G = _mlp(noise, hidden, L) self.D = _mlp(L, hidden, 1) def fit(self, W, steps=3000, seed=0, batch=256, lr=2e-4): g = _seed(seed, self.G, self.D) X = self._set_scale(W) oG = torch.optim.Adam(self.G.parameters(), lr=lr, betas=(0.5, 0.999)) oD = torch.optim.Adam(self.D.parameters(), lr=lr, betas=(0.5, 0.999)) bce = nn.BCEWithLogitsLoss() ones, zeros = torch.ones(batch, 1), torch.zeros(batch, 1) for _ in range(steps): xb = X[torch.randint(len(X), (batch,), generator=g)] fake = self.G(torch.randn((batch, self.noise), generator=g)) lD = bce(self.D(xb), ones) + bce(self.D(fake.detach()), zeros) oD.zero_grad() lD.backward() oD.step() lG = bce(self.D(fake), ones) oG.zero_grad() lG.backward() oG.step() return self def sample(self, n, seed=0): g = torch.Generator().manual_seed(seed) with torch.no_grad(): return self.G(torch.randn((n, self.noise), generator=g)).numpy() * self.scaleListing 16.1. A generator and a discriminator trained against each other. code/firm/genmkt/firm_genmkt.py A denoising diffusion model.
class Diffusion(_Scaled): """Denoising diffusion (Ho, Jain and Abbeel): noise the windows over T steps with a linear variance schedule, train a network to predict the noise from the noisy window and the step, and sample by running the chain backwards.""" def __init__(self, L=32, T=100, hidden=(128, 128)): self.L, self.T = L, T self.net = _mlp(L + 1, hidden, L) self.beta = torch.linspace(1e-4, 0.05, T) self.abar = torch.cumprod(1 - self.beta, 0) def fit(self, W, steps=3000, seed=0, batch=256, lr=2e-3): g = _seed(seed, self.net) X = self._set_scale(W) opt = torch.optim.Adam(self.net.parameters(), lr=lr) for _ in range(steps): xb = X[torch.randint(len(X), (batch,), generator=g)] t = torch.randint(self.T, (batch,), generator=g) e = torch.randn(xb.shape, generator=g) ab = self.abar[t][:, None] xt = ab.sqrt() * xb + (1 - ab).sqrt() * e loss = ((self.net(torch.cat([xt, (t[:, None] + 0.5) / self.T], 1)) - e) ** 2).mean() opt.zero_grad() loss.backward() opt.step() return self def sample(self, n, seed=0): g = torch.Generator().manual_seed(seed) x = torch.randn((n, self.L), generator=g) with torch.no_grad(): for t in range(self.T - 1, -1, -1): tt = torch.full((n, 1), (t + 0.5) / self.T) e = self.net(torch.cat([x, tt], 1)) b, ab = self.beta[t], self.abar[t] x = (x - b / (1 - ab).sqrt() * e) / (1 - b).sqrt() if t > 0: x = x + b.sqrt() * torch.randn(x.shape, generator=g) return x.numpy() * self.scaleListing 16.2. Noise, a denoiser, and the reverse chain. code/firm/genmkt/firm_genmkt.py Train on synthetic, test on real.
def tstr(W_train, W_test): """Fit a linear forecast of each window's last day from its features on W_train; score on W_test: the rank IC and the annualised Sharpe ratio of trading the sign of the forecast.""" from scipy.stats import spearmanr Xa, ya = window_features(W_train), W_train[:, -1] Xa1 = np.column_stack([np.ones(len(Xa)), Xa]) b = np.linalg.lstsq(Xa1, ya, rcond=None)[0] Xb = np.column_stack([np.ones(len(W_test)), window_features(W_test)]) f = Xb @ b pnl = np.sign(f - np.median(f)) * W_test[:, -1] return float(spearmanr(f, W_test[:, -1])[0]), float(pnl.mean() / pnl.std() * math.sqrt(252))Listing 16.3. The TSTR score: rank IC and Sharpe ratio on real windows. code/firm/genmkt/firm_genmkt.py - Run
ml_gen.judge()(about a minute on one core) andfig_gen.py.
What to change next. Condition the generators on the volatility regime; train the GAN longer and watch its TSTS Sharpe ratio and its two-sample accuracy; replace the block bootstrap’s 20-day blocks by 5-day blocks and watch the momentum disappear from its TSTR.
16.7 Build: generators and their judges
Purpose. Synthetic market data with a stated purpose, judged by tests that match it.
Interface. real_index(n_days, seed), windows; block_bootstrap, fit_garch, garch_windows; VAE, GAN, Diffusion with fit(W, steps, seed) and sample(n, seed); facts, scorecard, c2st(A, B, groups_a, groups_b, seed), tstr(W_train, W_test), nn_distance.
Rules. Every generator is fitted on the training period only; overlapping real windows are split by time; TSTR is averaged over samples; results inside the synthetic world are never reported alone.
Acceptance tests. code/firm/genmkt/tests/: the GARCH fit recovers planted parameters; the bootstrap copies blocks; the two-sample test is near 0.5 on two samples of one distribution and near 1 on two different ones; TSTR recovers a planted predictor; the networks train deterministically and reduce their losses; nearest-neighbour distance is zero for copies.
Stretch. Conditional generation on regimes; signature-based generators; Book 10’s agent-based market (chapter 27) as a generator with mechanisms, judged by the same tests.
Sources and further reading
- D. P. Kingma and M. Welling, “Auto-encoding variational Bayes”, arXiv:1312.6114, 2014.
- I. Goodfellow and co-authors, “Generative adversarial nets”, arXiv:1406.2661, 2014.
- J. Ho, A. Jain and P. Abbeel, “Denoising diffusion probabilistic models”, arXiv:2006.11239, 2020.
- M. Wiese, R. Knobloch, R. Korn and P. Kretschmer, “Quant GANs: deep generation of financial time series”, Quantitative Finance 20(9), 2020.
- H. Buehler, B. Horvath, T. Lyons, I. Perez Arribas and B. Wood, “A data-driven market simulator for small data environments”, arXiv:2006.14498, 2020.
- D. Lopez-Paz and M. Oquab, “Revisiting classifier two-sample tests”, arXiv:1610.06545, 2017.
- R. Cont, “Empirical properties of asset returns: stylized facts and statistical issues”, Quantitative Finance 1(2), 2001.
- C. Esteban, S. L. Hyland and G. Rätsch, “Real-valued (medical) time series generation with recurrent conditional GANs”, arXiv:1706.02633, 2017.
16.8 Exercises
Exercise 16.1 ★
A two-sample test scores 0.62 on 4 000 generated against 4 000 real windows (2 000 of each held out). Is the difference from 0.5 significant? What would 0.51 tell you?
Solution
Solution of Exercise 16.1.
Under the null the held-out accuracy on 4 000 windows has standard error : 0.62 is 15 standard errors above 0.5, a clear difference. 0.51 is 1.3 standard errors: no evidence of a difference, which is not evidence of equality; a stronger classifier or more samples might find one.
Exercise 16.2 ★
Which of the six stylised facts can a GARCH(1,1) model with symmetric Student-t shocks never produce, and why?
Solution
Solution of Exercise 16.2.
The leverage effect and gain-loss asymmetry: with symmetric shocks and a variance that responds to squared shocks, negative and positive returns raise volatility equally and the distribution of returns is symmetric. A GJR or EGARCH term and skewed shocks add them.
Exercise 16.3 ★
The fitted GARCH has . What is the half-life of a volatility shock in days?
Solution
Solution of Exercise 16.3.
A shock to the variance decays like ; the half-life is days.
Exercise 16.4 ★★
Why does the block bootstrap keep the momentum effect partly, and what block length would destroy it?
Solution
Solution of Exercise 16.4.
The effect links a day’s return to the previous twenty days’ sum; inside a 20-day block the ordering of real days is kept, so a window’s last day and its preceding days are often from the same block, and the relation survives there. Across block boundaries it is broken. Blocks of a few days destroy it; blocks much longer than twenty days keep it almost intact, and copy more of the history.
Exercise 16.5 ★★
Why did a random split make the two-sample test score real windows against real windows at 0.83?
Solution
Solution of Exercise 16.5.
Neighbouring windows share 31 of their 32 days. With a random split, almost every held-out window has a near-copy in the training half with the same label, so the classifier recognises the period, not the distribution: 0.83 for two halves of the same process. Splitting by year leaves no neighbours across the split.
Exercise 16.6 ★★
Find the flaw. “Our GAN’s paths pass all our stylised-fact tests, and a strategy trained on a million of them has a Sharpe ratio of 3 out of sample on another million, so it is robust.”
Solution
Solution of Exercise 16.6.
The test is inside the GAN’s world: a strategy trained and tested on generated paths measures the generator’s regularities, including artefacts. Here a GAN gives 10.7 inside its world and 0.52 on real data. Passing stylised-fact tests says nothing about predictability. Test the strategy on real data it was not trained on, and compare with training on real.
Exercise 16.7 ★★★
Coding. Train the GAN for 1 000 and for 8 000 steps. Report its excess kurtosis, two-sample accuracy and TSTS Sharpe ratio each time. Does longer training cure the collapse?
Solution
Solution of Exercise 16.7.
After 1 000 steps: excess kurtosis 2.53, two-sample accuracy 0.884, TSTS Sharpe ratio 11.2. After 8 000: 2.49, 0.734 and 9.9 (4 000 steps gave 1.98, 0.749 and 10.7). Longer training makes the windows harder to tell apart on average but does not restore the tails, which a discriminator on standardised windows rarely sees, and the invented predictability stays. The cures are architectural (heavy-tailed noise, a volatility channel as in Quant GANs) or a likelihood-based model.
Exercise 16.8 ★★★
Derive the VAE’s objective: show that , and write the KL term for a diagonal Gaussian and a standard normal prior.
Solution
Solution of Exercise 16.8.
by Jensen’s inequality, and the second term is . For and : , the term in the code.
16.9 Problem: Paths That Never Happened
Problem 16.1
Weekend problem — shape is not predictability
The chapter’s thirty years of returns, five generators and four tests.
Part I — The data and the generators.
- What structure does the real series have, and which parts can each generator represent?
- What are the fitted GARCH parameters, and what does the model miss?
- What do the VAE and GAN networks look like, and how long do they train?
- What is the diffusion model’s generating process?
Part II — Judging.
- How many stylised facts does each generator pass?
- What are the two-sample accuracies, and why must real windows be split by time?
- What are the TSTR information coefficients and Sharpe ratios against train-on-real?
- What does the memorisation check show?
Part III — Misleading.
- What Sharpe ratio does the GAN’s world give, and where does it come from?
- Why is GARCH’s TSTR noise, and how large is that noise?
- Which generator would you use for a stress test, and which to augment a training set?
- Why is the bootstrap unsuitable for sharing data outside the firm?
Part IV — The verdict.
- State the named result: each generator’s scorecard pass count and two-sample accuracy, and the TSTR gap in Sharpe ratio.
- What would convince you that a generator has learned a real predictability?
- How would you test a backtester with a generator?
- When is a mechanistic simulator better than a learned generator?
- How would you condition a generator on a regime, and what new risk does it bring?
- What would you log about a synthetic dataset used in research?
- What does the diffusion model’s weakness at the tails imply for stress tests?
- In one sentence: what does a generative model learn from twenty years of returns?
Solution
Solution of Problem 16.1.
Part I.
- Fat tails, volatility clustering, a leverage effect, six crashes and a 20-day momentum effect. The bootstrap keeps what fits in a block; GARCH-t keeps tails and clustering; the networks can represent any of it in principle.
- , , 3.6 degrees of freedom; it misses the leverage effect, the skew and the momentum.
- Multilayer perceptrons on standardised 32-day windows: a VAE with a four-dimensional code and a GAN with 16-dimensional noise and two layers of 128 units, 4 000 Adam steps each on one core.
- Start from Gaussian noise and apply 100 learned denoising steps, each removing the predicted noise and adding a little fresh noise.
Part II.
- Bootstrap 6, GARCH-t 4, VAE 6, GAN 3, diffusion 4 (real 6).
- 0.491, 0.539, 0.619, 0.749, 0.693 (real against real 0.545). Overlapping windows leak across a random split: 0.83 for real against real.
- IC and Sharpe: real 0.078 and 0.86; bootstrap 0.062 and 0.71; GARCH-t and ; VAE 0.040 and 0.76; GAN 0.052 and 0.52; diffusion 0.030 and 0.23.
- Bootstrap windows are close to training windows 14.5% of the time (copies), GAN windows 11.2% (collapse onto calm windows), the others near the 5% of real test windows.
Part III.
- 10.7, from linear autocorrelation and smooth volatility patterns the GAN invented.
- GARCH-t has no predictability, so the fitted forecast is a random combination of features whose real-data IC is random in sign: on average with a standard deviation of 0.65 across five samples.
- A stress test needs tails and clustering: bootstrap or GARCH-t, with the crashes; augmentation needs the predictability without artefacts: none of the networks here is good enough, and the bootstrap’s copies add little new.
- It reproduces real 20-day sequences: anyone with the real data can recognise them.
Part IV.
- Paths that never happened. Facts passed and two-sample accuracy: bootstrap 6 and 0.491, GARCH-t 4 and 0.539, VAE 6 and 0.619, GAN 3 and 0.749, diffusion 4 and 0.693. TSTR Sharpe against 0.86 trained on real: 0.71, , 0.76, 0.52 and 0.23; the GAN’s own world gives 10.7.
- TSTR close to train-on-real across several samples and fits, on a test period after the training data.
- Plant known effects in generated data, run the backtester, and check that it recovers them and nothing else.
- When the mechanism is the question (market impact, queue dynamics, a policy change), and when data are too few to learn a distribution.
- Feed the regime as an input to the generator; the new risk is regimes that are rare in training and poorly generated exactly where stress tests need them.
- The generator’s version and training data window, its seeds, the tests it passed, its purpose, and which results used it.
- Tail losses from diffusion samples would be understated; stress with the tails from the real data or a heavy-tailed model.
- Its shape, and only as much of its predictability as the architecture and the data force it to learn.
16.10 Interview questions
Interview question 16.1 ★ researcher, mle
Name the main stylised facts of daily returns. Which does a GARCH model reproduce?
Solution
Solution of Interview question 16.1.
Heavy tails, no linear autocorrelation, volatility clustering, the leverage effect, gain-loss asymmetry, aggregational Gaussianity, among others. GARCH with fat-tailed shocks reproduces tails, clustering and aggregation; asymmetric variants add leverage.
What the interviewer is looking for: four or more facts, and which a GARCH reproduces.
Interview question 16.2 ★★ mle
Explain the difference between a VAE, a GAN and a diffusion model in one paragraph each, and one failure of each.
Solution
Solution of Interview question 16.2.
VAE: an encoder to a latent distribution and a decoder, trained on a likelihood bound; samples are blurry or too Gaussian. GAN: a generator trained against a discriminator; mode collapse and unstable training. Diffusion: learned denoising of gradually noised data; slow sampling, and tails depend on training. Each fails differently on fat-tailed data.
What the interviewer is looking for: the training principle and a failure for each.
Interview question 16.3 ★★ researcher
How would you evaluate a generator of synthetic market data?
Solution
Solution of Interview question 16.3.
By purpose: stylised facts, a two-sample test with a time-blocked split, TSTR against train-on-real on later data, and a memorisation check; never by results inside the synthetic data alone.
What the interviewer is looking for: TSTR and a blocked two-sample test, tied to the purpose.
Interview question 16.4 ★★ researcher, trader
A colleague augments the training set with GAN paths and the backtest improves. What do you check?
Solution
Solution of Interview question 16.4.
Whether the improvement holds on real data not used to fit the generator; whether the generator was trained on the test period; whether synthetic paths carry artefacts the model learned (TSTS against TSTR); and whether the same effect appears with a bootstrap.
What the interviewer is looking for: real out-of-sample evaluation and generator leakage.
Interview question 16.5 ★★ mle
What is mode collapse, and how would you detect it in generated returns?
Solution
Solution of Interview question 16.5.
The generator covers only part of the distribution. Signs in returns: thin tails, too little diversity in volatility levels, generated windows close to a few typical real ones, a two-sample classifier winning on tail features.
What the interviewer is looking for: the definition and concrete diagnostics.
Interview question 16.6 ★★★ researcher
Design a classifier two-sample test for overlapping windows of a time series. What goes wrong with a random split?
Solution
Solution of Interview question 16.6.
Build features per window, split real windows by time blocks (with a gap longer than the window) so that no held-out window overlaps a training one, split generated windows at random, train, report held-out accuracy with its standard error. A random split puts near-copies on both sides and measures recognition of periods.
What the interviewer is looking for: blocked splitting with a gap, and why overlap inflates accuracy.