All books

Professional

Apps About Coach Log in Start reading

Quantitative Finance · Glossary

What is Classifier two-sample test, train-on-synthetic test-on-real?

Also known as: classifier two-sample test · train-on-synthetic test-on-real

Definition 16.5 Machine Learning for Markets · Chapter 16 — Generative Models and Synthetic Data

The classifier two-sample test trains a classifier to tell generated samples from real ones and reports its accuracy on held-out samples: 50% means the classifier cannot tell them apart (Lopez-Paz and Oquab, 2017). Train-on-synthetic test-on-real (TSTR) trains a downstream model on generated data and evaluates it on real data, comparing with the same model trained on real data (Esteban, Hyland and Rätsch, 2017).

factstwo-sampleTSTRTSTScloser than
generatorpassedaccuracyICSharpeSharpereal (%)
real (train on real, test on real)60.5450.0780.865.0
block bootstrap60.4910.0620.711.1914.5
GARCH-t40.539−0.019-0.019−0.10-0.10−0.09-0.095.7
VAE60.6190.0400.760.795.1
GAN30.7490.0520.5210.6811.2
diffusion40.6930.0300.231.451.4
Table 16.1. Five generators judged four ways. Two-sample accuracy of the real row: the training years’ even against odd years. Data: ml_gen.judge.
Volatility clustering: autocorrelation of absolute daily returns within 32-day windows, real training windows against 4 000 windows from each generator. Data: ml_gen.abs_acf.
Figure 16.1. Volatility clustering: autocorrelation of absolute daily returns within 32-day windows, real training windows against 4 000 windows from each generator. Data: ml_gen.abs_acf.
Read in context →