Target leakage is information in a feature, or in any step of the pipeline, that would not have been available at the decision time the prediction stands for, typically because it was derived from the target or recorded after it. Train–test contamination is information about the test data reaching the fit through the data themselves: overlapping labels or duplicated records on both sides of a split, or statistics computed on the whole sample.
| leak | leaky | honest version | caught by |
|---|---|---|---|
| period-end join of the surprise | 1.27% | (filing date) | truncation test |
| target encoding, same month | 8.97% | (none) | canary, truncation test |
| screening on the whole sample | 0.13% | (training months) | canary, truncation test |
| shuffled folds, overlapping labels | 3.47% | (purged folds) | fold-overlap check |
| best of 30 seeds on the test | (untouched months) | holdout, nested search |
ml_validation.leaks.