Anomaly detection scores observations by how unlike the bulk of the data they are, without labels for the anomalies. An isolation forest builds random trees that split on random features at random values; points that are isolated after few splits are anomalous (Liu, Ting and Zhou, 2008). The precision–recall curve plots, as the alarm threshold moves, the share of alarms that are true (precision) against the share of true cases that raise an alarm (recall); when positives are rare it is more informative than the ROC curve (Davis and Goadrich, 2006).
| detector | recall at 1% false positives | precision |
|---|---|---|
| order-to-trade ratio (rule) | 0.00 | 0.00 |
| fill-rate gap (rule) | 0.37 | 0.24 |
| cancellations after a fill (rule) | 0.51 | 0.31 |
| gap cancellations (rule) | 0.77 | 0.40 |
| isolation forest | 1.00 | 0.47 |
| autoencoder | 0.32 | 0.22 |
ml_regimes.anomalies.