Inference latency is the time from the moment a model’s inputs are available to the moment its output is, measured per call and reported as a distribution (median and tail), within the strategy’s latency budget (Book 13, chapter 1).
Quantitative Finance · Glossário