Alle boeken

Professioneel

Apps Over Coach Inloggen Begin met lezen

Quantitative Finance · Begrippenlijst

Wat is Time to detect, time to restore?

Ook bekend als: time to detect · time to restore

Definition 28.7 Research, Data and Risk Platforms · Hoofdstuk 28 — Observability and Incident Response

The time to detect of an incident is the time from its start to the first alert or report that a person acts on; its time to restore is the time from its start to the moment the service users receive is back within its objective.

The slow-consumer incident: from 14:00 the risk service processes messages slower than they arrive and the age of its data grows by 0.05 seconds each second. Its process stays alive throughout. The five-second threshold fires at 1.7 minutes, the burn-rate rule at 2.5, the 30- and 60-second thresholds after 10 and 20 minutes. Data: fig_observe.py.
Figure 28.3. The slow-consumer incident: from 14:00 the risk service processes messages slower than they arrive and the age of its data grows by 0.05 seconds each second. Its process stays alive throughout. The five-second threshold fires at 1.7 minutes, the burn-rate rule at 2.5, the 30- and 60-second thresholds after 10 and 20 minutes. Data: fig_observe.py.
The intermittent stalls: 20 seconds without data every minute from 11:48. The data age never exceeds 20 seconds, so thresholds at 30 and 60 seconds never fire; the burn rate over the last hour passes 14.4 — one fiftieth of the month’s budget in an hour — after 3.2 minutes, with the five-minute window confirming it. Data: fig_observe.py.
Figure 28.4. The intermittent stalls: 20 seconds without data every minute from 11:48. The data age never exceeds 20 seconds, so thresholds at 30 and 60 seconds never fire; the burn rate over the last hour passes 14.4 — one fiftieth of the month’s budget in an hour — after 3.2 minutes, with the five-minute window confirming it. Data: fig_observe.py.
Lees in het hoofdstuk →