MonOps
The whole lifecycle
Stage 0510:02:38 · your team

Closed when the signal says so

Recovery has to hold before anything is marked green.

Closing on the first good sample is how you end up reopening the same incident twenty minutes later. A restarted service looks healthy for a few seconds regardless of whether the underlying problem is fixed.

Recovery has to hold across four consecutive sweeps. A machine that recovers and immediately degrades never reaches closed, and the incident stays open with the flapping recorded on it.

Separately, incidents whose source stopped reporting entirely are reaped after ten minutes — an incident about a machine that no longer exists is noise, and leaving it open trains people to ignore the list.

Defaults at this stage

sweeps to confirm
4
stale incidents reaped
600s

All six stages, on every plan with agents.

Detection through post-mortem in one system. No second vendor for paging.

Deploy an agent