Incident Response
Detect, triage, mitigate, communicate, recover. Stopping user impact before understanding cause, and the roles that keep a severe incident coordinated.
Alert, acknowledge, triage, mitigate, recover, verify, learn — a defined sequence, so nobody has to invent one at 3am.
A shared shorthand for how much of the organisation to wake — and a local convention, not a fact about software.
The mandatory distinction: mitigation ends user impact, root cause analysis explains it, and they happen in that order.
Separating coordination from investigation so that neither starves the other — valuable at high severity, overhead at low.
An evidence-based sequence of changes, signals and actions — built from records, because memory reorders events with total confidence.
Different audiences need different things at different cadences — and none of them should have to interrupt the person fixing it.
Someone has to be reachable when production breaks. Done well it is the shortest feedback loop a team has; done badly it is the fastest way to lose people.
Load, frequency and recovery are properties of the system, and a rotation that cannot be sustained is a defect in the system rather than a shortcoming of a person.