Incident Simulator
A page, a timeline and a wall of signals. The cause is not revealed, and neither is the mitigation, because the skill being practised is deciding what to look at when you do not yet know — and noticing that the most alarming signal is often a consequence rather than a cause.
The clock times are the scenario's premise, not measurements. What transfers is the reasoning: what changed, in what order did the signals move, and does the symptom track the thing you suspect.
At 14:03 the error rate rose from 0.2% to 12%. A deployment went out at 14:01. You are on call.
Timeline
A change is not a signal, and a signal is not a cause. Read the ordering before you read the labels.
- 14:01changeDeployment of v2.14.0 begins (rolling, 25% batches)
- 14:03signalError rate rises from 0.2% to 12%
- 14:04signalPage fires on the checkout success-rate alert
- 14:05changeRollout is at 50% of instances
- 14:06actionResponder acknowledges
Signals
Mark the ones you think actually narrow the cause. A signal that would read the same under several different causes is not discriminating, however alarming it looks.