Incident Simulator

A page, a timeline and a wall of signals. The cause is not revealed, and neither is the mitigation, because the skill being practised is deciding what to look at when you do not yet know — and noticing that the most alarming signal is often a consequence rather than a cause.

SIMULATED

The clock times are the scenario's premise, not measurements. What transfers is the reasoning: what changed, in what order did the signals move, and does the symptom track the thing you suspect.

The page you just got

At 14:03 the error rate rose from 0.2% to 12%. A deployment went out at 14:01. You are on call.

Timeline

A change is not a signal, and a signal is not a cause. Read the ordering before you read the labels.

  1. 14:01changeDeployment of v2.14.0 begins (rolling, 25% batches)
  2. 14:03signalError rate rises from 0.2% to 12%
  3. 14:04signalPage fires on the checkout success-rate alert
  4. 14:05changeRollout is at 50% of instances
  5. 14:06actionResponder acknowledges
changesignalactionrecovery

Signals

Mark the ones you think actually narrow the cause. A signal that would read the same under several different causes is not discriminating, however alarming it looks.