Incident simulations
Six production incidents. You get the page that woke someone up, a timeline in the order things were observed, and evidence you pull up one item at a time. Observation order is not causal order — the first thing anyone noticed is almost never the cause.
p99 checkout latency hits 6 s while CPU, memory and database all look healthy.
SLO BURN · checkout-api · latency SLO (99% < 800 ms) burning at 14× · 2h budget remaining · p99 = 5,940 ms
A cache node reboots, the database takes eight times its normal load, and client retries turn a blip into a collapse.
CRITICAL · catalog-api · error rate 31% (threshold 2%) · p99 12.4 s · 5xx from 34 of 40 instances
Latency spikes worsen through the day, a restart fixes it, and by evening it is back.
WARN · recommendations-api · pod restart count 4 in 12 h (OOMKilled) · p99 1,840 ms (threshold 900 ms)
Traffic is flat, throughput halves, and tripling the worker count changes almost nothing.
WARN · notification-worker · oldest message age 2,410 s (threshold 300 s) · queue depth 184,000 and rising
The median is flat, p95 is flat, p99 is twice what it was, and there was no deploy.
WARN · search-api · p99 latency 1,910 ms (baseline 940 ms) · p50 and p95 within normal range · no deploy in 9 days
A marketing email drives 6× traffic in 90 seconds; new instances arrive four minutes later, to a fleet that has already collapsed.
CRITICAL · storefront-api · availability 61% (SLO 99.9%) · p99 timeout · 5xx 4,200/s