Postmortems
Blameless but accountable learning: contributing factors over single root causes, and action items specific enough to change the system rather than the people.
The written reconstruction of an incident — impact, timeline, detection, contributing factors, what went well and what changes — done blamelessly and still accountably.
Why complex systems rarely have one cause, and why "human error" is a question rather than an answer.
One investigative technique for pushing past the first plausible answer — useful, widely over-applied, and structurally unable to represent multiple interacting causes.
"Be more careful" is not an action item. "Add migration validation", "add a canary", "reduce the permission", "automate the verification" are.
The learning that only exists in aggregate — patterns across incidents, near misses, and making the record something people actually read.