A postmortem is not a report. It is a mechanism for turning one team’s bad week into the whole organization’s improvement — and most of them fail at that.
The failure mode is always the same: the document lists a timeline, names a root cause, assigns three action items, and gets archived. Six months later the same class of incident happens again.
What changed things for my teams was treating action items as engineering work with owners and deadlines tracked in the same backlog as features, capping them at three per incident, and reviewing overdue ones in a monthly reliability meeting that leadership actually attends.
The other shift: stop asking “what was the root cause” and start asking “what made this hard to detect, hard to diagnose, and hard to fix?” Those three questions surface the systemic issues — missing alerts, tribal knowledge, slow deploys — that a single root cause hides.