tezvyn:

The Second Story of an Incident

AI-drafted, machine-checkedintermediate

The first story blames human error and stops there; the second story asks why the action made sense to the person at the time and what systemic conditions enabled it. Seeking the second story is the heart of blameless, learning-oriented incident analysis.

WHY IT EXISTS Stopping an investigation at human error feels satisfying but explains nothing and fixes nothing, because the next person facing the same conditions will make the same choice. The concept of the second story exists to push analysis past blame into the systemic factors that made the error likely and even reasonable.

THE MENTAL MODEL Every incident has two stories. The first story is the convenient narrative: a person did the wrong thing, end of analysis. The second story assumes the person was competent and acting in good faith, then asks why their action made sense to them at the time, given what they could see, the tools they had, the time pressure, and how the system was designed. Human error becomes the starting point of inquiry, not its conclusion.

HOW IT WORKS To reach the second story, you reconstruct the responder's local rationality: what information was and was not visible, what the interface or alert implied, what defaults or ambiguous design nudged them, and what production pressure shaped the decision. From that you identify systemic contributors, a confusing UI, a missing guardrail, a misleading metric, and you fix those. This is the operational expression of a blameless culture, drawn from how high-reliability fields study accidents.

WHEN IT MATTERS It matters whenever an incident is quickly labeled operator error or fat-finger, which is precisely when the real lessons are about to be missed. Seeking the second story is what makes postmortems generative rather than punitive, and it directly protects psychological safety, since people disclose mistakes honestly only when they trust the analysis will examine the system, not punish them.

ONE CONCRETE EXAMPLE First story: an engineer ran a delete command against production and caused an outage, so the cause is human error. Second story: the staging and production terminals looked identical, the command had no confirmation prompt, and the engineer was rushing during an unrelated incident. The fixes that follow are distinct prompts for production, a confirmation guard, and clearer environment indicators, none of which the first story would ever surface.

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.