tezvyn:

Reviewing a large-scale cascading outage

AI-drafted, machine-checkedSource: interviewadvanced
WHAT IT TESTS

Scaling the review to a complex outage.

OUTLINE

dedicated facilitator, cross-team timeline reconciliation, map cascade chains and multiple contributing factors, layered action items.

WHAT THIS TESTS Whether you right-size the review process and handle the realities of cascading, cross-team failures rather than shoehorning them into a routine template.

A GOOD ANSWER COVERS A routine single-system incident often suits a lightweight template, one team, a short timeline, a clear trigger, and a handful of action items, sometimes self-reviewed. A multi-day, multi-team cascade demands more structure. Appoint a dedicated, neutral facilitator who did not respond to the incident. Invest heavily in reconstructing a single authoritative timeline by reconciling each team's logs, chats, and dashboards, since perspectives will conflict. Explicitly map the cascade: how a failure in one system propagated, what amplified it, and the several contributing factors and missing circuit breakers along the way, resisting the urge to name one root cause. Involve leadership and all affected teams, and produce layered action items spanning architecture, tooling, and process, with cross-team ownership and a follow-up review to confirm completion.

COMMON WRONG ANSWERS Using the routine template for a massive outage. Forcing a single root cause. Letting one team write it in isolation, missing cross-team interactions.

LIKELY FOLLOW-UPS How do you reconcile conflicting team timelines? How do you assign cross-team action items? When do you escalate to an org-wide review?

ONE CONCRETE EXAMPLE A three-day outage began with a database failover bug, which overloaded a cache, which made an API team retry aggressively, amplifying load org-wide. A neutral facilitator builds one merged timeline from five teams, maps each link in the cascade, and identifies multiple contributing factors: the failover bug, the missing cache circuit breaker, and the unbounded retries. The action items are layered, owned across the three teams, and a follow-up review weeks later verifies the architectural fixes shipped, unlike the quick template used for a routine single-service blip.

Read the original → sre.google

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.