tezvyn:

Error budget exhausted early: what now?

AI-drafted, machine-checkedSource: interviewintermediate
WHAT IT TESTS

whether you use the error budget as a decision tool, not punishment.

OUTLINE

invoke the error budget policy, shift focus from features to reliability, prioritize stability work, analyze what burned the budget.

WHAT THIS TESTS The question probes whether you understand the error budget as a governance mechanism that converts reliability into an objective, shared currency, and whether you can lead a constructive, data-driven response rather than a reactive or punitive one.

A GOOD ANSWER COVERS Note that exhausting the budget should trigger a previously agreed error budget policy, so the response is not improvised. Typical policy actions include pausing or slowing feature releases, prioritizing reliability and bug-fix work, and requiring stricter rollout safeguards until the budget recovers. Crucially, drive the decision with data: analyze what consumed the budget, whether a specific bad release, a flaky dependency, an under-provisioned component, or a single large incident, so the remedy targets the real cause. Frame the discussion with product and dev as an objective tradeoff between velocity and reliability that both sides agreed to, which depersonalizes it. The budget turns a values argument into a numbers conversation.

COMMON WRONG ANSWERS Ignoring the overrun and continuing as normal. Treating it as a reason to blame or penalize individuals. Imposing a rigid total freeze with no analysis, or conversely waving it away because revenue features matter more. Having no pre-agreed policy at all.

LIKELY FOLLOW-UPS Who has authority to override a freeze? How should the policy be written beforehand? What if a single incident, not chronic issues, spent the budget? How do you prevent recurrence next quarter?

ONE CONCRETE EXAMPLE Analysis shows eighty percent of the burn came from one botched migration that caused a multi-hour outage, not steady degradation. Rather than freezing all launches, the team invokes the policy to pause risky launches, fast-tracks the rollback safeguards and canary process that would have caught the migration, and presents product with the burn breakdown. Product agrees to delay the next risky feature two weeks while the safeguards ship, a balanced, evidence-based outcome the budget made possible.

Read the original → sre.google

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.