Defining SLOs for a new critical service
Cross-functional SLO design.
Start from user journeys, pick SLIs, involve product, engineering, and business stakeholders, set realistic targets iteratively.
WHAT THIS TESTS This evaluates whether you understand SLO definition as a cross-functional, user-centered exercise rather than a purely technical one, and whether you can name the right people and considerations.
A GOOD ANSWER COVERS Start by identifying the critical user journeys the service supports, because SLOs should protect what users actually do, not whatever happens to be easy to measure. For each journey, choose SLIs that faithfully capture that experience, such as availability and latency at the relevant entry point. Then set target values that are realistic, informed by historical performance and by what users genuinely need, and accept that the first targets are provisional and will be tuned as you gather data. Key stakeholders include product managers who own user expectations and business priorities, the engineers who build and operate the service, SRE for reliability expertise, and business or account leadership who understand contractual and competitive stakes. Beyond raw metrics, weigh user tolerance for failure, the rising cost of each additional nine, hard limits imposed by upstream and downstream dependencies, regulatory or contractual obligations, and competitive expectations.
COMMON WRONG ANSWERS Defining SLOs solely from whatever dashboards already exist, divorced from user journeys. Setting an aspirational ninety-nine point nine nine nine without regard to cost or feasibility, or excluding product and business voices so the targets do not reflect real priorities.
LIKELY FOLLOW-UPS How do you pick the measurement window? How do you handle a dependency whose own SLO caps yours? How often do you revisit SLOs? How do you avoid too many SLOs diluting focus?
ONE CONCRETE EXAMPLE For a new checkout service, you map the journey of adding to cart and completing payment. You pick availability and p95 latency SLIs at the checkout API. Looking at beta data showing about 99.7 percent success, and consulting product on revenue sensitivity, you set a starting SLO of 99.9 percent availability over twenty-eight days, deliberately modest to stay achievable. Product, the checkout engineers, SRE, and a finance stakeholder all sign off, and you schedule a review in one quarter to ratchet the target based on observed performance and customer feedback.
Read the original → cloud.google.com
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.