First-SRE 90-day plan at a startup
Whether you can introduce SRE incrementally and show value.
Listen and measure first, pick one high-impact service, define SLIs/SLOs and basic alerting, then show reduced toil or incidents to leadership.
WHAT THIS TESTS: Whether you can introduce SRE as an incremental, evidence-driven practice rather than a sweeping mandate, and whether you understand that the first SRE must earn trust and prove ROI before scaling process.
A GOOD ANSWER COVERS: Phase one, roughly the first thirty days, is listening and measuring: understand the architecture, sit with on-call, inventory recent incidents and recurring toil, and figure out which single service is most business-critical and most painful. Pick that one service as the beachhead rather than boiling the ocean. Phase two, days thirty to sixty, is instrumenting and defining: establish meaningful SLIs (latency, error rate, availability) for that service, set realistic SLOs with the team, replace noisy cause-based alerts with symptom and SLO-based alerts, and write runbooks for the top failure modes. Phase three, days sixty to ninety, is demonstrating value with hard numbers leadership cares about: reduced page volume, lower MTTR, fewer customer-facing incidents, recovered engineering hours from automating toil, and the service tracking against its SLO. Use that proof to propose expanding SRE practices to the next service.
COMMON WRONG ANSWERS: Imposing org-wide error budget policies, mandatory on-call rotations, and heavy process on day one, which generates resistance before any value is shown. Choosing a low-impact service, or measuring activity (tickets closed) instead of outcomes (reliability, toil reduction).
LIKELY FOLLOW-UPS: How do you pick the first service objectively? What if engineers resist being measured by SLOs? Which single metric would you show the CEO first?
ONE CONCRETE EXAMPLE: You target the checkout service because outages there directly cost revenue. You add latency and error SLIs, set a 99.9% SLO, and cut alert noise by 70%. After ninety days you show leadership that MTTR dropped from hours to minutes and on-call pages fell sharply, earning a mandate to extend SRE to the next critical service.
Read the original → rootly.com
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.