Primary vs. Guardrail Metrics in Experiments

This tests if you can balance improving a key metric with not harming the user experience. Define primary (the goal) and guardrail (don't harm) metrics. Give an example where a guardrail regression (e.g., latency) blocks a feature ship.
What's really being asked
This question assesses your practical experience with A/B testing and product sense. It's not about textbook definitions. The interviewer wants to see if you understand that optimizing for a single metric can have negative second-order effects. They are testing your ability to manage risk, define a holistic view of success, and make a sound business decision even when the primary success metric looks good. It separates candidates who just look at a dashboard from those who understand the system and user experience as a whole.
The full answer
First, clearly define a primary metric as the single, pre-registered metric an experiment is designed to move, directly tied to the hypothesis. It's the reason for the test. Second, define guardrail metrics as critical health indicators that you must not harm. These are your ship-blockers, such as latency, error rates, or user churn. They are not expected to improve. Third, provide a concrete example where a statistically significant win on a primary metric is correctly vetoed by a statistically significant regression in a guardrail metric, explaining the business rationale for the 'no-ship' decision.
The mistakes people make
A major red flag is confusing guardrail metrics with secondary metrics. Secondary metrics are 'nice-to-haves' that you hope might also improve, but they aren't ship-blockers. Another mistake is treating a guardrail regression as just another data point to 'consider' rather than a hard stop. A weak answer gives a generic example without numbers or fails to explain why the guardrail regression is unacceptable. For instance, saying 'latency went up' is weaker than 'P95 latency increased by 300ms, violating our 100ms regression budget for this page.'
What usually comes next
Expect questions like: 'How do you decide which guardrail metrics to choose for a given experiment?', 'What do you do if a guardrail metric shows a negative trend but it's not statistically significant?', or 'How would you handle an experiment with two competing primary metrics?'.
A concrete example
Let's say we're testing a new, personalized recommendation algorithm on an e-commerce homepage. The hypothesis is that better recommendations will increase engagement. The PRIMARY METRIC is 'click-through rate (CTR) on the recommendation module'. The GUARDRAIL METRICS are 'P95 page load time', 'session-level conversion rate', and 'support ticket volume'. The experiment runs and shows a 20% lift in CTR, a huge win. However, the guardrail metrics show that P95 page load time increased by 500ms and the overall session-level conversion rate dropped by 3%. The CTR win is negated by the terrible performance impact and the drop in actual purchases. The guardrail metrics prove the feature is harmful overall. The decision is 'no-ship' and to investigate the performance of the new algorithm.
Interview question
An A/B test shows a significant lift in the primary metric but also a statistically significant regression in a critical guardrail metric. What is the recommended action?
- a.Analyze the trade-off between the primary metric gain and the guardrail regression to determine if the overall business impact is positive.
- b.Ship the feature, as the primary metric's success outweighs the guardrail's negative impact, and plan to fix the guardrail issue later.
- c.Treat the guardrail regression as a secondary metric concern and only block the ship if other key secondary metrics also show negative trends.
- d.Do not ship the feature; investigate the guardrail regression to understand its cause and resolve it before considering deployment.Correct
Why? this is the answer
Guardrail metrics are defined as critical health indicators that are 'ship-blockers' if they show a statistically significant regression, even if the primary metric is a win. Therefore, the feature should not be shipped until the guardrail issue is resolved. Option C is incorrect because it confuses guardrail metrics with secondary metrics, which the card explicitly states is a common misconception and a 'major red flag'.
Just read this? Test yourself on what you have been reading.
Read the original → statsig.com
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Open roles that interview on analytics — each one lists the topics its interview covers.
See open roles