Skip to content
tezvyn:

Primary vs. Guardrail Metrics in Experiments

Source: statsig.comMediumHow cards are made

Primary vs. Guardrail Metrics in Experiments

Tests your grasp of risk management in A/B testing. A great answer defines a primary metric as the goal and a guardrail as a 'do no harm' check. A feature ships only if the primary improves without hurting guardrails.

What's really being asked

This question tests your understanding of risk management and product sense within experimentation. The interviewer wants to see if you think about the holistic impact of a change, not just its narrow, intended effect. It separates candidates who just optimize a number from those who understand how to protect the overall business and user experience.

The full answer

First, define a primary metric as the single, pre-defined measure of success that directly evaluates the experiment's hypothesis. It must be sensitive enough to detect the change. Second, define guardrail metrics as key health indicators that you monitor to ensure they do not get significantly worse. Their purpose is to prevent unintended harm. Third, explain the decision-making framework: a feature should only ship if the primary metric shows a statistically significant improvement AND no guardrail metrics show a statistically significant regression. A guardrail regression is a hard stop.

The mistakes people make

A major red flag is confusing guardrail metrics with secondary metrics. Secondary metrics are other positive outcomes you hope to see (a bonus), while guardrail metrics are negative outcomes you must prevent (a safety net). Another wrong answer is treating a guardrail regression as an acceptable trade-off. For senior roles, a significant guardrail regression should be treated as a blocker that requires investigation, not a simple business decision to ignore. Finally, failing to state that all metrics must be chosen before the experiment begins suggests a lack of scientific rigor.

What usually comes next

Expect follow-ups like: "How do you choose which guardrail metrics to monitor?" (Answer should cover core business KPIs, key funnel steps, and system performance like latency or error rates). Or, "What if a guardrail metric regresses but the change is not statistically significant?" (Discuss statistical power, practical significance, and making a product judgment call).

A concrete example

Let's test a new, more complex recommendation algorithm on an e-commerce homepage. The hypothesis is that it will increase product discovery. Primary Metric: Click-through rate (CTR) on recommended products. Critical Guardrail Metrics: Page load time, add-to-cart rate, and overall revenue per session. Ship Decision Scenario: The experiment shows a 10% statistically significant lift in CTR. However, it also causes a 150ms statistically significant increase in page load time. Decision: Do not ship. The engagement gain from better recommendations is not worth the user frustration and potential SEO penalty from a slower site. The page load time regression is a critical failure that stops the launch.

Interview question

An A/B test shows a significant increase in the primary metric, but also a statistically significant regression in a key guardrail metric. What is the correct course of action?

  • a.Rerun the experiment with a larger sample size to get a more accurate measurement of the guardrail metric's impact.
  • b.Ship the feature, as the improvement in the primary metric is the main goal and outweighs the guardrail regression.
  • c.Do not ship the feature. Investigate the cause of the guardrail regression, as it indicates unintended harm.Correct
  • d.Launch the feature, but re-classify the guardrail metric as a secondary metric in the final report.
Why?

The correct action is to not ship. A statistically significant regression in a guardrail metric is a 'hard stop' designed to prevent unintended harm. The most tempting mistake is to treat it as an acceptable trade-off (B), which defeats the purpose of having a guardrail.

Just read this? Test yourself on what you have been reading.

Read the original → statsig.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on analytics — each one lists the topics its interview covers.

See open roles