What experiment artifacts and metrics do you present to decide shipping?

This tests structured experiment discipline. A strong answer covers the pre-approved design doc, scorecard results for primary goal and guardrail metrics plus secondary breakdowns versus the hypothesis, and duration context.
What's really being asked
The interviewer wants to know if you treat experimentation as a disciplined engineering process or as a casual dashboard exercise. Shipping decisions should be anchored in artifacts created before the experiment launched, not improvised during the review. The question probes your ability to separate signal from noise, protect the business from regressions, and drive a binary decision with pre-defined criteria.
The full answer
First, present the experiment design document itself, which should include the original description, screenshots or feature context, and the hypothesis stated in a confirmable way. Second, bring the scorecard organized by metric tiers: primary goal metrics that the experiment hoped to move, primary guardrail metrics that must not regress, and secondary metrics that reveal trade-offs or cannibalization effects. Third, show the power analysis and duration context so the team knows the readout hit the required duration and that the minimum detectable effect was reasonable. Fourth, give a crisp recommendation that maps each primary and secondary metric result back to the original hypothesis, explicitly calling out whether the evidence supports shipping, killing, or extending the test.
The mistakes people make
A major red flag is walking in with a grab bag of every metric that moved and asking the room what they think. Another is ignoring guardrail metrics entirely or dismissing a negative top-line result because a secondary metric looked good. Retrospectively changing the hypothesis to match what won is also a fatal error. Finally, presenting results before the pre-calculated duration has been reached shows you do not understand statistical maturity.
What usually comes next
The interviewer may ask how you would handle a situation where the primary metric is flat but a secondary metric is strongly positive, or how you would structure a follow-up holdout to measure long-term effects. They might also probe what you do when guardrail metrics regress slightly but goal metrics are strongly positive, or how you prioritize which metrics belong in the primary versus secondary bucket.
A concrete example
Imagine you ran a checkout flow redesign. In the review, you present the design doc that hypothesized a two percent lift in purchase conversion with no regression in refund rate. The scorecard shows a two point five percent lift in purchases with a flat refund rate, while secondary metrics reveal average order value dropped one percent. You note the experiment ran for the full fourteen days dictated by the power analysis, then recommend shipping with a post-launch holdout to monitor whether the lower average order value eventually cannibalizes lifetime revenue.
Interview question
Which combination of artifacts best demonstrates disciplined experimentation rather than a casual dashboard review when deciding to ship an experiment?
- a.A live dashboard of every metric that moved significantly, inviting the team to interpret trends and vote
- b.Only the primary goal metric, arguing that guardrails and secondary metrics slow down decisions
- c.Retrospective slides that reframe the strongest secondary result as the original primary hypothesis
- d.The pre-approved design doc, tiered scorecard, duration context, and a binary recommendation tied to the original hypothesisCorrect
Why? this is the answer
Disciplined shipping decisions must be anchored in pre-launch artifacts and pre-defined criteria, not improvised during the review. Option A represents the common red flag of bringing a grab bag of metrics and asking the room what they think, which fails to separate signal from noise.
Just read this? Test yourself on what you have been reading.
Read the original → statsig.com
- #experimentation
- #ab-testing
- #metrics
- #product-engineering
- #data-analysis
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Open roles that interview on experimentation — each one lists the topics its interview covers.
See open roles