p-value is 0.08, significance is 0.05. Ship it?
This tests your ability to translate statistical risk for business partners. Explain that p=0.08 means an 8% chance of a false positive, quantify the cost of a bad decision, and suggest extending the test to increase power.
WHAT THIS TESTS: This isn't a stats quiz. It tests your ability to hold your ground on data integrity, communicate complex statistical concepts to a non-technical stakeholder (the PM), and collaboratively find a business-aware solution. The interviewer wants to see if you can balance rigor with pragmatism.
A GOOD ANSWER COVERS: A good answer hits four points. First, translate the p-value: "A p-value of 0.08 means there's an 8% chance we'd see this result, or a more extreme one, even if the new checkout flow has zero actual effect on conversion." Second, state the risk clearly: We set our significance level at 0.05, meaning we were only willing to accept a 5% chance of a false positive (a Type I error). Shipping now means accepting a higher 8% risk. Third, frame the business trade-off: "Is the potential upside of this feature worth taking an 8% risk that it's useless and we've wasted engineering effort?" Fourth, propose a solution: "The most statistically sound action is to extend the test. We need more data to increase the test's power and see if the p-value drops below 0.05. Let's calculate how many more users or days we need."
COMMON WRONG ANSWERS: A major red flag is immediately agreeing with the PM. Saying "it's directionally correct" shows you don't appreciate the risk of a false positive. Another mistake is being overly rigid and just saying "No, the p-value is > 0.05" without explaining why or offering a path forward. This shows a lack of business partnership. A subtle error is confusing p-value with the probability of the hypothesis being true; it's the probability of the data given the null hypothesis is true.
LIKELY FOLLOW-UPS: "How long should we extend the test?" (This requires a power calculation). "What if the PM says we can't extend the test because of a hard deadline?" (This pushes you into a risk/reward discussion: what's the cost of shipping a dud vs. the cost of missing the deadline?). "What if the change was very cheap to build and is easy to revert? Does that change your answer?" (Yes, it lowers the cost of being wrong, which might make an 8% risk acceptable).
ONE CONCRETE EXAMPLE: "Let's say launching this adds 0.5 FTE of maintenance cost per year, about 150k. By shipping with a p=0.08, we are accepting an 8% chance we are taking on that 150k cost for zero gain. If we run the test for three more days, we can get that risk down to our agreed-upon 5% threshold. This seems like a worthwhile trade-off to avoid a costly mistake."
Read the original → en.wikipedia.org
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.