A/B test p-value is 0.08, PM wants to ship. What now?
Tests if you can translate statistical risk into business terms for a PM. A good answer defines the 8% false positive risk, weighs it against the cost of shipping, and suggests next steps like running the test longer instead of just saying no.
WHAT THIS TESTS: This question tests your ability to translate a statistical concept (p-value) into concrete business risk for a non-technical stakeholder. It assesses your pragmatism and your capacity to be a data-informed partner to product, rather than a rigid implementer. The interviewer wants to see if you can defend statistical rigor while navigating a realistic business scenario that requires nuance and collaboration.
A GOOD ANSWER COVERS: First, acknowledge the PM's goal but gently reframe the conversation around risk. Second, explain the p-value in simple terms: "A p-value of 0.08 means there is an 8% chance we would see this result (or something more extreme) purely by random luck, even if our new checkout flow has zero real impact." This is the risk of a Type I error, or a false positive. Third, frame the decision as a cost-benefit analysis. Quantify the cost of being wrong (e.g., 50 engineering hours to ship and maintain a useless feature) against the potential upside of the change. Fourth, propose concrete next steps instead of a hard 'no'. The best option is usually to run the test longer to gather more data and increase statistical power, which could drive the p-value below the 0.05 threshold if the effect is real.
COMMON WRONG ANSWERS: One major red flag is being overly dogmatic: "The p-value is greater than 0.05, so the result is not significant and we cannot ship." This shows a lack of business sense. The opposite error is immediately agreeing with the PM based on the 'directionally correct' argument, which shows a lack of statistical rigor. The most critical mistake is misinterpreting the p-value, such as saying "There's a 92% chance the new version is better." This is fundamentally incorrect and a common misunderstanding of frequentist statistics.
LIKELY FOLLOW-UPS: An interviewer might ask, "When would it be acceptable to ship with a p-value of 0.08?" (A good answer: when the cost of implementation is near-zero and the consequences of being wrong are negligible). They might also ask, "What if the PM suggests we just change our significance level to 0.10?" (This is a form of p-hacking; the significance level must be set before the experiment begins).
ONE CONCRETE EXAMPLE: With a p-value of 0.08, there's an 8% chance, or roughly 1-in-12, that we are celebrating a false positive. If shipping this new checkout flow costs two engineers one week of work (80 hours), we are betting 80 hours of effort on a 1-in-12 shot that it's all for nothing. Is the potential conversion lift worth that risk? If we run the test for another week, we might get a clearer signal, reducing the risk of wasting that engineering investment.
Read the original → en.wikipedia.org
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.