A/B test p-value 0.08, PM wants to ship. How do you advise?
Tests statistical rigor versus business pragmatism. A strong answer covers pre-registered thresholds, false positive risk, statistical power, confidence intervals, and the business cost of being wrong. Red flag: shipping without quantifying downside risk.
WHAT THIS TESTS: This question tests whether you understand that a p-value is a measure of consistency with the null hypothesis under the assumed model, not the probability that the variant is better. It also tests whether you can defend statistical protocol without being a rigid obstructionist, and whether you understand the practical difference between statistical significance and business significance.
A GOOD ANSWER COVERS: First, explain that the 0.05 threshold was chosen before the experiment to control the Type I error rate, and abandoning it after seeing the data is p-hacking that inflates the false positive rate. Second, interpret the p-value precisely: if the null hypothesis were true, there is an 8% chance of observing a result at least this extreme, not an 8% chance the null is true. Third, discuss statistical power and sample size, noting that a larger sample might clarify whether the effect is real or noise. Fourth, examine the confidence interval for the effect size to see if the plausible range includes zero or negative values. Fifth, weigh the business cost of shipping a null or negative change against the opportunity cost of waiting, and propose concrete next steps such as running a larger test or targeting a high-risk segment first.
COMMON WRONG ANSWERS: Saying the result is directionally correct so it is safe to ship ignores that the observed lift may be entirely due to random variation. Arguing that 0.08 is close enough to 0.05 misunderstands frequentist inference; the threshold is arbitrary but must be pre-specified to preserve the error rate. Conversely, refusing to ship anything that misses the threshold without discussing power, sample size, or business impact shows a lack of practical judgment. Conflating the p-value with the probability that the variant is better is a fundamental misinterpretation.
LIKELY FOLLOW-UPS: The interviewer may ask how you would design the next experiment to detect a smaller effect size, or what sample size would be needed given a baseline conversion rate and minimum detectable effect. They may also ask how you would handle a situation where the business cost of a false negative is higher than the cost of a false positive, or how Bayesian methods would change your recommendation.
ONE CONCRETE EXAMPLE: Suppose the baseline checkout conversion is 10% and the observed lift is 2% with p=0.08. The confidence interval might span from minus 0.3% to plus 4.3%. If the new flow requires significant engineering maintenance, shipping means accepting a real chance of harming conversion. If the change is free to maintain, a small-scale rollout with close monitoring could be a reasonable compromise.
Read the original → en.wikipedia.org
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.