How do you determine sample size for a conversion lift experiment?
Tests fluency with statistical experiment design. A strong answer frames N as a function of alpha, power, baseline rate, and MDE, noting that shrinking the MDE or raising power inflates N. Red flag: picking N from traffic instead of risk tolerance.
WHAT THIS TESTS: This question probes whether you can design a valid frequentist experiment from first principles rather than relying on a calculator. The interviewer wants to see that you understand power as the probability of detecting an effect given that some prespecified effect actually exists, and that you can articulate how business constraints map to statistical parameters.
A GOOD ANSWER COVERS: First, name the four levers that determine sample size for a two-proportion test: the significance level alpha, the desired power one minus beta, the baseline conversion rate, and the minimum detectable effect. Second, explain the directional relationships: increasing power or shrinking the MDE both increase N, and for conversion rates the relationship is roughly inverse-square with the MDE. Third, discuss the business tradeoffs: alpha controls the false positive rate, power controls the false negative rate, and the MDE should be anchored to the smallest lift that is practically worth launching. Fourth, acknowledge real-world constraints, such as using a sequential testing framework or accepting a higher MDE if traffic is limited, rather than pretending the sample size is arbitrary.
COMMON WRONG ANSWERS: A major red flag is treating sample size as a fixed input based on available traffic rather than a derived output based on risk tolerance. Another is conflating statistical significance with practical significance, for example claiming a 0.1 percent lift is meaningful just because the p-value is below 0.05. Recommending that you simply run the experiment until you reach significance is a serious error because it invalidates the Type I error rate through optional stopping. Finally, stating that power equals one minus alpha reveals a fundamental misunderstanding of the distinction between false positives and false negatives.
LIKELY FOLLOW-UPS: The interviewer may ask how you would handle a low-traffic product, which tests whether you know about variance reduction techniques like CUPED or stratification. They might ask about sequential testing or group sequential designs that allow early stopping without inflating alpha. Another common follow-up is how you would choose the MDE in practice, which should tie to the minimum ROI needed to justify engineering and operational costs. You might also be asked to sketch the formula or explain why the standard error of a proportion depends on p times one minus p.
ONE CONCRETE EXAMPLE: Suppose your baseline conversion rate is 10 percent and you want to detect a relative 5 percent lift, meaning an absolute MDE of 0.5 percentage points. Using a two-sided test with alpha at 0.05 and power at 0.80, the required sample size per variant is roughly sixty-three thousand users. If the business insists on detecting a 2 percent relative lift instead, the absolute MDE drops to 0.2 percentage points and the required sample size balloons to roughly three hundred ninety thousand per variant, illustrating the quadratic cost of precision. If you only have one hundred thousand total users, you must either accept lower power, raise alpha, or adopt a variance reduction technique.
Read the original → en.wikipedia.org
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.