tezvyn:

Determine Sample Size for a 2% Lift A/B Test

AI-drafted, machine-checkedSource: kissmetrics.iointermediate

This tests your grasp of statistical power and the business trade-offs in experimentation. A great answer defines baseline conversion rate, minimum detectable effect (MDE), and statistical power. A red flag is ignoring the business context of MDE.

WHAT THIS TESTS: Your ability to move beyond basic A/B testing definitions and articulate the practical trade-offs between statistical rigor and business velocity. The interviewer wants to see if you can translate a business goal (a 2% lift) into the statistical inputs needed for a valid experiment, and explain the costs and risks associated with those choices.

A GOOD ANSWER COVERS: A great answer identifies three, and often four, key parameters. First, the Baseline Conversion Rate (BCR), which is the current performance of the control group, ideally based on at least 30 days of data (e.g., 3.2%). Second, the Minimum Detectable Effect (MDE), which is the smallest lift you care about detecting. For a 2% lift, the MDE is 2%. This is a business decision, not a statistical one. Third, Statistical Significance Level (alpha), the probability of a false positive (Type I error), usually set at 5%. Fourth, Statistical Power (1-beta), the probability of detecting a real effect and avoiding a false negative (Type II error), typically set at 80%. The relationship is that a smaller MDE, lower alpha, or higher power all require a larger sample size.

COMMON WRONG ANSWERS: A major red flag is suggesting a fixed time duration, like "run the test for two weeks." This ignores the statistical requirements and leads to underpowered tests. Another common mistake is treating the MDE as a purely statistical value rather than a business decision based on the expected return versus the cost of implementation. For example, is a 2% lift worth the engineering effort? Candidates also often forget to mention power or only mention significance, showing an incomplete understanding of experimental design. Naming only one or two of the parameters is a sign of junior-level knowledge.

LIKELY FOLLOW-UPS: "How would you handle this if you don't have enough traffic to detect a 2% lift in a reasonable timeframe?" (Answer: Increase the MDE, run the test longer, or use alternative methods like sequential testing). "What happens if you 'peek' at the results before the required sample size is reached?" (Answer: It dramatically inflates the false positive rate, leading you to ship changes that have no real effect).

ONE CONCRETE EXAMPLE: To detect a 2% relative lift on a 3.2% baseline conversion rate (lifting it to 3.264%), with standard 95% significance (alpha=0.05) and 80% power, you would need approximately 393,000 users per variation. If you only get 10,000 users per day to this page, this test would need to run for about 40 days per variation. This illustrates the trade-off: detecting small effects requires a massive sample size and significant time investment.

Read the original → kissmetrics.io

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.