tezvyn:

Calculate Sample Size for a 2% A/B Test Lift

AI-drafted, machine-checkedSource: kissmetrics.iointermediate

This tests if you connect statistical inputs to business goals. A good answer defines baseline rate, minimum detectable effect (MDE), and power, then explains MDE as a cost/benefit trade-off.

WHAT THIS TESTS: This question tests your ability to move beyond just running a test to designing a valid one. It checks if you understand that sample size isn't a magic number but a result of conscious business and statistical trade-offs. The interviewer wants to see if you can connect engineering work to business impact and risk management, specifically avoiding false positives and false negatives.

A GOOD ANSWER COVERS: A good answer hits three points in order. First, it names the primary inputs for a sample size calculator: the baseline conversion rate, the minimum detectable effect (MDE), and the desired statistical power and significance levels. Second, it explains the role of each: baseline sets the current performance, MDE is the smallest change you care about (the 2% lift), and power/significance control your risk of false negatives (Type II error) and false positives (Type I error). Third, it frames the MDE as the key business decision, explaining the trade-off: a smaller MDE requires a much larger sample size, increasing test duration or cost, which must be justified by the potential business value of that smaller lift.

COMMON WRONG ANSWERS: A major red flag is treating these parameters as purely statistical inputs without connecting them to business decisions. For example, saying "we need to detect a 2% lift" without explaining the business trade-off is a weak answer. Another red flag is forgetting one of the key inputs, most commonly statistical power (or its counterpart, the risk of a Type II error). Simply saying "we need more traffic" is too vague and misses the point of the question, which is about the parameters that determine how much traffic is needed.

LIKELY FOLLOW-UPS: Expect questions like: "How would you handle this calculation if the site has very low traffic?" or "What happens if you stop the test early because the results look good?" or "How does the baseline conversion rate affect the required sample size?"

ONE CONCRETE EXAMPLE: To detect a 2% relative lift on a baseline conversion rate of 3.2%, you must also define power and significance. Using the standard 80% power and 95% significance (alpha=0.05), you input these values into a calculator. If the result is 500,000 users per variant and you only get 50,000 unique visitors a week, the business must decide if a 10-week test is worth the potential 2% gain. This makes the trade-off between statistical certainty and business velocity explicit.

Read the original → kissmetrics.io

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.