How do you determine sample size and duration for an A/B test?

This tests your grasp of statistical power and business trade-offs. A good answer defines the four inputs (baseline, MDE, significance, power) to calculate sample size, then uses traffic to find duration.
WHAT THIS TESTS: This question assesses your ability to design statistically valid experiments that are also practical for the business. It's not a pure statistics quiz; it's about demonstrating you understand the trade-offs between speed, risk (false positives/negatives), and the cost of detecting an effect. The interviewer wants to see if you can move beyond simple rules of thumb (like "run tests for two weeks") to a more rigorous, data-driven approach.
A GOOD ANSWER COVERS: A strong answer follows a clear, logical sequence. First, state that the goal is to determine the sample size needed to achieve statistical power, and duration is a function of that sample size and traffic. Second, define the four key inputs for a power calculation: 1) Baseline Conversion Rate (BCR), the existing performance of the control; 2) Minimum Detectable Effect (MDE), the smallest relative lift that the business cares about, which is a business decision, not a statistical one; 3) Statistical Significance Level (alpha), the tolerance for a false positive (Type I error), typically set at 95% (p < 0.05); and 4) Statistical Power (1-beta), the probability of detecting a true effect (avoiding a false negative or Type II error), typically set at 80%. Third, explain that these inputs are used in a sample size calculator. Fourth, clarify that you divide the required sample size per variation by the average daily traffic to that page to estimate the duration, often rounding up to complete full business cycles (e.g., full weeks) to account for user behavior variance.
COMMON WRONG ANSWERS: The most common mistake is suggesting a fixed duration, like "I'd run it for two weeks." This completely ignores the sample size required for statistical validity. A low-traffic page might need months to reach significance. Another red flag is "peeking" at results and stopping the test as soon as it hits significance, which dramatically increases the false positive rate. Setting an unrealistically small MDE (e.g., 1%) without acknowledging the massive sample size it would require is another sign of inexperience. A senior candidate should be able to articulate why a 1% MDE might require a test to run for an entire year and is therefore impractical.
LIKELY FOLLOW-UPS: Expect questions that test your ability to handle real-world constraints. For example: "What if your calculation says the test needs four weeks, but the PM wants results in one?" (This tests your ability to negotiate trade-offs: you can increase the MDE, decrease the power, or accept a higher alpha). Another is, "How do you decide on the MDE?" (The answer should revolve around business impact and ROI—what is the smallest lift that justifies the engineering cost and potential risk of the change?).
ONE CONCRETE EXAMPLE: Let's say we have a checkout page with a 10% baseline conversion rate (BCR). The business decides the minimum lift worth pursuing is a 5% relative improvement, so our MDE is 5%. This means we want to be able to detect a new conversion rate of 10.5%. We use standard levels for significance (95%) and power (80%). Plugging these into a power calculator gives a required sample size of approximately 31,000 users per variation. If the page gets 4,000 users per day, the total required sample of 62,000 would take 15.5 days to collect. Therefore, you would recommend running the test for three full weeks (21 days) to ensure you gather enough data and capture three full weekly business cycles.
Read the original → cxl.com
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.