Why is stopping an A/B test early problematic?

Tests understanding of the 'peeking problem' in A/B testing. A good answer defines peeking, explains how it inflates false positive rates, and contrasts it with waiting for a pre-determined sample size. A red flag is not explaining the statistical mechanism.
What's really being asked
This question tests your understanding of a critical A/B testing pitfall and your ability to guide non-technical stakeholders. The interviewer wants to see if you can identify the practice, name the statistical concept, explain the consequences in business terms (wasted effort, bad decisions), and articulate the correct process. It's a test of data literacy and technical leadership, not just statistical trivia.
The full answer
First, name the problem: This practice is called "peeking" or a form of p-hacking. Second, explain the statistical mechanism: In a standard (Frequentist) A/B test, statistical significance (e.g., p < 0.05) is calculated for a single point in time, after a pre-determined sample size is reached. Each time you "peek" at the results, you are running a new, independent statistical test. This dramatically increases the cumulative probability of a Type I error (a false positive). Third, quantify the risk: A test designed for a 5% false positive rate, if peeked at 10 times, can see its actual false positive rate inflate to over 20%. This means you'd incorrectly declare a winner in 1 out of 5 experiments. Fourth, state the solution: The correct approach is to use a sample size calculator before the test to determine the required number of users, and only call the experiment after that sample size is reached. A bonus is mentioning that alternative methods like Bayesian statistics are designed to handle continuous monitoring.
The mistakes people make
One red flag is simply saying "it's bad to stop early" without explaining the mechanism of inflated false positive rates. Another is confusing peeking with the multiple comparisons problem (testing 20 different metrics at once) or the Texas Sharpshooter Fallacy (forming a hypothesis after seeing a pattern). While related, peeking is specifically about repeatedly checking a single metric over time. A junior answer identifies the problem; a senior answer explains the 'why' and prescribes the correct process.
What usually comes next
How would you explain this to that product manager in a non-technical way? Are there any statistical frameworks where continuous monitoring is allowed? How do you calculate the required sample size for an experiment?
A concrete example
Imagine you're running a test to see if a new button color increases sign-ups. You aim for 95% confidence (a 5% chance of a false positive). If you check the dashboard once after collecting your target sample of 20,000 users, you have a 5% chance of being wrong if you see a significant result. However, if you check every day for 14 days, random noise will almost certainly make the p-value dip below 0.05 at some point. Stopping the test then means you acted on noise, not a true effect. The cumulative false positive rate can easily jump from 5% to over 25%, meaning you'd ship a useless or even harmful change one out of every four times you followed this peeking practice.
Interview question
A product manager wants to stop an A/B test early because the results are already statistically significant. What is the primary statistical risk of this 'peeking'?
- a.The test won't capture weekly user behavior patterns, leading to a sample that isn't representative.
- b.The small sample size means the observed effect is likely an overestimation that won't hold up at scale.
- c.Each check is a new statistical test, which cumulatively inflates the false positive rate above the intended level.Correct
- d.It introduces the multiple comparisons problem by analyzing too many different success metrics simultaneously.
Why? this is the answer
This is correct because each 'peek' is a new statistical test, and repeatedly testing increases the cumulative probability of a false positive (Type I error). While failing to capture weekly patterns is also a risk of short tests, it is a separate sampling issue, not the specific statistical error caused by peeking.
Just read this? Test yourself on what you have been reading.
Read the original → docs.growthbook.io
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Open roles that interview on a/b testing — each one lists the topics its interview covers.
See open roles