Skip to content
tezvyn:

Why is stopping an A/B test when it hits significance problematic?

Source: docs.growthbook.ioMediumHow cards are made

Why is stopping an A/B test when it hits significance problematic?

Tests your understanding of the 'peeking problem' in A/B testing. A great answer defines peeking, explains how it inflates the Type I error rate (false positives), and states the need for a predetermined sample size.

What's really being asked

This question tests your knowledge of a common and critical A/B testing pitfall. The interviewer is checking if you can identify the practice by its statistical name, 'peeking,' and explain the mathematical reason why it's invalid. It separates candidates who have a superficial understanding of statistical significance from those who grasp the mechanics of frequentist hypothesis testing and its practical application.

The full answer

An excellent answer identifies the issue and explains the consequences in four steps. First, name the problem: this is called 'peeking' or the 'peeking problem.' Second, explain the statistical impact: every time you check the dashboard, you are effectively running a new hypothesis test. This dramatically increases the cumulative probability of a Type I error (a false positive). Third, quantify the effect: with a standard 95% confidence level (p-value < 0.05), the false positive rate is 5% for a single check. Peeking just 10 times can inflate this rate to over 15%. Fourth, state the correct procedure: determine the required sample size before the experiment begins and only analyze the results once that sample size is reached.

The mistakes people make

Vague explanations like 'the results aren't stable yet' or 'you need to wait for the data to mature' are red flags. While true, they don't demonstrate an understanding of the underlying statistical principle. Another common error is confusing peeking with the 'multiple comparisons problem,' which refers to testing many different metrics at once, not one metric many times. A senior candidate is expected to know the term 'peeking' and distinguish between these concepts. Finally, suggesting an invalid fix, like waiting a few more days after significance is hit, shows a misunderstanding of the core problem.

What usually comes next

Expect follow-ups like: 'How would you explain this concept to a non-technical PM?' to test communication skills. Or, 'Are there any statistical methods that do allow for continuous monitoring?' which opens a door to discuss sequential testing or Bayesian experimentation. They might also ask, 'How do you calculate the required sample size?' to test your foundational knowledge.

A concrete example

Imagine you run a test with a 95% confidence threshold (alpha = 0.05). This means you accept a 5% chance of a false positive if you check the results only once at the end. However, if you peek at the results daily for 20 days, the probability of seeing a p-value under 0.05 at least once due to random chance alone can skyrocket to over 20%. This means that in 1 out of 5 experiments where this peeking occurs, you would incorrectly declare a winner and ship a feature with no actual positive effect, wasting engineering resources.

Interview question

What is the primary statistical risk associated with "peeking" at A/B test results and stopping the experiment as soon as significance is observed?

  • a.It inflates the cumulative probability of a Type I error, leading to an increased rate of false positives.Correct
  • b.It creates a multiple comparisons problem by testing the same metric repeatedly over time.
  • c.It makes the observed effect size unstable and difficult to interpret accurately.
  • d.It primarily increases the chance of a Type II error, meaning a real effect might be missed.
Why?

Peeking at results multiple times effectively runs a new hypothesis test each time, dramatically increasing the cumulative probability of a Type I error (false positive). While repeated checks might feel like a 'multiple comparisons problem,' that term specifically refers to testing many different metrics simultaneously, not one metric many times.

Just read this? Test yourself on what you have been reading.

Read the original → docs.growthbook.io

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on a/b testing — each one lists the topics its interview covers.

See open roles