Skip to content
tezvyn:

Why not stop an A/B test when it looks significant early?

Source: docs.growthbook.ioMediumHow cards are made

Why not stop an A/B test when it looks significant early?

Tests whether you understand repeated looks inflate false positives. The term is peeking: checking daily can turn a 5% Type I error rate into roughly 15% by day 3. Red flag: citing "low sample size" without stating that early stopping invalidates the p-value.

What's really being asked

This question probes whether you understand that frequentist statistical inference depends on a fixed experimental protocol. In a standard t-test or z-test, the p-value is mathematically derived under the assumption that the sample size and stopping rule are determined before the experiment begins. If you violate that protocol by checking repeatedly and stopping conditionally, you are no longer playing the game the p-value was designed for.

The full answer

First, name the problem: it is called peeking, also known as optional stopping or interim monitoring without correction. Second, explain the mechanism: every time you look at the data, you give the random noise another chance to cross the significance threshold. With a nominal 5% alpha and daily checks, the actual false positive rate can exceed 15% within just a few days because the probabilities accumulate. Third, distinguish it from ordinary sample size concerns: even if the effect is real, early estimates are volatile and often regress to the mean as more users enter the experiment. Fourth, offer a solution: pre-commit to a runtime and sample size, use a proper power analysis, or adopt a sequential testing framework such as group sequential tests or always-valid p-values if the business truly needs early stopping.

The mistakes people make

Saying "the sample is too small" without addressing why the p-value itself is invalidated. Calling it the multiple testing problem; while related, that term usually refers to testing many metrics or subgroups, not repeated temporal looks at the same metric. Suggesting to lower the alpha threshold ad hoc without a formal correction method like O'Brien-Fleming or spending functions. Agreeing to stop early because "the PM needs an answer fast," which signals you will trade statistical rigor for speed.

What usually comes next

How would you design a test if leadership insists on daily readouts? What is the difference between peeking under frequentist versus Bayesian frameworks? Can you explain why a 5% daily alpha does not equal a 5% experiment-wise alpha? How do sequential testing or always-valid p-values solve this mathematically?

A concrete example

Imagine you run a 30-day experiment with a 5% significance level and you peek every day. The probability of at least one false positive across those 30 looks is roughly 1 minus 0.95 to the 30th power, which is about 79%. Even though each individual day carries a 5% risk, the compound risk is catastrophic. In practice, with correlated daily observations the exact inflation is slightly less, but the directional problem remains severe: by day 3 you might easily have a 15% to 20% chance of declaring a winner when there is no true effect.

Interview question

You run a 30-day A/B test and peek daily, stopping early if p < 0.05. What is the principal statistical issue?

  • a.Statistical power drops because the test was stopped before the planned sample size was reached
  • b.The overall false positive rate inflates because each look gives random noise another chance to appear significantCorrect
  • c.The Type I error rate is unaffected because each daily look still uses the same 5% alpha level
  • d.Early users differ systematically from later users, creating a selection bias that exaggerates the treatment effect
Why?

Peeking gives random noise multiple chances to cross the significance threshold, compounding the false positive rate far above the nominal 5% level. Distractor D is tempting but wrong because a 5% alpha is only valid for a single pre-specified analysis, not for repeated uncorrected looks.

Just read this? Test yourself on what you have been reading.

Read the original → docs.growthbook.io

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on ab-testing — each one lists the topics its interview covers.

See open roles