Explain statistical significance in copy A/B tests and why one day fails.

This checks if you distinguish signal from noise. A strong answer defines statistical significance as confidence a difference is real, warns that one-day samples are small and skewed by variance, and cites false positive risk.
What's really being asked
The interviewer wants to know if you understand that A/B testing is a randomized experiment governed by statistical hypothesis testing, not a simple comparison of raw numbers. They are checking whether you can separate true signal from random noise, and whether you appreciate why sample size and time matter for valid inference.
The full answer
First, define statistical significance as the probability that an observed difference in conversion rates between variant A and variant B reflects a genuine effect rather than random fluctuation. Second, explain that a one-day sample is usually too small to achieve adequate statistical power, and that daily traffic patterns, day-of-week effects, and external events can temporarily skew results. Third, mention that proper A/B tests require a predetermined sample size or run time, a chosen significance level such as 0.05, and ideally a power analysis to control false negatives. Fourth, note that stopping early and picking the leader inflates the false positive rate because random early leads regress toward the mean as more data arrives.
The mistakes people make
Claiming that the higher conversion rate is automatically the winner regardless of sample size or variance. Saying that statistical significance is just a fancy way to say the numbers look big. Arguing that copy tests do not need rigor because words are cheap to change. Proposing to run the test until one version wins without accounting for peeking or multiple comparison problems.
What usually comes next
How would you calculate the required sample size before launching the test? What would you do if the test reaches significance but the absolute lift is tiny? How do you handle seasonality or traffic spikes during a test? Can you explain the difference between statistical significance and practical significance?
A concrete example
Imagine variant A has a 2.0 percent conversion rate and variant B has a 2.4 percent conversion rate after one day with only 200 visitors per variant. The 0.4 percentage point gap looks promising, but with such a small sample the confidence interval likely spans zero and the result is not statistically significant. If you ship B based on that single day, you might be chasing noise. After running for two weeks to reach ten thousand visitors per variant, the same 0.4 percentage point gap might hold with a p-value below 0.05, giving you confidence the copy change actually moved the needle.
Interview question
Variant B leads variant A by 0.4 percentage points after one day in a copy A/B test. What is the main reason to keep running before declaring a winner?
- a.You can stop early as soon as the dashboard shows one variant ahead by a visible margin
- b.A 0.4 percentage point lift is too small to have meaningful business impact
- c.Copy changes are low risk, so there is no harm in shipping the current leader after just one day
- d.The small sample makes it impossible to tell whether the gap reflects a true difference or random noiseCorrect
Why? this is the answer
A one-day sample usually lacks the statistical power to distinguish a genuine copy effect from random fluctuation, so the confidence interval likely spans zero and the result is not statistically significant. Option B is tempting because it sounds like cautious data analysis, but it confuses statistical significance with practical significance—the problem is not that the lift is too small to matter, but that we cannot yet trust the difference is real.
Just read this? Test yourself on what you have been reading.
Read the original → en.wikipedia.org
- #ab-testing
- #statistics
- #conversion-optimization
- #hypothesis-testing
- #experimentation
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Open roles that interview on ab-testing — each one lists the topics its interview covers.
See open roles