Skip to content
tezvyn:

How do you determine sample size and duration for an A/B test?

Source: cxl.comMediumHow cards are made

How do you determine sample size and duration for an A/B test?

This tests statistical power literacy. A strong answer names baseline rate, MDE, alpha, and beta; explains the duration versus sensitivity trade-off; and notes traffic allocation. A red flag is ignoring power or stopping early when results look significant.

What's really being asked

This question probes whether you understand that an A/B test is a controlled experiment, not a dashboard you watch until the numbers feel good. Interviewers want to see that you treat sample size and duration as derived quantities rooted in statistical power, not arbitrary calendar events. They also care whether you recognize real world constraints like traffic volume, business cycles, and the cost of wrong decisions.

The full answer

First, name the four inputs to a standard power analysis. These are the baseline conversion rate or mean, the Minimum Detectable Effect you want to be able to spot, the significance level alpha usually set at five percent, and statistical power typically eighty percent. Second, explain that sample size per variant is what the formula yields, and duration equals total required visitors divided by daily eligible traffic. Third, discuss the core trade off. A smaller MDE or higher power inflates sample size and lengthens the test, while a larger MDE shrinks it but raises the risk of missing small yet valuable gains. Fourth, bring in practical guardrails. Mention running for full business cycles to capture weekday versus weekend behavior, avoiding peeking or optional stopping, and reserving a holdout if the business cannot afford a fifty-fifty split.

The mistakes people make

The biggest red flag is saying you run the test for a week or two because that feels standard. Another is ignoring power entirely and claiming a large sample size alone guarantees validity. Some candidates confuse MDE with the observed lift, or suggest lowering alpha after the test starts to salvage significance. Stopping early because the p-value crossed the threshold mid experiment is a severe error that destroys the false positive rate.

What usually comes next

An interviewer might ask how you would handle low traffic, which pushes you toward larger MDEs, sequential testing, or Bayesian methods with acceptable error bounds. They might also ask what you do if the test reaches the planned duration but the p-value is exactly point zero six, which tests whether you understand that the protocol should have been locked beforehand.

A concrete example

Suppose your checkout page converts at ten percent baseline and you want to detect a relative fifteen percent uplift, which is an absolute one point five percentage point lift. Using a two tailed test at alpha equals point zero five and power of point eight, you need roughly six thousand two hundred visitors per variant, or about twelve thousand four hundred total. If your site sees one thousand eligible visitors per day and you run a fifty-fifty split, you will need roughly twenty five days. If the business demands results in two weeks, you must either accept a higher MDE, lower power, or run on a higher traffic page.

Interview question

Your power analysis says an A/B test needs 25 days, but the business demands results in 14. What is the valid statistical response?

  • a.Keep the original MDE and power targets but run for 14 days, because daily traffic is high enough to ensure validity.
  • b.Stop at 14 days if the p-value is significant, since early significance still reflects a true effect.
  • c.Accept a larger minimum detectable effect or lower power to fit the shorter window.Correct
  • d.Replace the planned MDE with the observed lift from the first week to justify stopping at 14 days.
Why?

The card explains that when business constraints shorten the available window, you must adjust the statistical inputs and accept a larger MDE or lower power, because duration is derived from those inputs. Stopping early simply because the p-value crosses the threshold mid-experiment is a severe error that destroys the false positive rate.

Just read this? Test yourself on what you have been reading.

Read the original → cxl.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Every open role lists the topics its interview covers, so you can prepare for the real thing rather than guessing.

See open roles