Skip to content
tezvyn:

Explain statistical power and respond to extending a null A/B test

Source: cxl.comHardHow cards are made

Explain statistical power and respond to extending a null A/B test

This tests statistical power and p-hacking judgment. A strong answer defines power as detecting a true effect, rejects extending the test to chase significance, and requires pre-registered sample sizes. Agreeing to run until it hits significance is a red flag.

What's really being asked

The interviewer wants to know if you understand that statistical power is not a post-hoc dial you can turn but a design-time property, and whether you have the statistical maturity to push back on p-hacking requests from stakeholders. They care about your ability to balance business pressure with methodological rigor.

The full answer

First, define statistical power clearly as the probability that the test correctly rejects the null hypothesis when a specific alternative is true, typically expressed as one minus beta. Second, explain that power is determined before data collection by three levers: sample size, significance level alpha, and the minimum detectable effect or MDE. Third, tell the stakeholder that extending the test after observing a null result is a form of optional stopping that destroys the validity of the p-value by inflating the Type I error rate, so you cannot simply run until significance appears. Fourth, describe the constructive path forward: report the observed effect size and its confidence interval, assess whether the null result is practically meaningful, and if the feature is still strategically important, design a new adequately powered study rather than recycling the current one.

The mistakes people make

A major red flag is agreeing to extend the test or suggesting that more data will eventually reveal the truth. Another is confusing statistical power with sample size alone, ignoring effect size and alpha. Some candidates incorrectly believe that a post-hoc power calculation justifies continuing the experiment; in reality, observed power is a function of the observed p-value and is not a valid reason to keep running. Finally, framing the issue as purely statistical without translating it into business risk, such as wasted engineering time or polluted decision-making, signals weak stakeholder management.

What usually comes next

The interviewer may ask how you would calculate power for a given MDE, how sequential testing or group sequential methods change the stopping rule, or what you would do if the stakeholder insists and has executive authority. They might also probe whether Bayesian approaches or ROPE-style decision rules could be viable alternatives in this context.

A concrete example

Suppose you ran a checkout flow A/B test for two weeks with ten thousand users per variant and observed a conversion lift of one percent with a p-value of 0.2. The stakeholder wants to run it for another two weeks. You respond by showing that the MDE powered at eighty percent was two percent, so the study was never designed to detect a one percent lift reliably. You explain that continuing would give a roughly twenty to thirty percent false positive rate under optional stopping, not the nominal five percent. You then propose reporting the confidence interval of minus 0.5 percent to plus 2.5 percent and running a new test with twenty thousand per variant if the business still believes a one percent lift is worth detecting.

Interview question

When a stakeholder asks to extend a null A/B test to chase significance, which response best balances statistical rigor with business stakeholder management?

  • a.Explain that optional stopping inflates the false positive rate, report the confidence interval, and design a new adequately powered study if the feature remains strategic.Correct
  • b.Calculate post-hoc power from the observed effect size to determine how many additional weeks are required to reach 80% power.
  • c.Agree to extend the test while lowering the alpha threshold to 0.01 to compensate for the larger total sample size.
  • d.Refuse the request on methodological grounds and decline to translate the null result into any business risk assessment.
Why?

B is correct because optional stopping invalidates the p-value by inflating Type I error, so the valid path is to report the confidence interval and design a new pre-registered study. A represents the post-hoc power fallacy: observed power is mathematically determined by the p-value and provides no independent justification for collecting more data.

Just read this? Test yourself on what you have been reading.

Read the original → cxl.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on statistics — each one lists the topics its interview covers.

See open roles