tezvyn:

Explain statistical power and respond to extending a null A/B test

AI-drafted, machine-checkedSource: cxl.comadvanced
Explain statistical power and respond to extending a null A/B test

This tests statistical power and p-hacking judgment. A strong answer defines power as detecting a true effect, rejects extending the test to chase significance, and requires pre-registered sample sizes. Agreeing to run until it hits significance is a red flag.

WHAT THIS TESTS: The interviewer wants to know if you understand that statistical power is not a post-hoc dial you can turn but a design-time property, and whether you have the statistical maturity to push back on p-hacking requests from stakeholders. They care about your ability to balance business pressure with methodological rigor.

A GOOD ANSWER COVERS: First, define statistical power clearly as the probability that the test correctly rejects the null hypothesis when a specific alternative is true, typically expressed as one minus beta. Second, explain that power is determined before data collection by three levers: sample size, significance level alpha, and the minimum detectable effect or MDE. Third, tell the stakeholder that extending the test after observing a null result is a form of optional stopping that destroys the validity of the p-value by inflating the Type I error rate, so you cannot simply run until significance appears. Fourth, describe the constructive path forward: report the observed effect size and its confidence interval, assess whether the null result is practically meaningful, and if the feature is still strategically important, design a new adequately powered study rather than recycling the current one.

COMMON WRONG ANSWERS: A major red flag is agreeing to extend the test or suggesting that more data will eventually reveal the truth. Another is confusing statistical power with sample size alone, ignoring effect size and alpha. Some candidates incorrectly believe that a post-hoc power calculation justifies continuing the experiment; in reality, observed power is a function of the observed p-value and is not a valid reason to keep running. Finally, framing the issue as purely statistical without translating it into business risk, such as wasted engineering time or polluted decision-making, signals weak stakeholder management.

LIKELY FOLLOW-UPS: The interviewer may ask how you would calculate power for a given MDE, how sequential testing or group sequential methods change the stopping rule, or what you would do if the stakeholder insists and has executive authority. They might also probe whether Bayesian approaches or ROPE-style decision rules could be viable alternatives in this context.

ONE CONCRETE EXAMPLE: Suppose you ran a checkout flow A/B test for two weeks with ten thousand users per variant and observed a conversion lift of one percent with a p-value of 0.2. The stakeholder wants to run it for another two weeks. You respond by showing that the MDE powered at eighty percent was two percent, so the study was never designed to detect a one percent lift reliably. You explain that continuing would give a roughly twenty to thirty percent false positive rate under optional stopping, not the nominal five percent. You then propose reporting the confidence interval of minus 0.5 percent to plus 2.5 percent and running a new test with twenty thousand per variant if the business still believes a one percent lift is worth detecting.

Read the original → cxl.com

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.