tezvyn:

WAU is flat despite positive A/B tests; why and how to diagnose

AI-drafted, machine-checkedSource: statsig.comadvanced
WAU is flat despite positive A/B tests; why and how to diagnose

This tests distinguishing real impact from statistical artifacts. Strong answers cite false positives from low base rates, peeking, novelty, and local-global mismatches. Diagnose with long-term holdouts, audits, and causal bridges.

WHAT THIS TESTS: Your ability to distinguish between statistically significant local effects and genuine business impact at scale. Interviewers care about your fluency in experimental design flaws, statistical power and false positive rates, and causal inference across metric hierarchies. They want to see you treat the flat north star metric as the ground truth and the A/B tests as potentially flawed measurements, not the other way around.

A GOOD ANSWER COVERS: Four hypotheses in priority order. First, false positives driven by low base rates of true improvements; if only ten percent of tested ideas actually work, then with standard eighty percent power and five percent alpha more than a third of statistically significant results can be false positives. Second, human bias and sequential testing; teams often let experiments run longer when they are close to significance or stop at a momentary peak, which inflates effect sizes beyond the true sustained lift. Third, novelty effect and user habituation; short-term engagement bumps fade as users adapt to the new feature. Fourth, local to global metric mismatch; A/B tests typically optimize proximal metrics like click-through rate or session frequency, but these do not automatically compound into a distal north star like Weekly Active Users due to cannibalization, funnel leakage, or zero-sum engagement shifts. For diagnosis, propose a concrete workflow: audit the experiment program for selective reporting and peeking, establish a one percent long-term holdout against the shipped feature to measure sustained lift, segment results by user tenure and time since exposure to quantify novelty decay, and build a causal bridge analysis that decomposes why movement in local metrics failed to move the global metric.

COMMON WRONG ANSWERS: Blaming engineering instrumentation or data bugs without first interrogating the statistical design. Treating every p-value below five percent as a guaranteed real effect. Claiming the global metric is simply a lagging indicator without articulating the causal chain and expected lag duration. Suggesting to re-run the same tests with larger sample sizes without addressing pre-registration or the base rate of true effects. Ignoring the possibility that local metric gains sum to zero at the global level because users were already active.

LIKELY FOLLOW-UPS: How would you size and maintain a permanent holdout group without sacrificing velocity? When would you use a Bayesian approach versus a frequentist approach given a low base rate of true effects? How do you balance organizational pressure to ship positive tests against the risk of false positives? What would you do if the long-term holdout showed no lift but user surveys reported higher satisfaction?

ONE CONCRETE EXAMPLE: Imagine a growth team runs twenty notification experiments in a quarter. Only two ideas are genuinely good. At five percent alpha and eighty percent power, they see roughly one true positive and one false positive from the bad ideas alone, yet they ship ten features based on significant results. Post-launch, WAU stays flat because five launches were false positives, three suffered novelty decay within two weeks, and two successfully boosted click-through rate but cannibalized organic app opens rather than creating net new active users. A long-term holdout would have revealed this within thirty days.

Source: statsig.com

Read the original → statsig.com

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.