WAU is flat despite positive A/B tests; why and how to diagnose

This tests distinguishing real impact from statistical artifacts. Strong answers cite false positives from low base rates, peeking, novelty, and local-global mismatches. Diagnose with long-term holdouts, audits, and causal bridges.
What's really being asked
Your ability to distinguish between statistically significant local effects and genuine business impact at scale. Interviewers care about your fluency in experimental design flaws, statistical power and false positive rates, and causal inference across metric hierarchies. They want to see you treat the flat north star metric as the ground truth and the A/B tests as potentially flawed measurements, not the other way around.
The full answer
Four hypotheses in priority order. First, false positives driven by low base rates of true improvements; if only ten percent of tested ideas actually work, then with standard eighty percent power and five percent alpha more than a third of statistically significant results can be false positives. Second, human bias and sequential testing; teams often let experiments run longer when they are close to significance or stop at a momentary peak, which inflates effect sizes beyond the true sustained lift. Third, novelty effect and user habituation; short-term engagement bumps fade as users adapt to the new feature. Fourth, local to global metric mismatch; A/B tests typically optimize proximal metrics like click-through rate or session frequency, but these do not automatically compound into a distal north star like Weekly Active Users due to cannibalization, funnel leakage, or zero-sum engagement shifts. For diagnosis, propose a concrete workflow: audit the experiment program for selective reporting and peeking, establish a one percent long-term holdout against the shipped feature to measure sustained lift, segment results by user tenure and time since exposure to quantify novelty decay, and build a causal bridge analysis that decomposes why movement in local metrics failed to move the global metric.
The mistakes people make
Blaming engineering instrumentation or data bugs without first interrogating the statistical design. Treating every p-value below five percent as a guaranteed real effect. Claiming the global metric is simply a lagging indicator without articulating the causal chain and expected lag duration. Suggesting to re-run the same tests with larger sample sizes without addressing pre-registration or the base rate of true effects. Ignoring the possibility that local metric gains sum to zero at the global level because users were already active.
What usually comes next
How would you size and maintain a permanent holdout group without sacrificing velocity? When would you use a Bayesian approach versus a frequentist approach given a low base rate of true effects? How do you balance organizational pressure to ship positive tests against the risk of false positives? What would you do if the long-term holdout showed no lift but user surveys reported higher satisfaction?
A concrete example
Imagine a growth team runs twenty notification experiments in a quarter. Only two ideas are genuinely good. At five percent alpha and eighty percent power, they see roughly one true positive and one false positive from the bad ideas alone, yet they ship ten features based on significant results. Post-launch, WAU stays flat because five launches were false positives, three suffered novelty decay within two weeks, and two successfully boosted click-through rate but cannibalized organic app opens rather than creating net new active users. A long-term holdout would have revealed this within thirty days.
Interview question
At 5% alpha, 80% power, and a 10% true-effect base rate, what should a team infer from roughly two significant wins per quarter?
- a.The flat north star metric proves instrumentation bugs are causing false positives
- b.They should treat the wins as real and wait for the global metric to catch up as a lagging indicator
- c.Approximately one of the two significant results is likely a false positive due to the low base rateCorrect
- d.They need larger sample sizes to confirm the wins are real and reduce variance
Why? this is the answer
With a 10% base rate and standard 80% power and 5% alpha, roughly one-third of significant results are expected to be false positives, meaning about one of two quarterly wins is likely spurious. Distractor A repeats the common error of seeking larger samples without addressing the underlying base rate of true effects.
Just read this? Test yourself on what you have been reading.
Read the original → statsig.com
- #growth
- #experimentation
- #statistics
- #causal-inference
- #metrics
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Open roles that interview on growth — each one lists the topics its interview covers.
See open roles