Explain Simpson's Paradox with a user engagement example

This tests your understanding of statistical pitfalls in A/B testing. A good answer defines the paradox, gives an example where a feature fails in aggregate but wins in every segment, and attributes it to a confounding variable.
What's really being asked
This question tests your statistical literacy and practical experience with A/B testing. The interviewer wants to see if you can identify confounding variables that can invalidate experiment results. It's a check for data-driven decision-making maturity and your ability to look beyond top-line metrics, a critical skill for senior roles that own product outcomes.
The full answer
First, a clear definition: Simpson's Paradox is a statistical phenomenon where a trend appears in several different groups of data but disappears or reverses when these groups are combined. Second, the mechanism: it's caused by a lurking or confounding variable that is not distributed evenly across the groups being compared. Third, construct a hypothetical scenario with numbers. For example, testing a new feature. Fourth, show the aggregate result where the feature appears to fail. Fifth, show the segmented results where the feature succeeds in every single segment. Finally, explain the reversal by naming the confounding variable (e.g., user tenure, geography, device type) and showing how its imbalanced distribution created the misleading aggregate result.
The mistakes people make
A candidate who confuses the paradox with other statistical issues like small sample sizes, regression to the mean, or simple volatility. Providing a definition without a concrete, numerical example is a major red flag; the example is the most important part. An answer where the math doesn't actually demonstrate a reversal of the trend (e.g., the trend just weakens). Failing to explicitly name the confounding variable that drives the effect. Suggesting the solution is to just trust the segmented view. The correct approach is to acknowledge the experiment is flawed and must be re-run with proper controls (like stratified sampling).
What usually comes next
How would you prevent this from happening in your experiment design? (Answer: Use stratified sampling to ensure the confounding variable, like user tenure, is balanced across test and control groups). You've run this experiment and discovered the paradox. What do you do now? Do you ship the feature? (Answer: The result is inconclusive. Do not ship. The best course is to design and re-run the experiment, controlling for the confounder. If that's impossible, you might make a decision based on the segmented results but must explicitly call out the high uncertainty and risk).
A concrete example
We test a new "AI Summary" feature for articles. The success metric is "user clicks 'Save Article'".
Aggregate result
Control (no feature): 35% save rate. Test (with feature): 31% save rate. The feature looks like a failure.
Segmented result
The confounding variable is user type. For New Users (who have a naturally low save rate): Control gets 10% save rate, Test gets 12%. The feature wins. For Power Users (who have a naturally high save rate): Control gets 80% save rate, Test gets 82%. The feature wins.
The paradox explained
The experiment had a traffic imbalance. The Test group was composed of 90% New Users and 10% Power Users. The Control group was the opposite: 10% New Users and 90% Power Users. The Test group's aggregate average was dragged down by the high proportion of low-saving New Users, making the feature appear to fail even though it was a success for everyone.
Interview question
An A/B test shows a new feature decreases overall user engagement, but engagement increases within every individual user segment. What is the most likely cause of this discrepancy?
- a.Users in the test group were subject to a novelty effect that will fade over time.
- b.The overall metric is flawed and should be ignored in favor of segment-level results.
- c.The experiment suffered from a confounding variable unevenly distributed across test and control groups.Correct
- d.The observed effect is due to random statistical variance from an insufficient sample size.
Why? this is the answer
This scenario describes Simpson's Paradox, which occurs when a trend reverses between aggregate and segmented data due to an uneven distribution of a confounding variable. Option B is a common misconception, as the paradox indicates a flawed experiment design that requires re-evaluation, not simply ignoring the aggregate.
Just read this? Test yourself on what you have been reading.
Read the original → statsig.com
- #analytics
- #a/b testing
- #statistics
- #data interpretation
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Open roles that interview on analytics — each one lists the topics its interview covers.
See open roles